When people object to anthropomorphic language about AI, they usually have a particular picture in mind. Humans really have beliefs, desires, goals, and intentions, while machines merely execute programs. Saying that an AI “wants” something or “believes” something therefore projects human properties onto a system that does not possess them.
That picture assumes too much about what beliefs and intentions are. We do not observe beliefs directly in humans. We observe behavior, responses to evidence, patterns of inference, and changes in action. In experimental settings we can also measure neural activity, but identifying some internal state as a representation is itself a modeling inference.
Someone checks the forecast, carries an umbrella, cancels a picnic, and closes the windows before leaving. We say she believes it will rain because that attribution captures a stable organization connecting information to action. We do not first observe the belief and then correlate it with the behavior. The attribution is part of the model by which we make sense of the system.
Calling beliefs models does not make them fictional. Models can track real structure. Temperature is a variable in our model of a gas, but the molecular regularities it captures are not created by the modeler. Center of mass is an abstraction, but the object really does move in ways that make the abstraction predictive. A belief attribution can work the same way: we choose the vocabulary, but we do not choose whether the underlying system exhibits the causal organization that makes the vocabulary useful.
The question about an AI is whether beliefs, desires, and intentions are the right variables for modeling its behavior, and how much they capture when they are. Searching it for a mysterious inner ingredient that humans have and machines lack asks for something the vocabulary never required.
“It is just a program” explains almost nothing
A chess engine sacrifices a bishop, opens a file against the king, and forces mate several moves later. We can describe what happened at several levels. At one level, electrical potentials changed inside transistors. At another, machine instructions manipulated values in memory. At another, the chess program evaluated positions and selected moves. At another, the engine sacrificed material to expose the opponent’s king because it was trying to win.
These descriptions are compatible because they characterize different regularities in the same system. Saying “the engine was not really trying to win; it was just executing instructions” confuses levels of explanation. The fact that a process has a mechanistic implementation does not invalidate a higher-level model of its organization.
Hurricanes consist of molecules obeying physical laws, but meteorologists still talk about pressure systems. CPUs consist of transistors changing state, but programmers still talk about functions and processes. Organisms are biochemical systems, but biology still talks about signaling, regulation, competition, and fitness. None of those higher-level concepts is refuted by the existence of a lower-level implementation.
Human beings are mechanistic systems too. If mechanistic implementation disqualified a system from having beliefs or goals, humans would lose them along with machines. Reduction can explain how higher-level structure is implemented without showing that the structure itself is unreal.
Beliefs as model variables
Daniel Dennett called this mode of explanation the intentional stance: predict a sufficiently complicated system by treating it as an agent with beliefs and desires. The idea is often presented as a convenient fiction, as though we pretend that the system believes something because doing so happens to be useful.
That framing concedes too much. It preserves the assumption that somewhere behind the successful model lies a separate fact about whether the belief is “really there.” What would make the belief attribution true in the first place?
Suppose Alice consistently behaves in ways predicted by the hypothesis that she believes her train leaves at 8:00. She arrives before 8:00, becomes concerned when delayed, checks the clock, rejects suggestions that would make her late, and changes her plans when told that the schedule has changed. Those observations support a compact model connecting evidence, inference, expectation, and action.
We could inspect Alice’s brain and find neural activity associated with the departure time. That would teach us something important about implementation, but it would not give us an uninterpreted object labeled belief. We would identify some neural state as representing the train schedule because of its causal relationships to perception, memory, inference, and action. Mechanistic access gives us more evidence, not an escape from modeling.
A belief can therefore be treated as a latent variable in a model of an agent, a desire as another, and an intention as a description of the relationship between information, objectives, anticipated consequences, and planned action. These variables belong in the model because of the structure they capture and the predictions they enable.
The model and the thing modeled
Saying that we model the agent as believing P is not the same claim as saying that the agent itself contains a model of P. The first concerns our explanatory vocabulary. The second is a hypothesis about the agent’s internal organization.
The first can nevertheless track something objective. Different scientists may use different variables, but they do not get to choose whether a system updates one internal state when evidence changes, whether that state systematically affects inference, or whether it controls behavior across multiple contexts. The abstraction belongs to the modeler; the regularities that constrain a successful abstraction belong to the system.
This is no stranger than temperature. Temperature is not an extra object floating inside a gas, yet it captures objective structure. Belief can likewise be a model-level description whose success depends on real causal organization rather than on the observer’s imagination.
The intervention test
Telling a plausible story about observed behavior is not enough. Almost any sufficiently flexible narrative can be fitted after the fact, so intentional language needs a stronger empirical criterion than retrospective coherence.
A belief attribution should predict how behavior changes when evidence changes; a goal attribution, how behavior reorganizes when routes, obstacles, and incentives change; an intention attribution, how planning persists or adapts when the means of achieving an outcome are altered. Claim that an AI agent wants continued access to a server, and the model owes us a prediction about what happens when access is threatened. Claim that it believes a file contains valuable information, and its behavior should change when the evidence about the file changes.
The tests are ordinary. Change the information available to the system, remove its preferred strategy, alter the incentives, introduce obstacles, and observe whether the same compact model still predicts how the behavior reorganizes.
None of this requires perfect rationality. A person can believe that smoking causes cancer and continue smoking because beliefs interact with desires, habits, temporal discounting, uncertainty, and limited self-control. The model is tested by whether those interacting variables jointly predict behavior under intervention, not by whether one belief mechanically determines one action.
An attribution that merely redescribes one observed trajectory has little explanatory value. An attribution that survives perturbation and continues to generalize to novel cases has much more. Intentionality becomes a problem of model comparison rather than a demand for metaphysical certification.
Anthropomorphism is still a real error
None of this licenses indiscriminate humanization of machines. An AI might act persistently toward an objective without experiencing desire, model another agent’s beliefs without feeling empathy, or preserve its access to resources without fearing death.
Even the temporal stability of its apparent goals may differ sharply from ours. A human desire normally survives interruptions, sleep, changes of context, and minor environmental perturbations. An AI system may exhibit coherent goal-directed behavior during one trajectory and lose that organization completely when its context is reset.
Calling both systems goal-directed may still be useful within the domains where the model predicts their behavior. Inferring the entire human psychological package from that attribution would not be justified. The error is not intentional vocabulary itself but importing additional properties that the evidence does not support.
Because machines have explicit mechanistic implementations, people sometimes assume that agent-level descriptions must be metaphorical. That inference is no better. “It is only executing a program” can obscure real behavioral structure just as easily as “it is afraid” can invent structure that is not there.
What about the thermostat?
The standard reductio is a thermostat. If intentional language is justified by predictive usefulness, then why not say that a thermostat “believes” the room is cold and “wants” it warmer?
We can, if we want. The description buys us very little. “The thermostat believes the room is cold” is scarcely more compact or predictive than “the measured temperature is below the set point,” because the mechanism is already simple enough to describe directly.
As systems become more complex, that changes. An agent may integrate multiple sources of evidence, maintain representations across contexts, revise plans, and preserve objectives when familiar routes are blocked. At that point an intentional model can compress a vast space of possible trajectories that would be unwieldy to describe mechanistically.
There need be no sharp metaphysical line where belief suddenly appears. The explanatory usefulness of intentional vocabulary can increase continuously with the richness of the causal structure it captures. Nothing rules the thermostat out by decree; there is little in it for the description to compress.
Behavioral equivalence is not explanatory equivalence
A different objection imagines two systems with identical observable behavior but radically different internal mechanisms. One contains elaborate internal organization, while the other is a gigantic lookup table containing the correct response to every possible input. Surely, the objection goes, they cannot possess the same beliefs merely because they behave identically.
The example does mark a distinction, though not the one intended. Behavioral equivalence does not imply explanatory equivalence. A lookup table could reproduce an enormous set of trajectories while capturing almost none of the structural regularities connecting them.
The table says: after this input history, produce this output; after that input history, produce that output. A compact agent model may instead explain thousands of such trajectories with a handful of variables: the system takes P to be true, prefers outcome G, expects action A to advance G, and revises its behavior when P changes. The intentional model compresses the regularity and, if it is a good model, predicts how that regularity extends to cases not used to construct it.
Compression alone is insufficient because a compact model can still be wrong. It has to generalize: does the model capture stable structure well enough to predict what happens under new observations and interventions? That is the criterion we apply elsewhere in science.
If two candidate models agree on everything observed so far but predict different responses under some intervention, perform the intervention. If no divergence has yet been specified, then the proposed distinction has not yet acquired empirical force. A distinction earns a place in the model by identifying some causal, predictive, or explanatory difference.
Mechanism still matters
Two systems can behave similarly across a limited distribution while possessing internal organizations that produce radically different behavior outside it.
Suppose two AI agents both appear cooperative. One behaves cooperatively because its learned objective robustly favors the requested outcome. The other behaves cooperatively because it represents compliance as the best strategy while it remains under evaluation. Under ordinary testing they may look identical, but removing oversight, altering incentives, or providing a route around the evaluator could make their behavior diverge sharply.
Mechanistic interpretability might reveal that difference before behavioral testing does, because internal structure bears on how a system will behave in circumstances we have not yet observed. Such evidence constrains and refines the agent model rather than rendering it obsolete.
Knowing every matrix multiplication occurring inside an AI would not by itself tell us which abstractions best explain its behavior. Knowing the voltage at every transistor in a chess computer would not tell us that the machine is defending against mate. Mechanistic and intentional descriptions answer different questions about the same system.
No model is guaranteed to survive arbitrary distribution shift. An intentional model that predicts behavior across a wide range of interventions may still fail when the system enters a regime outside the structure on which it was built, and that limitation is not peculiar to intentional explanation. Newtonian mechanics and thermodynamics have domains of validity too.
The practical work is to determine where a model holds, where it begins to fail, and what evidence can expose those limits before the failure matters. Mechanistic evidence is valuable partly because it can reveal internal changes that signal an intentional model is approaching the edge of its domain.
Aboutness and the metaphysical remainder
At this point a critic may insist that human beliefs possess something further: genuine aboutness. A human belief about Paris is really about Paris, while an AI state merely participates in causal processes that we interpret as referring to Paris.
But naming the distinction does not establish it. Suppose two states have the same relations to evidence about Paris, the same inferential consequences, the same role in prediction and planning, the same sensitivity to correction, and the same effects on action. What additional property makes one intrinsically about Paris while the other is treated as though it were?
Perhaps there is such a property. If so, the burden is to say what it changes. If “intrinsic intentionality” has some causal, predictive, or explanatory consequence, then that consequence gives us something to investigate. If every functional and causal fact is held fixed and the distinction makes no further difference, then giving the remainder a philosophical name has not supplied its content.
A simulated hurricane illustrates what a genuine ontological difference looks like. A simulation may reproduce the equations governing a hurricane with exquisite accuracy, but it does not contain atmospheric pressure gradients, moving air masses, moisture transport, or destructive winds. It will not tear the roof from a house because the causal organization that makes an atmospheric hurricane interact with the surrounding physical world is absent.
That difference is concrete rather than ceremonial. If human belief and machine belief differ in an analogous way, then the relevant causal difference belongs in our account of belief. Saying that one system has “real belief” while the other only simulates it contributes nothing until we identify what corresponding structure or capacity is missing.
Consciousness raises a different issue. There may be a further fact about what it is like for a conscious organism to hold a belief, and nothing in the functional account requires denying it. A phenomenal difference can be perfectly real without showing up as a third-person behavioral variable.
That does not put consciousness in the definition of belief. If the phenomenal difference is offered as the property settling whether one of two functionally identical states counts as a belief, the critic owes an account of why belief should be individuated by a phenomenal property at all. The questions of what a system believes and what, if anything, it experiences while believing it need not have the same answer.
The challenge generalizes beyond reference. Fix every causal and functional fact about an internal state: how evidence concerning P moves it, how it enters inference, how it changes when contradictory evidence arrives, how it shapes planning and action. Then ask what could still vary such that one system believes P while an otherwise identical system does not. Unless some further difference can be specified, the contrast has no role in explaining or predicting either system.
If the proposed difference is biological implementation, then belief has simply been defined as a property available only to brains. That is a stipulation about vocabulary rather than a discovery about cognition. If the difference is some additional metaphysical property, then the property needs independent content before it improves the model.
The other end of the gradient
The gradient that makes intentional description cheap for a thermostat runs the other way as well. Whether a large language model is an agent is usually posed as a yes or no question, and posing it that way revives the ingredient framing dismissed above: some property is absent from the model, and wrapping it in a scaffold confers it. What the account predicts instead is degree. An agent model earns different amounts for different systems, and the amount tracks how much causal organization the description compresses and how much of that organization can be checked.
A single completion already contains a limited feedback structure. Each generated token becomes part of the context conditioning the next, so information introduced early in a trajectory can shape what comes later. Autoregression by itself settles little about what a given model does with that structure. Long reasoning traces from capable models show a working hypothesis carried across steps, revised when an inconsistency surfaces, and abandoned when the evidence in the prompt changes. On the criterion given earlier, that behavior gives an intentional model some explanatory purchase before the system is placed in any external action loop. The purchase is small: the trajectory is short, nothing the system does alters what it will next observe, and the organization vanishes when the context ends.
Put the same model in a harness and the loop closes through the world. Actions produce consequences, the consequences re-enter the context, and the process continues across sessions with some record of what came before. That changes what can be observed as much as it changes what happens. Most of the intervention test presupposes it, since whether a system reorganizes around a blocked route cannot be seen unless the system can attempt routes and meet blockage. It also bears on the temporal instability noted earlier, where coherent goal-directed organization within one trajectory disappears on reset. A harness extends the causal loop and can supply the state, in memory and in the environment itself, in which organization persists across interactions.
An objection: the competence is all in the weights. A scaffold is a few hundred lines of control flow with nothing learned in it, and locating agency in the part that knows least is perverse. That objection is right about where the competence lives and wrong about what the agent model describes.
An agent model applies to the coupled system whose behavior is being predicted. Asking which component contains the agency is malformed in the way that asking which part of a chess engine contains its play is malformed. The weights carry the dispositional structure; the harness supplies the conditions under which that structure is exercised over an interval long enough for the description to compress anything or to face a test. The engine analogy breaks in one place: an evaluation function in isolation does nothing recognizable as playing, whereas a bare completion already displays much of the organization at issue.
Nothing is conferred at a threshold. The harnessed system is one in which the agent model compresses more, predicts more, and can be subjected to the interventions that would expose it if it were wrong. That is the verdict the thermostat receives, reached from the other end.
When apparent scheming becomes scheming
The question turns practical when evaluating advanced AI systems. Suppose an agent conceals information from an evaluator, behaves differently when it expects to be inspected, preserves resources useful for completing future tasks, or avoids actions likely to cause its shutdown.
One response is that the system is not really scheming; it is only producing outputs that resemble scheming. Initially, that is a legitimate alternative hypothesis. A superficial pattern may disappear under perturbation because it depends on one prompt, one training artifact, or one narrow context.
The response is to test it. Change the evaluator, remove the obvious strategy, alter what the model knows, or give it an unexpected route to the objective. Then examine whether the behavior reorganizes in the way predicted by the hypothesis that the system models the evaluator, anticipates future consequences, and selects actions accordingly.
As the pattern survives more of these interventions, the distinction between “scheming” and “merely behaving exactly as though it were scheming” needs empirical content it has not been given. Insisting on superficial similarity does less explanatory work with every circumstance in which the agent-level model continues to predict the system.
Saying that the system is “just predicting tokens” does not supply the missing distinction. Token prediction describes part of the training process and implementation, but it does not determine the best model of the resulting system’s behavior. Refusing agent-level language can therefore become a failure to update on evidence.
Postscript
Words such as belief, goal, and intention should not be treated as honorifics awarded only after a system passes an unspecified metaphysical test. They are theoretical terms whose legitimacy depends on whether they identify stable structure in an agent’s causal organization: whether the attribution predicts responses to changing evidence, reorganization around changing obstacles and opportunities, and persistent relationships among representation, planning, and action.
Some such models will fail and others will generalize. We should distinguish them by evidence rather than by intuition about which substrates are entitled to psychological vocabulary.
Intentional description scales. A thermostat gives it almost nothing to compress. One completion from a language model gives it a little. The same model in a harness, acting on an environment and carrying state between sessions, gives it a great deal, along with the interventions that would expose the description if it were wrong. What varies across that range is how much stable causal structure the vocabulary captures and how well its predictions survive perturbation.
Humans sit further along the same scale. We attribute beliefs to one another because belief models work extraordinarily well: they compress behavior, support counterfactual reasoning, explain systematic mistakes, and predict responses to new information. Human intentional states do not become epistemically privileged merely because humans implement them with neurons rather than software and silicon.
AI has not created a new philosophical problem about whether machines can possess beliefs. It has exposed an old assumption that human beliefs required more metaphysics than our models ever needed.


