Alice believes that the bridge is closed.
That sentence looks like “Alice weighs seventy kilograms,” and it is not built the same way. Weighing her recovers a quantity that maps fairly directly onto a physical description of her body. Inspect her brain and we do not find a single state labeled belief that the bridge is closed. We find neural activity, memories, perceptual representations, learned associations, action tendencies, and perhaps some explicit representation concerning the bridge. The belief appears only once those facts are organized under a particular model of Alice.
I have argued elsewhere that this is the right way round. Belief, desire and intention are latent variables in a model of an agent’s causal organization, objective to the extent that the organization constrains which models succeed. The Intentional Gradient makes that case and defends it against the levels objection, the lookup table, and the demand for intrinsic aboutness.
Two complications sit outside that account. Alice may herself contain models: she models the world, may model her own internal states, and may explicitly represent what she thinks she believes. Once the referent is itself a modeler, the distinction between map and territory becomes recursive. And a model of Alice can reach Alice. She can learn what she has been classified as and change accordingly, at which point the model’s continued success stops being straightforward evidence that the classification was right to begin with.
Three locations of structure
Some properties belong plainly to the representation. A map may be printed on A4 paper; Paris is not. A statistical model can be overfit, expensive to run, or expressed in a particular coordinate system. Those are model-side properties.
Other structure is tightly constrained by the referent. A processor contains transistors in particular physical states; a glass contains molecules in a particular configuration. These descriptions still depend on theory and measurement, so calling them metaphysically “intrinsic” buys little. The system itself sharply constrains which descriptions survive contact with observation.
A third class is indexed to the modeling scheme. A gas has a temperature, a population has a fitness distribution, an economy has an inflation rate. None of these appears as a microscopic label attached to the constituent parts, and each depends on a particular abstraction, coarse-graining, or statistical construction. Yet a physicist cannot assign whatever temperature she likes to a gas and defend the choice by saying that temperature belongs to her model. Model dependence is not arbitrariness.
Belief sits in the third class, and the analogy stops carrying weight sooner than it appears to. Thermodynamics has a degree of mathematical and empirical convergence that intentional psychology may never achieve. Several models of Alice may predict her equally well while partitioning her internal organization differently: one representing her as holding belief P, another using a latent decomposition with no variable corresponding neatly to belief at all, two intentional models disagreeing about where one belief ends and the next begins. That does not make belief incoherent. It sets the burden: weaker ontological commitment where successful models diverge, stronger where independent models and interventions recover approximately the same structure.
An old problem with a new edge
Dennett’s intentional stance treats beliefs and desires as predictive explanatory posits, and his account of “real patterns” allows that a pattern can be real without being a primitive object in the lowest-level physical description, provided it supports genuine compression and prediction. The argument here inherits that and puts more weight on underdetermination and recursion. Competing models may carve the same agent differently, and the referent may itself contain models of the very predicates we are using. Once that happens the relation between model and referent becomes causal, not merely descriptive.
Ian Hacking’s “looping effects” named the phenomenon: classifications of people can change the people classified. Intentional predicates are often available to their own referents. Alice can hear that she is risk-averse, incorporate the classification into her self-model, and behave differently afterwards. Belief is a useful case for extending model-based realism into a setting where models can enter the systems they describe.
What “Alice believes P” compresses
Ordinary belief talk hides several claims under one predicate. Alice might contain a representation of P, say that P is true when asked, use P as a premise in reasoning, behave in ways well predicted by treating P as true, and revise her actions when presented with evidence against P. Her self-model might also classify her as someone who believes P.
These properties often travel together, which is why ordinary language can ignore the distinctions. If Alice hears that the bridge collapsed, says that it is closed, stops planning routes across it, warns other people away, and reverses all of that when credible evidence arrives that a temporary crossing has opened, the compressed sentence works extremely well.
They can also separate. Alice may explicitly endorse the importance of exercise while repeatedly sacrificing it to trivial conveniences. A political belief can be stated by someone whose predictions are incompatible with it. An internal representation of P may sit largely isolated from the systems responsible for action, and an agent may deny believing something that nevertheless predicts her behavior unusually well.
So the binary predicate believes discards structure. A more faithful model would treat belief as multidimensional: representation, endorsement, inferential use, behavioral control, reportability, cross-context stability, sensitivity to evidence, self-attribution. Nothing guarantees that these dimensions align, and ordinary belief attribution works because, in ordinary cases, enough of them do.
Representing P is not the same as representing yourself as believing P
An agent can represent P without representing itself as believing P. A dog may behave as though food is behind a door without entertaining the proposition “I believe food is behind the door.” A human may rely on a tacit assumption without ever formulating it. An artificial system may contain a world-model state corresponding to P while lacking any metacognitive representation of itself as a believer.
Self-attribution is a second-order capacity layered on first-order world representation, and building it into the definition of belief would exclude many plausible cases of belief-like cognition. The recursion begins only when the agent models its own modeling. Alice may represent that the bridge is closed; she may also represent that she believes the bridge is closed. Those are distinct internal structures with different causal roles.
The model may be inside the agent
Suppose Alice explicitly represents:
I believe the bridge is closed.
That representation is part of Alice, with some physical realization in her nervous system. An artificial agent could make it more visible by storing bridge_closed = 0.93 alongside a second-order entry such as self.beliefs["bridge_closed"] = true. It looks as though the question is settled: we wanted to know whether the belief was merely in our model or inside Alice, and here is an explicit belief representation inside the agent.
Now suppose that representation is stale, strategically generated, or weakly integrated with the rest of the system. Alice reports that she believes the bridge is closed while her navigation continues to plan routes across it. The artificial agent stores goal = X, consults it in some contexts, and overrides it through another subsystem whenever X conflicts with a stronger learned policy.
Humans supply the pattern in bulk. A person can confabulate reasons for her choices, misidentify her preferences, or maintain explicit commitments that do little work in actual decision-making. Alice can sincerely say that she prefers tea while choosing coffee in nearly every stable context, paying more for it and becoming disappointed when only tea is available. Introspection gives unusually direct access to some internal states, and a self-model is still a model that can misrepresent its referent.
So self-representation and model-supported attribution come apart. The first is a property of the agent’s internal state; the second is a judgment about which higher-level description best captures the organization of the agent as a whole. Locating an explicit representation strengthens the evidence without settling the ontology, and the strength of an attribution depends on how much of the system is organized around it, under which conditions. When the two diverge, the divergence is itself something to explain.
When multiple models disagree
The objection to face is several successful models explaining the same agent with different latent structures. One model predicts Alice by assigning her belief P and preference Q; another predicts her equally well using latent variables that do not resemble beliefs or preferences at all.
Then ontological commitment should weaken. A property that appears only under one convenient factorization may be useful without being deeply attributable to the referent, because predictive success alone does not guarantee that a model’s internal vocabulary maps onto independently stable structure.
The case strengthens when the same organization reappears across multiple successful descriptions. Different models trained on different data recover approximately the same distinctions. Internal measurements correlate with the same latent variable. Interventions on corresponding structures alter downstream behavior in predicted ways. The attribution generalizes to contexts that were not used to construct the model.
Internal evidence discriminates further, because two agents can produce the same outputs for different reasons. If a latent state correlates with particular reports, predictions and actions, and intervening on that state systematically changes the downstream pattern, confidence in the attribution increases, on the same intervention test that licensed intentional vocabulary in the first place. Internal localization is not mandatory; it is one more way to check whether a model tracks stable organization or fits a surface regularity.
Convergence is strong evidence only when it is independent. Five models that inherit the same conceptual vocabulary, training data, assumptions or measurement pipeline do not provide five discoveries. Nor does convergence tell us much if all five have altered the referent in the same way and then rediscover the structure they jointly helped produce. The criterion is independent convergence under varied models and interventions: the more a higher-level property survives changes in representation, data source, method and causal pathway, the stronger the warrant for attributing corresponding structure to the referent.
Objectivity here does not require scientific consensus. Researchers may disagree because the evidence is incomplete, the models are immature, or the system is underdetermined. Those failures reduce our warrant for a particular attribution. They do not show that the referent imposes no constraints.
The recursion is causal
Once self-models enter the picture, several levels interact: Alice as a physical agent, Alice’s model of the world, Alice’s model of herself, our model of Alice, our model of Alice’s world model, our model of Alice’s self-model. Ordinary language slides among these because in familiar cases they agree well enough.
They can also change one another. If Alice comes to think of herself as someone who hates public speaking, the description may alter which invitations she accepts. Her resulting behavior supplies new evidence to observers, who update their models of her, and their responses feed back into her self-model.
So the act of modeling can modify the referent, and it runs in both directions. A confirming loop inflates apparent fit. The description takes hold, the agent organizes behavior around it, and the model looks better calibrated than it was when it was made.
A defeating loop looks like the opposite. Alice can reject a description and train herself away from the pattern it captured. She can discover that her stated preference differs from her revealed choices and reconcile them in the direction of the statement rather than the choices. The model then stops predicting. The natural reading is that it was overfitted; the other is that it was right and the referent reacted, which makes the failure evidence about Alice rather than about the model. Nothing observed after the description was delivered separates the two, and a model that predicted well and then stopped is not automatically a model that was wrong. The relation between representation and referent is a feedback loop rather than a one-way act of description.
What the loop requires
Reflection looks like the ingredient that makes this possible, and it is not. The loop needs a causal channel from description to referent that is sensitive to descriptive content: change what the model says, and the referent changes differently.
That condition separates two things measurement routinely confuses. A thermometer warms the liquid it measures, and quantum measurement disturbs the system measured, but neither depends on what the description says. Swapping in a different model with identical apparatus leaves the perturbation unchanged. Probe disturbance destroys the baseline without manufacturing evidence. You may have no access to the unmeasured system, which is a provenance problem of its own, but nothing has been made true by having been asserted.
Content-sensitive coupling is the other case, and it is common. Goodhart’s law has this structure: a measure adopted as a target stops measuring what it measured, because the content of the measure is what gets optimized against. The Lucas critique is the same point about econometric models, whose estimated relationships shift once policy is set from them. Hacking’s looping effects are the version that runs through the classification of people.
Nor does the loop require anything that would pass the intentional gradient’s tests. An adaptive controller maintains a model of the plant it regulates and sets its outputs from that model. The plant’s dynamics change because of what the model says about them, and the controller then observes a plant its own description helped shape. No comprehension, no self-representation, no understanding of any kind. It is a standard engineering configuration with the full provenance pathology.
Agency can close the loop. It is neither necessary nor sufficient, since an agent can be handed a description and ignore it. The requirement is content-sensitive causal coupling, and self-representation, comprehension, endorsement and reflection are all dispensable. A description has to be causally effective, not understood.
What agency changes is where the loop closes. Detach the adaptive controller and the plant can in principle be observed unmodeled, because the loop ran outside the referent. With a reflective agent the relevant controller is part of the referent, so removing it changes the system whose unmodeled baseline we were trying to recover. That governs whether a baseline is recoverable in principle, not whether the phenomenon occurs.
When the model manufactures its evidence
Performativity is a problem for any realism grounded in predictive success. A model can succeed because it discovered structure already present in the referent. It can also succeed because exposure to the model caused the referent to acquire that structure.
A psychologist classifies Alice as highly risk-averse. Alice takes the diagnosis seriously, incorporates it into her self-model, begins avoiding uncertain situations, and behaves exactly as the original model predicted. That success does not establish that the attribution described Alice before she encountered it. The model may have manufactured its own evidence.
What is missing is provenance: whether the structure preceded the classification, arose independently of it, or was partly produced by the act of classification. Longitudinal evidence, observations taken before exposure, interventions on the self-description and persistence after the description is withdrawn can help separate discovery from construction.
For reflective agents provenance may be only partially recoverable, because the loop has been closing inside them from the start. Humans encounter models of themselves from infancy: family expectations, personality categories, reputations, diagnoses, and ordinary descriptions such as “shy,” “clever,” or “difficult.” By adulthood much of the structure we are trying to model may already be the accumulated result of previous modeling.
The causal distinction survives; the evidence needed to reconstruct the history may not. Provenance constrains what we can claim rather than promising that every attribution decomposes cleanly into discovered and constructed components.
“Does Alice now behave as a risk-averse agent?” can have a well-constrained answer even when “Was Alice already risk-averse before being classified that way?” has a different and perhaps unrecoverable one. A model can correctly describe a structure it helped create, provided we do not confuse present fit with evidence of prior existence. For reflective agents, the history of a property can matter as much as its current fit.
When a property crosses from model to referent
Scientific language routinely takes variables introduced in models and speaks as though the corresponding properties belong to the modeled system, sometimes as harmless shorthand and sometimes because the projection reflects robust convergence. The modeler chooses a representational scheme; the referent determines whether it succeeds. A property stable across measurements, alternative formulations, contexts and interventions warrants stronger ontological commitment than one that appears only under a narrow construction.
Apply this to Alice and coffee. Modeling her as preferring coffee to tea predicts her choices across cafés, prices, presentation orders and social settings. It predicts what she will sacrifice to obtain coffee, how she responds when coffee is unavailable, and how her behavior changes after relevant new information. Independent models recover approximately the same preference structure, and internal evidence points toward corresponding organization.
At that point “Alice prefers coffee” is more than convenient shorthand. It is a compact attribution constrained from both sides: the model supplies vocabulary and abstraction, and the referent determines whether that vocabulary keeps working under observation, intervention, alternative description and provenance analysis. The objectivity lies in the constraints Alice imposes on successful models of Alice, together with our ability to tell independent convergence from shared assumptions, and pre-existing structure from structure produced by the modeling.
A concrete AI case
The weaker criterion is what makes the AI case immediate. A language model is trained partly on human descriptions of language models, and post-training and system instructions further shape how it describes its own capabilities, limitations and goals. Those self-descriptions then enter future datasets, evaluation regimes and training decisions. The training pipeline is a content-sensitive channel: change what we say about these systems and what gets built changes differently. None of that requires the system to understand the descriptions, or to be an agent at all. The descriptions only have to be data. Routing the argument through reflection would beg the question, making the loop a property of agents and then applying it to systems whose agency is the matter in dispute.
So a system saying “I am merely a language model” or “I want X” is not transparent access to an underlying self-model. The statement may reflect internal organization, learned discourse about systems of its kind, explicit behavioral training, or some mixture of the three.
The feedback also sits a level higher than in Alice’s case. The descriptions need not concern the particular system that later instantiates them, because descriptions of the class enter the process that constructs later members of the class. Alice reads a model of Alice, and the loop closes within one referent. The AI case is a lineage-level loop, in which models of earlier systems help shape their successors.
That makes provenance harder rather than easier. When a psychologist classifies Alice we can at least ask what she was like beforehand, since the classification arrives after the agent exists. For a model trained on the accumulated public discourse about models, the relevant descriptions predate the referent entirely, and there is no unmodeled predecessor state to compare against.
Not every attribution to such a system stands or falls together. A model that answers that Paris is the capital of France across paraphrases, routes a trip plan accordingly, integrates the fact into further inference, and holds the answer stable under adversarial framing and under intervention on the relevant internal states gives a first-order attribution something to work with. “I am a language model” is a different kind of claim. It is self-representation, and its provenance is known to be heavily shaped: it is the output most directly targeted by post-training, and the descriptions it echoes were in the corpus before the system existed. Reading the second as a window onto the first is the error the recursion makes easy.
The intentional gradient scales with how much causal structure a description compresses, and it assumes the structure was there to be compressed. For any referent coupled to its own description, that assumption is the part to check.
Asking whether Alice “really believes P” does not guarantee a single microscopic fact waiting to be uncovered. It asks whether an intentional description captures stable enough structure across a recursively modeling agent, whether that structure survives alternative models, interventions, contexts and independence tests, and where the structure came from. The last question is usually absent from passive thermodynamic description. It becomes unavoidable whenever a model participates causally in producing the structure it later describes.


