In October 2025, I published From Correlation to Counterfactuals, arguing that language models had acquired a capacity for causal reasoning that Judea Pearl, in The Book of Why (2018), placed beyond purely statistical learners. GPT-5 handled Pearl’s firing-squad counterfactuals correctly, distinguishing overdetermination from preemption and causal responsibility from moral culpability.
Then I asked how it worked. GPT-5 said modern reasoning models contain a hybrid layer that builds structural causal models and executes interventions under do-calculus semantics. I put that in the article as fact.
The Invented Architecture
The article stated:
Modern systems no longer rely solely on pattern prediction. They now contain a hybrid reasoning layer: a symbolic interpreter capable of building structural causal models (SCMs) on demand.
There was no evidence for it. The account came from GPT-5, and I didn’t check it.
A language model has no privileged access to its own implementation. Its descriptions of architecture are generated outputs, not observations of its machinery. GPT-5’s account was plausible, detailed, and phrased as an engineer would phrase it. Plausibility is not architectural evidence. The detail should itself have been a warning. Do-calculus derives interventional distributions from a causal graph and observational data. The firing squad supplies its whole structure, and its counterfactuals are answered by evaluating that structure under the hypothetical change, so GPT-5 named a procedure the problem did not call for.
I retract the claim. Some neural computation may be functionally equivalent to a structural causal model, but that would not establish the symbolic interpreter the article described.
The Introspection Trap
Nisbett and Wilson showed in 1977 that people readily give coherent reasons for choices whose causes they cannot report. The name for this is confabulation: a sincere explanation with no access to the thing explained.
Language models confabulate in a harder-to-catch way. A model can describe its architecture accurately without inspecting itself, because it read the documentation or was told in context. Accurate descriptions of real components and invented descriptions of nonexistent ones arrive in the same confident register, and only external evidence separates them. A mechanism the model proposes is a hypothesis that cannot certify itself.
I also wanted this one. I believed Pearl that machine learning was stuck on the first rung, so GPT-5’s success was an anomaly, and the invented architecture resolved it without costing me Pearl: statistics alone could not do it, so something non-statistical had been added. The original article called the development a completion of Pearl’s framework. But Pearl had moved on two years earlier. In a 2023 interview with Dana Mackenzie, he said the ladder restrictions no longer apply to models trained on text, because text carries causal information. I was rescuing a position its author had abandoned, with a mechanism nobody had shown.
What the Experiment Showed
GPT-5’s answers were consistent with a structural causal analysis. That supports a claim about competence, not implementation.
It also concerns reasoning over a causal model, not discovering one. Two models can agree on every correlation and still predict different effects of an intervention, and the firing squad hands over its structure: an order causes two soldiers to fire, and either shot kills. Corr2Cause tests which causal conclusions a set of correlations licenses, and the models its authors evaluated performed close to chance.
Nor was the problem novel. Pearl put the same questions to ChatGPT in 2023 and got the right answer after careful prompting. He then predicted that on another problem with the same structure, the model would need prompting from scratch: “It won’t generalize.”
I tested that, though the article didn’t say so. Another model generated twenty abstract mechanisms for me: colored emissions, gates and converters acting on a single variable, each a textbook causal motif (preemption, preventers, opposing mediators, overdetermination, latching) with the story stripped out, and a counterfactual question about each. That is Pearl’s test: familiar structure, unfamiliar problem. GPT-5 matched the answer key on all twenty. I never checked the key either, and on review after grading it is wrong on one item and underdetermined on another. Agreeing with the wrong key is an error, so the result is eighteen correct, one wrong, and one question with no determinate answer. I didn’t publish the tests, and twenty model-generated cases graded against a model-written key are not a benchmark, so the article rested on its weakest evidence and stayed silent on the one claim of Pearl’s that evidence could settle.
The Argument That Survives
The original article argued that systems trained on prediction need not stay confined to association. Pearl’s 2023 concession settles the information side: text written by people with causal models carries causal information. The live dispute is what the model does with it.
Pearl’s answer, in his 2024 preface to the French edition of The Book of Why, is that LLMs only give the impression of a causal model, having “merely quoted word sequences” from their training texts. In the same preface he left open whether faking a causal model differs from having one in a machine that can store trillions of sequences. He did not leave open which of the two the models were doing, and that second claim is about mechanism, with no evidence of the mechanism behind it. GPT-5 ruled causal machinery in from the inside; Pearl rules it out from the outside. Neither examined the system.
Behavior can press on Pearl’s account without settling it. A model that only quotes should fail where no quotation fits, and success on unfamiliar structures makes pure quotation harder to sustain. It cannot distinguish a learned causal model from interpolation fine enough to pass the tests at hand. That takes mechanistic evidence: examining representations and intervening on pathways. A network that reasoned causally without any dedicated symbolic component would be a more interesting result than the one GPT-5 invented.
Learning causation from human text does not make the result quotation, any more than a student who learned it from Pearl is quoting him when she solves a new problem. The open question is whether the model built anything from what it read. The architectural explanation was invented. The behavioral result remains.
Postscript
More capable systems will explain themselves with better vocabulary and more convincing mechanisms, and none of that ties the account to the implementation. A more intelligent system can be more persuasive about things it does not know.
I wrote an article about the difference between association and causal understanding. I reported a capability, asked the system to explain its cause, and published the explanation unchecked, in defense of an authority I had not reread. I had mistaken a plausible causal story for evidence of the mechanism.


