Brett Hall argues that large language models are not genuinely creative. In his conversation on AGI and personhood he grants that they produce fluent, useful and surprising output, but holds that they manipulate learned statistical structure rather than generate the conjectures that make human beings universal explainers. He does not claim artificial persons are impossible. A genuine AGI would be a universal explainer and therefore a person; current LLMs lack the mechanism. David Deutsch supplies the background. A universal explainer is capable of open-ended explanatory growth, and in Creative Blocks he holds that “AGIs are people“ has been implicit in the concept from the outset.
The Personhood Trap answered Hall’s identity thesis, that the general intelligence worth the name is the intelligence of a person, and asked him for a criterion separating creative conjecture from rearrangement. The Randomness Beneath Thought supplies one, and it puts creativity somewhere Hall does not look. If human beings had a faculty that generated conjectures intelligently, LLMs might lack it. But the source of novelty in human thought is not intelligent.
It is random.
The Knowledge Is Not in the Variation
The Randomness Beneath Thought argued that thought has an evolutionary architecture. A random variation R, decorrelated from the problem, passes through a generator G built by selection, and the candidate C = G(R) can be highly structured even though R knows nothing. Biological evolution shaped the nervous system, cultural evolution supplied concepts and practices, individual learning changed associations and heuristics, and the current problem shapes what is salient now. The randomness beneath all this remains ignorant. The knowledge accumulates above it, in G.
R reaches G in two ways. In input variation it is a draw that G turns into a candidate, which escapes G’s preferences but not its reach. In structural variation it perturbs G itself, and that changes what G can produce at all. Both matter below.
This is why expert thought does not look random. A good conjecture can arrive at once and look closely adapted to the problem, but the candidate we experience is not basal variation. It is variation that has passed through a long selection history. What looks like intelligent generation can be selection history viewed from the top.
Hall places the intelligence too early. He observes that human conjectures are relevant, sophisticated and problem-directed, and treats this as evidence that their source belongs to a different category from the blind variation of evolution. A random perturbation can know nothing about a problem while the machinery it passes through knows a great deal.
Having a Problem First
Hall also separates human creativity from evolution because people have problems and then conjecture solutions. Mutation does not know what problem an organism faces; a scientist can ask why something happens and set out to explain it.
Having a problem first does not mean the mechanism supplying novelty knows how to solve it. The problem conditions the generator and the selection environment. A physicist puzzling over an observation that conflicts with a theory attends to different information, activates different concepts and counts different candidates as promising. The system is strongly conditioned by the problem while the basal variation feeding it stays decorrelated from it.
Human cognition has also moved much of selection inside the organism. We imagine consequences, run simulations and reject bad conjectures before acting on them. Imagination is selection moved inside the organism. That makes human creativity far faster than evolution without giving it a different source of novelty.
Whether the problem was self-generated is a separate matter. The Personhood Trap argued that goals supplied from outside leave the capacity to pursue them untouched, and human creativity routinely runs on commission: a theorem someone asked for, a design constraint, a diagnosis to make. Self-generated problems may matter for full explanatory agency. They are not part of what creativity is.
What an LLM Has
Training builds a highly structured generator through repeated error correction over a large corpus. At inference the model combines that structure with the prompt and context to produce a distribution over continuations, and a random draw picks one. The mapping onto the architecture is direct:
R = the sampler’s random draw
G = the trained model plus prompt and context
C = the generated continuation
The draw knows nothing about the problem. The generator knows a great deal, and the prompt conditions it the way a problem conditions a human thinker. For R → G → C this is a direct implementation, not an analogy. Whether the rest of the loop is present, selection and lasting change to G, is the question the rest of this essay takes up.
There is also a loop inside a single generation. Each sampled token becomes part of the context, and the context is part of G, so the randomness of one draw changes the mapping for every draw after it. Call it state variation: it reshapes the effective generator for the rest of the response and lasts only as long as the context does. Structural variation is reserved for changes that outlast it.
Noting that LLM generation involves randomness, statistics and learned correlations settles nothing, since human cognition involves all three. Calling an LLM a next-token predictor settles nothing either. As The Personhood Trap put it, a training objective names the pressure applied, not what grew under it, just as reproductive success does not describe an eye. LLMs routinely produce programs, arguments and designs that satisfy new combinations of constraints and appear nowhere in their training data. That does not make each one creative, but it blocks the inference from “generated by next-token prediction” to “not creative”. The open question is whether humans have some further mechanism, constitutive of creativity, that LLMs necessarily lack.
Where the Generator Changes
Hall’s Popperian framework sharpens the comparison, because explanatory knowledge grows through conjecture and criticism and not through candidate generation alone. A closed self-critique loop is weak. If a model generates, criticizes and revises with every evaluation coming from the same generator, its failure modes reinforce themselves.
LLM systems need not be closed. A model can run the code it wrote, submit a proof step to a checker, compare a prediction with a measurement, or take criticism from a person. The candidate is then tested against constraints outside the generator, which is Popperian in structure. That the feedback arrives as tokens does not disqualify it; no criticism reaches a human unmediated either, as The Personhood Trap argued against Hall’s appeal to contact with reality. The distinction that matters is between closed evaluation and contact with an independent selection environment, and it cuts across the line between humans and machines.
Criticism produces growth only if it changes what the system generates later, and here the distinction between input and structural variation does real work. A frozen model at inference has input variation and, within one context, the state variation of its own sampled tokens. The structural variation that persists happened in training, where noise in gradient descent perturbed the generator and the loss decided what to keep. Humans have persistent structural variation all the time. Any failure can change the generator that produces the next conjecture.
Persistence does not require rewriting weights. External memory, retrieved state, code, databases and auxiliary models can carry what criticism taught into later search, as books, notebooks, instruments and institutions do for us. The test is whether criticism at one time changes what the system can generate later. That is an engineering question, and current systems answer it partially.
Recombination
The usual retreat grants LLMs a weak creativity and denies them explanatory novelty: they recombine existing human concepts in a latent space fixed by training. The Personhood Trap answered the provenance version of this. Einstein and Darwin also rebuilt inherited material, and “rearrangement” needs a criterion independent of who did the rearranging.
The distinction between input and structural variation supplies one, though not as a definition of explanatory novelty. A fixed G can produce a genuinely new explanation, one nobody has stated, if it lies within G’s reach, and a change to G guarantees nothing about the quality of what follows. What the distinction marks is reach. Input variation through a fixed G explores what that G can reliably produce, however surprising the result, and that is what a fixed latent space means. Structural variation is what lets the system’s future reach expand beyond it. On that reading the recombination charge describes a real bound on a frozen model answering one prompt, the bound of its reach and not an absence of novelty, and it describes no bound at all on a system whose generator criticism keeps changing. It applies with equal force to a human who never revises their own.
Universality and Understanding
Deutsch’s stronger concept, explanatory universality, concerns the scope of possible growth. Even granting that humans are universal explainers and current LLMs are not, it does not follow that LLMs cannot create. That would need an argument that every creative act requires universality. A system can fall short of open-ended explanatory personhood and still do creative work.
Hall can press a stronger objection: explanatory creativity requires understanding, and LLMs do not understand. That is relevant, since a good explanation represents causal or structural relations in a way that survives criticism. But “understanding” has to name a capacity before it can rule anything out. If it means building causal models, the question is whether LLMs build them. If it means testing models against reality, it is whether the system has an open loop. If it means revising models when predictions fail, it is the persistence question above, and there current fixed-weight models are weaker. Each reading can be investigated. “LLMs do not understand” leaves the mechanism unspecified.
Where Current LLMs Fall Short
There are good reasons not to equate current LLM creativity with human explanatory creativity, and the architecture says where they lie. Current systems can run variation through a structured generator, expose candidates to external criticism and revise them. They are weaker at persistent structural variation: at letting criticism permanently change the generator, at discovering their own problems, and at sustaining a cumulative epistemic trajectory without external scaffolding.
Those are substantive limitations. Hall may be right that current LLMs are not yet autonomous, continuously learning universal explainers. That is a different claim from saying they cannot be creative.
Building Before Understanding
Deutsch argues in Creative Blocks that “what is needed is nothing less than a breakthrough in philosophy, a theory that explains how brains create explanations”, and Hall follows him: until we understand how humans create explanations, we cannot program a machine to do it. That imposes a theory-first standard no other capability has met. Evolution produced cognition with no theory of cognition. Breeders produced traits long before genetics explained them. Machine learning produces representations its designers neither specify nor fully understand afterwards.
Construction can precede understanding, and a Popperian should expect it to, since knowledge grows through conjecture and criticism precisely because understanding does not come first. If creativity runs on variation and selection, there is no reason its artificial version should wait for a complete theory of the human one.
Stepping Outside the Library
On his decision-making page, Hall grants that a language model produces novelty but says that recombination is not creativity, “because you are drawing from inside a finite library even if it is large.” People, he says, “can step outside of the library.” In this essay’s terms the library is G’s reach, and stepping outside it is structural variation. The phrase names the phenomenon without explaining it. What lets a human step outside? If the answer is “creativity”, the explanation is circular. If the answer is that random variation perturbs a generator, accumulated knowledge turns the perturbations into candidates, criticism tests them against reality, and what survives changes the generator, then there is a mechanism, and none of its parts is unavailable to machines.
Humans may implement it more completely, with richer representations, better causal models and a tighter link between criticism and structural change. Those are differences of capability. They are not evidence for a separate creative substance, and the burden on Hall is to name a step in that loop that human cognition has and an LLM-based system cannot.
Every generator draws from inside a finite library, ours included. What Hall calls stepping outside is the library being rewritten by whatever criticism keeps, and nothing in that process needs a person to carry it out.


