Where do thoughts and choices come from? My answer is an evolutionary process with randomness at the bottom. Variation that knows nothing about the problem supplies the novelty, and selection, accumulated over many timescales, supplies the structure that makes the result look directed.
The standard objection is that thought shows no sign of randomness. A mathematician sees a promising transformation, a composer hears where the phrase has to go, a driver takes the gap in traffic without deliberating. The better adapted the mind, the less its working resembles random search. If cognition is evolutionary, as I argued in Evolution Is All You Need, where is the randomness?
It looks for randomness at the level of conscious thought, and that is the wrong level. What reaches consciousness is already highly processed. The idea we experience is the surviving output of lower-level processes whose alternatives are mostly invisible to introspection. The directedness of a thought tells us nothing about whether its novelty was directed at its source.
Looking at the Wrong Level
Consider an experienced mechanic who hears a knock on a cold start that fades as the engine warms and immediately suspects piston slap. The hypothesis is new to this engine and this morning, and it is specific, relevant, and informed by years in the shop. Nothing about it looks random.
But the hypothesis that enters consciousness is not where the process began. It is the output of a system already shaped by layers of prior selection. The mechanic’s nervous system is a product of biological evolution. The concepts available for describing the problem come from a technical culture that has kept useful distinctions and discarded useless ones. Years of teardowns have changed which sounds become salient, and recent context activates some associations while suppressing others. Candidate fragments may compete below awareness before one becomes the thought: check the piston-to-wall clearance.
The idea looks directed because we see the survivor and mistake the endpoint of selection for the source of variation.
Blind and Random
Evolutionary Rationality traced the variation-and-selection schema to Donald Campbell’s blind variation and selective retention, and Gerald Edelman’s Neural Darwinism later proposed an analogous process within the nervous system itself, as selection among variable neuronal groups. Campbell chose “blind” with care. A variation is blind if it is generated without foresight of its success, and on his account blind variation need not be random: a systematic search that tries every option in order is blind too. Dean Keith Simonton, who has carried the theory forward, treats blindness as a matter of degree, with sightedness measuring how well a generator’s variants anticipate their own utility.
I want to keep the stronger word. By random I do not mean that conscious candidates are sampled uniformly, that expert thought resembles dice rolling, or that neural events must be indeterministic in some metaphysical sense. I mean variation that is decorrelated from the system’s current model of the problem, and not only from the eventual success of the variant.
The difference matters because Campbell’s systematic search is blind to success but bound to its representation. It enumerates whatever the current state knows how to enumerate. A generator of that kind can find any answer its representation contains and none that it lacks. Random variation introduces perturbations whose occurrence the representation does not specify, and selection can then capture useful consequences of those perturbations and build them into future structure. Selection can only act on what has been generated.
In brains this variation has a plausible physical floor. Synaptic transmission is probabilistic: a spike often fails to release transmitter at a given synapse, and ion channels open and close stochastically. These are molecular events, at a scale where the randomness is quantum or thermal in origin. Whether a particular thought turns on a particular fluctuation is a harder empirical question, but the noise is there to be used. In a branching universe, that means each descendant of a thinker receives a different perturbation.
The argument does not depend on the physics. The criterion is decorrelation, not indeterminism: a pseudorandom number generator is deterministic, but its output carries no information about the problem a model is working on, which is all the argument needs.
Selection supplies direction after the fact.
Random at the Bottom, Directed at the Top
There is no contradiction between basal randomness and highly structured thought. Selection converts random variation into accumulated structure. A successful variant is retained, and future variation then occurs within a system altered by that success.
Biological evolution already demonstrates the principle. Mutations need not anticipate their usefulness, yet selection produces eyes, immune systems, echolocation, and brains adapted to particular environments. The apparent purpose lives in the accumulated structure, not in foresight at the source of each mutation.
Some nervous systems go further and generate variation on purpose. Juvenile zebra finches learn their song by trial and error, matching their own output against a tutor’s, and their early song is highly variable. Bence Ölveczky, Aaron Andalman and Michale Fee traced the variability to a basal-ganglia loop that feeds the motor pathway through a nucleus called LMAN. Silencing LMAN made juvenile song immediately and reversibly stereotyped, and lesions of Area X, another part of the loop, permanently remove the rapid pitch variation that serves as the bird’s exploration. The same circuit also biases output away from vocal errors, so what reaches the motor pathway is noise passed through learned structure, the architecture this essay describes. Birdsong is motor learning rather than thought, though cortico-basal-ganglia loops run through prefrontal cortex too. It is the clearest case of a brain building a dedicated generator of variation for selection to work on.
Cognition inherits the same architecture across nested timescales. Natural selection shapes cognitive machinery. Cultural selection shapes concepts, practices, languages, and institutions. Individual learning shapes associations, representations, and heuristics. Ongoing cognition selects among candidate interpretations, memories, actions, and expressions. Each layer stores information extracted by earlier selection, and that stored information shapes what later variation becomes, so outputs grow more structured the more layers they pass through.
The randomness is still there, buried beneath accumulated selection. Buried means passed through more and more selected structure. It does not mean made less random.
Selection Moves Information Into the Generator
Selection does more than choose among current candidates. It changes the machinery that turns variation into candidates, and this is the step that explains why cognition becomes increasingly directed.
Write the random variation as R and the cognitive machinery as G, so that a candidate is C = G(R). R is independent of what the system knows about the problem. C need not be: it can be highly structured and depend heavily on everything G encodes. Learning changes G. It does not need to change R.
R enters in two places. In input variation it is a draw that G turns into a candidate. In structural variation it perturbs G itself: noise in synaptic change, two concepts activated together by accident, a rule misremembered and applied anyway. Input variation can only yield candidates G already has room for. It escapes G’s preferences but not its reach. Structural variation changes what G can produce at all, and it is the deeper source of novelty. Biological mutation is structural variation, since it acts on the genome, which is the generator, rather than on an input passed through it. Selection then decides which perturbed generators to keep, so learning changes G partly by retaining G’s own accidents.
Suppose a random perturbation yields a useful candidate and selection retains it. The retained structure changes G, and with it the distribution over candidates: the same kind of perturbation now becomes some candidates more readily and others less. In Simonton’s vocabulary, this is how a generator acquires sightedness.
The cycle runs:
random variation → transformation by G → candidate → selection → altered G
and then a fresh, independent perturbation enters the altered machinery. The higher we move through the stack, the more structure accumulated selection places between random variation and observable cognition. The randomness at the base remains uninformed while its consequences become increasingly informed. An expert and a novice could draw on the same stochastic source and produce very different candidates, because their G differs. That is the precise sense of Evolutionary Rationality’s claim that an expert does not search the same distribution as a novice: the distribution that differs is over candidates, the output of G, while the perturbations feeding it are drawn alike.
The same architecture covers choice. A decision is a candidate like any other. Randomness can affect which alternatives become available for evaluation without making the resulting choice arbitrary, because the evaluation is performed by structure accumulated over the chooser’s history, and that history is the chooser’s own.
Compressed Selection History
Compressed selection history lives in G. Evolutionary Rationality introduced the term: what presents as intuition is the residue of earlier error correction.
An expert often cannot explain where a useful thought came from. A physician recognizes a pattern; a writer knows how an argument should be framed. That opacity is sometimes taken as evidence that intuition belongs to a different category from rational thought. It is what we should expect when previous search has altered the machinery that filters and propagates current variation. The historical search does not need to be replayed, because its consequences are built into the system doing the thinking.
Expertise therefore changes more than what someone knows. It changes what occurs to them. Two people can know the same formal rules and reason very differently, because one notices the relevant symmetry and the other never considers it.
Creativity Is Not a Special Faculty
The same architecture runs through cognition that nobody calls creative, which is why creativity needs no separate faculty. Perception is not a passive reading of uniquely determined sensory data. Sensory input is noisy and ambiguous, and competing interpretations are constrained by context, prior learning, attention, and expectation until one organization stabilizes. Bistable figures expose the competition: the Necker cube’s input never changes, yet the interpretation flips. Seeing an object is already seeing a selected representation.
Memory retrieval has the same structure. A cue activates overlapping traces, associations, and reconstructions; some become accessible and the rest drop out of competition. So does language production. At any moment many continuations are possible, and semantics, syntax, habit, goals, and recent activation constrain them until one is spoken. Speech errors show the losers. A blend or a spoonerism is a rejected candidate that leaked into output.
Planning, reasoning, and decision-making work over candidate trajectories, decompositions, analogies, counterexamples, and options, most of which never reach explicit consideration. Formal reasoning does not escape the architecture. Logic can determine whether a proposed inference is valid, but it does not tell us which theorem to attempt, which representation to choose, which lemma to invent, or which contradiction to notice. Generating those candidates still requires variation. Creativity is the name we give the process when the candidate that survives is new.
Prediction Is Part of the Machinery
Predictive-processing theories might seem to offer a different picture, in which the brain generates structured predictions from an internal model rather than producing arbitrary candidates. They describe part of the same machinery. Evolution, development, culture, and learning determine which models exist, which regularities they encode, which predictions receive high probability, and how prediction errors modify them. Prediction is one of the mechanisms through which accumulated selection shapes what later variation becomes, and a well-trained predictive system should produce strongly directed candidates, as this account predicts.
One strand of the Bayesian literature supports the account directly. Adam Sanborn and Nick Chater argue that a Bayesian brain need not represent or compute probabilities at all; it approximates inference by drawing samples from a learned model of the world. On that view the learned model is the compressed selection history. The randomness is the draw that performs each sample, not the sampled result: the draw is independent of the model, and the model turns it into a candidate that depends heavily on what the model has learned.
Pointing to the sophistication of the current generator leaves unexplained how the sophistication arose. The generator contains information because selection put it there.
Why Thought Looks Non-Random
Introspection exposes little of this. We experience the recognized object, the recovered memory, the chosen word, the decision. We do not experience the lower-level variation, the rejected candidates, or the history that shaped which candidates became likely. A thought arrives coherent and relevant, so it is natural to imagine that cognition produced it in something like its final form.
A deep selection process should produce exactly that phenomenology. The more successful the accumulated selection, the less random the surviving output will appear. Expertise intensifies the effect: bad possibilities become less salient, useful representations appear sooner, and search becomes less visible as more of it is compressed into the machinery conducting it. The appearance that motivated the objection is a prediction of the account.
Language models offer a clean instance. Suppose a model assigns the next token cat 0.6, dog 0.3 and aardvark 0.1. The sampler draws a uniform number between 0 and 1, and that number knows nothing about cats, dogs, or the sentence. The trained distribution maps it onto a token. The draw decides which point is sampled and training decides what that point means. Turn the sampling temperature toward zero and the model returns its single most probable continuation every time; raise it and the output grows more varied and, past a point, incoherent. The comparison does not require brains and transformers to implement the same algorithms. It shows random sampling and directed-looking output coexisting in a system whose mechanics we can inspect.
What the Compression Costs
The compression that makes expertise fast has a price. A generator shaped by past success makes some possibilities easy to produce by making others hard, and the hard ones include possibilities nobody has tested and found wanting.
Merim Bilalić, Peter McLeod, and Fernand Gobet tracked the eye movements of expert chess players facing a position with a familiar mate and a shorter, less familiar one. Players who had found the familiar solution said they were looking for a better one, while their gaze kept returning to the squares that belonged to the solution they already had. The search they reported was not the search their attention carried out. This is the Einstellung effect, and on this account it is compressed selection history working against its owner: the first idea comes from the generator’s history, and so does the attention that is supposed to check it.
The Bias Before Bayes described model-generator bias, where one preference shapes the hypotheses considered and is then multiplied by correct inference. Here we can see where such a generator comes from. Its biases are its selection history, the same information that makes it competent.
This is why randomness at the base is essential rather than vestigial. Variation the generator does not dictate is the only route back into regions its history has made improbable. A deterministic mechanism can be told to search unfamiliar territory, but the rule that tells it to do so belongs to the current generator, and its exploration is generated by existing structure. A random perturbation is an event whose particular occurrence that structure does not dictate. A system that annealed model-independent variation to zero would be confined to trajectories its existing structure generates.
The route has a condition. Input variation still passes through G, and it reaches a region only if G gives that region some weight. Selection can drive that weight low enough that, for practical purposes, no amount of input variation gets through, which is what the chess players’ gaze looks like. Structural variation could reopen the region, but a generator shaped by long success is also a generator stabilized against its own perturbations.
So when expertise stops improving on its own first ideas, more noise through the same machinery is not enough. The remedy has to come from a different mapping: a change of representation, or the problem as posed by someone trained elsewhere. The perturbations that reach the new mapping are as ignorant as the old ones. What has changed is the structure they pass through. Thought draws its novelty from randomness that knows nothing about the problem, and its direction from that structure, which accumulates above the randomness and never inside it.


