Recursive self-improvement is usually discussed as a future capability of artificial intelligence. At some threshold, an AI system becomes able to improve its own architecture, its more capable successor repeats the process, and the loop accelerates into superintelligence. Whether this can happen, or whether diminishing returns will prevent it, is one of the central disputes in AI forecasting.
But recursive self-improvement has been happening for thousands of years. It produced the civilization that is now building artificial intelligence.
Humanity’s capacity to solve problems has grown over four thousand years far beyond anything biological change could account for. We predict eclipses, manipulate individual atoms, design organisms, and build machines that discover new mathematics. We got here because intelligence has been improving the machinery through which intelligence operates.
Writing gave knowledge an external memory. Mathematics made reasoning cumulative and transferable. Printing accelerated the spread of discoveries. Scientific instruments extended observation beyond biological limits. Computers automated calculation, simulation, and eventually substantial portions of reasoning itself. Each improved our ability to produce further knowledge, on top of whatever it added to the stock.
That is recursive self-improvement.
The Intelligence That Built Intelligence
Not every accumulation of knowledge is recursive self-improvement. Adding another book of findings to a library increases the stock of knowledge; a better method of proof increases the capacity to produce more of it. The loop closes when discoveries improve the methods, tools, or institutions that make further discoveries.
Nothing about this requires a single agent inspecting and rewriting its source code. It requires a causal feedback loop in which a system’s outputs increase its capacity to produce further improvements, whether the improvement occurs in neurons, software, external memory, or the institutions coordinating them.
One might object that humanity has improved only its tools, leaving its cognitive architecture unchanged. The objection treats the skull as a privileged boundary. A mathematician with pencil and paper solves problems the same mathematician cannot solve mentally, and the capability belongs to the combined system.
Human civilization is such a system, distributed across people, accumulated knowledge, institutions, and technology. Discoveries are preserved, transmitted, and applied to improve the methods and instruments behind later discoveries. Its continuity lies in the ongoing reproduction of collective research capacity rather than in any particular individual or institution.
Recursive improvement need not accelerate. A system can keep improving its capacity for improvement while meeting harder problems and rising costs. An intelligence explosion requires feedback strong enough to sustain acceleration over a consequential interval. Human civilization establishes the existence of recursive improvement. It does not establish that an explosion is inevitable.
Diminishing Returns to What?
In Where’s the “intelligence explosion”?, Ramez Naam argues that the loop is real and too weak to sustain itself. OpenAI’s researchers consumed 124 times as many tokens per person while running 1.6 times as many experiments. Naam converts the experiment figure into a return on capability: roughly 2 to 3 percent more research productivity per point of Epoch’s Capabilities Index, against a self-sustaining threshold of 15 to 19 percent taken from a model by Tom Cunningham and colleagues. Within that model, the loop would need to be five to ten times stronger to cross the threshold.
Naam also draws on Are Ideas Getting Harder to Find?, in which Bloom and colleagues document falling research productivity across several domains; sustaining Moore’s Law took ever more researchers. These findings cannot be dismissed as measurements of obsolete methods. A system that keeps improving its methods still faces diminishing returns if improvements get harder to find faster than its productivity rises.
Diminishing returns to research effort and diminishing returns to research intelligence are different propositions. Adding researchers of fixed ability may yield declining marginal returns. Making researchers more capable changes which problems can be solved and how well effort is directed: a better researcher might skip unproductive experiments, find a new algorithm, or spot a conceptual error that has constrained a whole field. Stronger researchers also meet harder problems, so the relationship has to be measured.
Naam does try to measure it, which is what makes his estimate worth arguing with. The difficulty is the unit. His return on capability is estimated from experiments per researcher, a measure of research throughput rather than research productivity. The route he names as a way to strengthen the loop is better researchers running fewer experiments and learning more from each, and that route registers in his metric as nothing, or as a decline. The error runs the other way too, since automation can multiply experiments without multiplying discoveries. A proxy that can miss in either direction leaves the estimate unsigned on the quantity the case for acceleration rests on: how much additional discovery capability buys.
The history of computing shows how the two come apart. Semiconductor research needed growing effort to sustain transistor improvements while computers transformed simulation, statistics, experimentation, and engineering design across other disciplines, and helped researchers build better computers. The productivity of one research process can decline while its outputs improve the broader research system.
Extrapolation is the second difficulty. Naam’s return is fitted over about sixteen capability points within one period of research practice. Projecting it forward assumes the return per point stays fixed as the researchers improve, when the researchers’ judgment is the thing improving. The technology for producing discoveries is itself subject to discovery. That makes long-term extrapolation unusually difficult.
The Researcher Joins the Loop
On October 6, OpenAI published 722 mathematical manuscripts in 372 families of results, produced by an unreleased internal model as part of its model-development evaluations. Many come with Lean formalizations, and some build on earlier results from OpenAI’s models.
OpenAI posed roughly 4,000 problems and published the results it judged significant, so the manuscript count measures selected output, not productivity or importance. Formal verification establishes that a result follows from its assumptions, not that it is novel or significant, and OpenAI warns that some unformalized results may contain errors. Even so, the work is evidence that frontier AI is moving from exercises into original research.
Naam would not be surprised. He expects narrow superintelligence in formal mathematics because a proof checker supplies unlimited verified training signal, and he does not take superhuman mathematics to imply competence at open-ended research. Verification is common ground between mathematics and chess; what separates them is where the results go. A chess engine’s discoveries stay on the board. Optimization, learning theory, numerical analysis, and complexity theory are part of the machinery AI systems are built with, so an AI that discovers useful mathematics may discover improvements to the methods that produce its successors. That is evidence of a prerequisite, and none of it shows the model has improved a successor.
AI already participates in its own development through more concrete channels. Synthetic data supplies training examples for successor models. Models rank and evaluate outputs used in reinforcement learning. Coding systems contribute to training software, numerical kernels, and experiment automation. Each channel has failure modes: synthetic data can propagate errors, automated evaluators can reward shortcuts, and faster training code can raise efficiency without improving the learning algorithm. What they share is the shape of the loop. A model helps build its successor, and the successor may do that work better.
The loop need not close within a single model. Humans may set goals, artificial researchers develop algorithms, automated experiments evaluate them, and training pipelines fold the results into later models. This is continuous with how technological civilization has always improved itself. What is changing is how much of the research artificial researchers do.
The Speed of the Loop
Civilization’s recursive improvement has run mainly through biological researchers, who take decades to educate, coordinate through slow institutions, and need rest. A trained model can be instantiated repeatedly without repeating the education, can share artifacts and capabilities at machine speed, and can run continuously within computational and economic limits. Additional research capacity can be created through computation rather than birth and schooling.
More capable artificial researchers can raise the productivity of each unit of research effort. Automation shortens the interval between discovering an improvement and building it into a successor. Parallelism lets many researchers pursue competing hypotheses at once, even when each cycle takes as long as before.
The loop is still physical. Computation needs energy, fabrication capacity, and capital; training runs take time; physical experiments do not always speed up when more researchers are assigned to them. But some advances need more hardware and others raise what existing hardware can do. A more efficient algorithm or training procedure yields a more capable system without a proportional increase in infrastructure. Physical limits constrain the available trajectories of improvement. They do not tie capability growth to the construction rate of data centers.
Nor does more research capacity guarantee proportionally more discovery. Researchers duplicate one another’s work, exhaust promising approaches, and meet problems whose difficulty grows faster than their capability.
Postscript
Civilization’s recursive improvement has been irregular. Disciplines have hit diminishing returns, institutions have failed, and technologies have stagnated. The record shows that the improvement persists; it says nothing about the mathematical form of its growth.
What changes now is that the machinery produced by recursive improvement can take part in further improvement. The line between human-assisted AI research and AI-assisted human research is becoming hard to draw; both describe one system in which biological and artificial researchers improve the methods, tools, and models used by their successors.
The strength of that feedback is measurable, and Naam asks for the right quantity: how much useful research each new model adds with resources held constant, and how that research translates into better models. His evidence supports skepticism toward claims that an explosion is already underway. It is a weaker basis for forecasting research systems in which artificial intelligence does a growing share of the work.
For four thousand years, humanity has improved its capacity to improve. Through most of that history the intelligence doing the research was biological, expensive to reproduce, slow to educate, and hard to coordinate. Computation is relaxing those constraints. Whether that produces an intelligence explosion is an empirical question, and measurements of yesterday’s research system may tell us little about tomorrow’s.
Recursive self-improvement did not begin with artificial intelligence. It built artificial intelligence. We are now finding out what happens when its products become researchers in their own right.


