Max Weinreich’s The Crisis of AI-Generated Mathematics identifies a real problem. Research mathematics has used the production of mathematical work as evidence that the producer possesses mathematical competence and understanding. AI breaks that connection.
A paper has traditionally done several jobs at once. It contributes a result, communicates an argument, establishes priority, and certifies something about its author. On Weinreich’s account, publishing shows that a mathematician practiced the deepest form of the art and arrived at “a complete understanding of a piece of mathematics.” AI-assisted production weakens the inference, because a mathematician can now produce sophisticated work without having discovered it and possibly without understanding it.
Weinreich is right about the decoupling. His mistake is treating the collapse of the proxy as a crisis for mathematics itself. It is a crisis for mathematical credentialing.
The bundled proof
Before powerful automation, producing a difficult artifact usually required possessing the skill associated with it. A sophisticated program demonstrated programming competence. A proof demonstrated mathematical ability. Useful output and difficult human achievement shared a production bottleneck, so institutions could treat one as evidence of the other.
As I argued in The Human Premium, these are different forms of value. A proof has artifact value because it establishes a result. Producing it may have achievement value because doing so displays skill or creativity. It may also have credential value because its production provides evidence about the mathematician. That essay separates two further kinds, relational and status value, which carry most of the argument in art and little of it here. Human scarcity kept the three correlated. AI makes them independently variable.
Weinreich responds by treating human production as partly constitutive of the value of mathematics. Solving more important problems, he argues, is not really the goal of the field; the practice of doing mathematics is an end in itself. That identifies a legitimate good, though not the only one.
Doing mathematics can be intellectually rewarding, and discovering a proof can be an extraordinary achievement. Struggling with a problem produces forms of understanding that reading the finished result does not. None of this fixes the epistemic value of the result.
A difficult human proof can be an impressive achievement and establish something unimportant. A machine-generated counterexample can be mathematically important and represent no human achievement at all. The two dimensions were never identical. A shared bottleneck kept them in step.
Knowledge and understanding
Suppose an AI produces a valid counterexample to a longstanding conjecture. Mathematical knowledge has grown, since a question that was open is now closed. Weinreich can reply that this does not by itself increase human mathematical understanding. He would be right, and he would be making a different claim. Mathematical knowledge and human mathematical understanding are both valuable, and neither reduces to the other.
His Jacobian Conjecture example makes the distinction concrete. An AI generates a counterexample, Terence Tao then uses AI to work out what it says, and the future on offer has machines producing the mathematics, machines explaining it, and humans surviving as intermediaries.
The danger is real. Provenance does not establish it.
A mathematician might ask an AI for a proof, paste it into a paper, and understand almost none of it. Another might use AI to explore the proof, test alternative explanations, reconstruct crucial steps, identify the mechanism behind the result, and then extend it. Both are “AI-assisted.” Their epistemic states have almost nothing in common.
AI does differ from earlier mathematical tools in that it can substitute for cognitive work rather than facilitate it. That strengthens the case for assessing understanding directly. The same system supports understanding in one use and replaces it in another, and the label records neither.
A theorem does not inherit its truth from the biography of whoever found it. Human understanding remains valuable and has to be evaluated on its own.
The certification problem
Weinreich’s institutional diagnosis is the durable part of the paper, and the collapse of proof-of-work credentials it describes reaches well beyond mathematics. If a paper no longer demonstrates that its nominal author can independently produce or understand its mathematics, publication cannot serve as the credential it once did. Hiring committees, grant bodies, and graduate admissions have all inferred qualities of mathematicians from finished work, and AI weakens every one of those inferences.
His remedy is to protect the old correlation. Departments might reserve positions for mathematicians who avoid AI, or weight AI-free papers more heavily in promotion. That preserves the proxy and leaves the property the proxy was meant to track unmeasured.
If a department cares whether someone understands a proof, test understanding. If it cares about independent reasoning, test independent reasoning. Judgment, originality, explanatory ability, problem selection, and the capacity to recognize a plausible failure can each be assessed more directly, through oral defense, adversarial questioning, reconstruction of critical steps, extension of a result under altered assumptions, or disclosure of what the machine contributed.
The cost is real. The bundled paper was an imperfect proxy, but a cheap and scalable one, and direct assessment consumes expert attention. Academia now faces a tradeoff between low-cost noisy signals and expensive high-fidelity evaluation, and the likely response is layered: routine screening for low-stakes decisions, spot checks, intensive assessment where the stakes justify it. High-fidelity assessment will also be distributed unevenly, since institutions with more expert time and money will do it better. That is a cost of unbundling, not a reason to keep a proxy whose reliability has collapsed.
Measure the property you care about, not the historical bottleneck that used to correlate with it.
The competence problem
A stronger argument for preserving unaided mathematics starts from how judgment is acquired. Mathematicians learn to recognize a fragile proof partly by building and breaking proofs themselves. The expertise required to supervise a process is often created by performing it.
Consider an ecosystem in which machines grow steadily more capable while successive generations of humans do less of the formative work through which mathematical judgment develops. Automated production improves. The surrounding human community loses the independent competence needed to evaluate it. That failure mode deserves to be taken seriously.
What it supports is protected formative practice rather than a prohibition on machine-generated mathematics. Students may need to prove theorems without AI even when AI could solve them instantly. Researchers may need occasional controlled unaided exercises. Institutions may need to maintain redundant human expertise where automation performs the immediate task more cheaply.
The incentive problem is harder than the epistemic one. Redundant competence looks most wasteful exactly when machines appear most reliable, so institutions will underinvest in it until rebuilding the capability is difficult. Preserving human expertise has the logic of resilience: its value is easiest to ignore while nothing has gone wrong. Recognizing the need does not summon the incentives, and the cost-minimizing equilibrium under existing incentives may contain less human competence than the epistemically resilient one.
Human achievement value also has a recruitment function. People give years to mathematics partly because discovery is difficult, uncertain, and personally meaningful, and if machines routinely reach the frontier first, some of that motivation goes with it. Withholding tools does not recreate the old frontier, since a protected sandbox is not the edge of human knowledge. Mathematical culture may still need to leave people room to formulate problems, pursue conjectures, and find things out for themselves if it is to keep producing mathematicians.
The criterion is functional. Preserve unaided practice where the practice is causally necessary to produce capabilities and incentives we still need. Do not preserve scarcity because scarcity once generated prestige.
Unbundling does not mean an assembly line
Separating roles does not require assigning them permanently to different agents. A human mathematician may still discover, prove, explain, verify, and extend a result, with AI contributing at any of those stages or at none of them. What has disappeared is the assumption that one finished paper reveals who performed which intellectual work.
Weinreich nearly reaches the same conclusion
His most interesting proposal points toward the alternative. Traditional authorship, he suggests, might give way to something like mathematical co-ownership: a mathematician who demonstrates authoritative understanding of an existing work acquires standing in relation to it without having originally produced it, and journals validate that understanding.
The proposal concedes the central institutional point. Authorship and understanding have come apart, and once understanding is certified separately from production, prohibition loses its credentialing rationale. His motivational case survives the concession, that AI cuts mathematicians out of their deepest experiences, and it has to be argued on its own terms. Machines can generate results. Humans can understand them, explain them, criticize them, recognize their significance, and extend them.
Who found the proof? Who understands it? Who verified it? Who can generalize it? Who is answerable for the claim that it is correct? Under overwhelmingly human production, one author list answered all of those questions at once. It no longer does.
A separate argument
Weinreich extends his case to AI companies, autonomous agents, infrastructure vulnerability, and catastrophic risk. Those claims need their own support.
One connection holds. If humans lose independent mathematical competence while advanced systems require human mathematical oversight, deskilling becomes an AI-risk factor in its own right. That reinforces the case for preserving competence, and it carries none of the broader conclusions. If advanced AI poses an unacceptable systemic danger, that is a reason to constrain advanced AI whether or not it proves theorems. Problems in mathematical credentialing cannot supply the argument.
Postscript
A mathematical paper once bundled truth, discovery, understanding, achievement, authorship, and prestige, because producing advanced mathematics almost always required advanced mathematical ability. AI makes those variables independently adjustable.
Restoring the correlation by making mathematical production artificially difficult is the wrong repair. Judgment, understanding, explanation, criticism, problem selection, and responsibility for what gets asserted still have to come from somewhere, and each of them can be measured directly. Where those capacities require unaided practice to develop, protect the practice deliberately. Where machines produce genuine mathematical knowledge more efficiently, the knowledge is not diminished by the ease with which humans came to hold it.
Weinreich is right that something is ending. A paper can no longer be assumed to certify the intellectual history that produced it. Mathematics does not depend on that assumption.
AI has not made mathematics artificial. It has made our proxies visible.


