The Provenance Fallacy
Why an AI score cannot tell you who thought, judged, or stood behind an essay
Every Axio essay is written with AI assistance.
This is neither a confession nor a revelation. Axio has already described its writing process: under the Gemini Protocol, the human directs the inquiry and adjudicates the result, GPT develops and integrates the argument, and Gemini attacks it from outside the system’s existing assumptions. Drafting, criticism, revision, and synthesis may cycle several times before publication.
This week Substack launched an AI detector, built with Pangram, that lets readers scan eligible posts, notes, comments, and replies. The tool assigns a percentage described as “human-written” or “AI-assisted.” Axio essays will presumably score near the top: 100 percent, or close to it.
We will save readers the scan: every Axio essay is AI-assisted.
That fact is relevant, but it is not dispositive.
A percentage without a quantity
“Percentage AI” sounds like an objective measurement. What exactly is being measured?
Consider an essay in which a human chooses the subject, develops the thesis, rejects weak arguments, checks the evidence, resolves objections, and approves every final claim, while an AI generates most of the sentences. Is the essay 90 percent AI because the model supplied most of the tokens, or 10 percent AI because the human supplied the governing judgment? Now reverse the process: a model originates the argument, and a human rewrites every sentence by hand. The prose may read as fully human even though the machine supplied the central idea.
The percentage could in principle measure the proportion of tokens a model generated, the proportion of surviving language a model generated, the proportion of ideas a model contributed, the proportion of editorial decisions a model made, or the proportion of human effort a model displaced. These are different quantities. A detector examining the finished text measures none of them directly. It detects statistical resemblance and presents the result as provenance.
Grant the detector perfect accuracy at what it measures. It measures resemblance to known classes of human and machine prose. It does not measure authorship.
Writing is part of thinking
Writing is not the transcription of a finished idea. The act of composing a sentence often changes the thought being expressed: a phrase exposes an ambiguity, a paragraph reveals a missing premise, an attempted explanation forces a distinction that was not previously visible.
The Gemini Protocol does not assume that the human begins with a complete argument and delegates the wording. The argument emerges through repeated interaction among human prompting, model generation, criticism, revision, and final judgment. Thought occurs throughout the process. The human role is therefore not merely to supply the initial idea. It is to govern a recursive process through which the idea becomes more precise, more defensible, and sometimes substantially different from where it began.
Authorship is governance
The Gemini Protocol does not treat intellectual production as a contest between one human and one machine. It assigns different functions to different intelligences. The human defines the objective, chooses the questions, sets the constraints, resolves disagreement, and authorizes publication. GPT preserves coherence and develops the argument. Gemini supplies adversarial criticism, rhetorical calibration, and inductive contrast. The final work is produced through structured collaboration under human control.
The models make substantive contributions. They sometimes produce distinctions, objections, formulations, and structural insights that would not otherwise have appeared, and pretending they are merely sophisticated spellcheckers would be false. But contribution is not control.
For collaborative work, authorship attaches less to who produced each sentence than to who governed the inquiry and accepts responsibility for the result. Who chose the objective? Who determined what counted as evidence? Who decided which objections succeeded? Who selected the final claims? Who authorized publication? Who accepts responsibility for errors? At Axio, those governing functions remain human.
As intellectual production becomes more collaborative, authorship moves away from the narrow question of who typed each sentence and toward direction, adjudication, and responsibility.
The finished-text fallacy
Pangram sees only the final sequence of words. It does not see the process that produced them.
The same passage could be an unreviewed model output pasted directly into a post, a human argument polished by a model, a model draft repeatedly reconstructed by a human, an extended human–AI dialogue condensed into prose, or a multi-model adversarial process such as the Gemini Protocol. These workflows have radically different implications for reliability, originality, and responsibility. Their final texts may be statistically indistinguishable.
A detector can infer that a machine was probably present. It cannot reconstruct the distribution of thought, judgment, and control behind the text.
Human voice still matters
Some readers are not seeking argument alone. They are seeking a particular human presence: temperament, humor, vulnerability, style, and the sense of contact with another mind. That preference is legitimate.
But unaided composition is neither necessary nor sufficient for human voice. Editors have always reshaped prose. Speechwriters have long expressed the judgments and persona of public figures, sometimes faithfully and sometimes deceptively. Translators can preserve or distort an author’s sensibility. A human can write generic, lifeless prose without assistance, while an AI-assisted essay can retain a recognizably individual point of view.
The question is whether the published voice faithfully represents the person who stands behind it. AI assistance may weaken that connection. It may also sharpen it. A detector cannot tell which occurred.
The default pressure of an LLM is often homogenizing. It smooths irregularity, removes productive friction, and pulls prose toward a generic median. Preserving a distinctive voice therefore requires deliberate resistance: strong constraints, aggressive revision, and a willingness to reject fluent but characterless language.
Why provenance becomes a substitute for criticism
Readers have legitimate reasons to care about production methods. They may object to undisclosed delegation, distrust unverified model output, or want direct access to a particular person’s unaided expression.
But “AI-generated” is increasingly used as a substitute for evaluating the work itself. Instead of identifying a false premise, invalid inference, fabricated citation, empty abstraction, or derivative claim, the critic produces an AI score. The provenance judgment supplies the appearance of criticism without the burden of argument. That is attractive because determining whether prose resembles model output is cheap, and determining whether an argument is sound is expensive.
Substack’s feature converts the shortcut into a platform affordance. A percentage can be screenshotted, circulated, and treated as forensic proof even though it says nothing directly about truth, originality, verification, voice, or editorial control.
Slop is real
AI has made grammatical prose nearly costless, and the predictable result is an enormous increase in low-effort, repetitive, unverified content. Substack has a legitimate problem to solve.
To its credit, the company’s own launch post acknowledges the central point: machine involvement does not define slop, and slop does not require a machine. It locates the real problem in the mismatch between a reader’s expectations and what actually produced the text. That framing is correct, and it undermines the percentage more than the company acknowledges, because a percentage cannot report expectations, care, verification, or control. A classifier may be useful for spam control, ranking, or aggregate platform analysis. A tool can be operationally useful while remaining inadequate as a public verdict on an individual essay.
The defining vice of slop is that it transfers the cost of thought from the producer to the reader: no serious selection, no verification, no compression, no accountable judgment, and no reason to exist beyond occupying attention. Humans can produce slop unaided. AI-assisted systems can produce rigorous work.
The opposite of slop is not human authorship. It is editorial responsibility.
Postscript
Alongside the detector, Substack shipped a better instrument: a “How I make this” statement, a space where a creator describes the production process directly. That field asks the questions worth asking. Was AI used for research, drafting, editing, criticism, or substantive argument development? Were factual claims checked? Were citations verified? Does a named person accept responsibility for the final publication? Those questions track reliability and authorship far better than a synthetic percentage, because they interrogate the process instead of inferring it from textual residue.
Axio’s statement is simple. Axio does not promise unassisted human prose. It promises selected arguments, sustained criticism, explicit reasoning, and accountable publication.
Substack also lets creators disable scanning of their own posts. Axio will not. Every essay is AI-assisted, and readers need no detector to discover it.
The detector’s question is settled. The argument is not. A percentage beside the prose is no substitute for the hard work of reading.


