Last year, in The Boghossian Principle, I used the phrase “academic cargo cults” for disciplines that reproduce the visible machinery of scholarship while weakening the norms of criticism and error correction that make that machinery epistemically useful. The argument there was procedural: a system that protects its foundational assumptions from serious challenge cannot reliably correct itself.
There is another way to test the same diagnosis from the output side. Instead of inspecting the machinery, inspect what the machinery leaves behind. If the process works, there should be durable additions to knowledge, visible revisions of once-credible claims, and constraints strong enough to prevent preferred conclusions from surviving indefinitely by reinterpretation.
A recent exchange about Gender Studies suggested a simple version of that test. One participant defended the field on the grounds that science has not yet answered every question about sex and gender. The reply was: “Name three verifiable, robust discoveries from Gender Studies in the last 25 years.”
Defenders in the exchange apparently struggled to supply them. That proves almost nothing; a social-media thread is not a literature review. So the task is to answer the question properly.
Once the standard is made fair enough to include the humanities and interdisciplinary scholarship, Gender Studies can produce plausible candidates. Feminist scholarship helped expose male-default assumptions in biomedical research. Intersectionality identified a classificatory failure in antidiscrimination law and social analysis. Historical scholarship organized around gender has recovered evidence and social structures that conventional histories often neglected.
Those examples force a correction to the original challenge. A discipline does not need to discover a new particle or empirical law to contribute knowledge, and it does not lose credit merely because another field supplied the final measurement. A conceptual distinction, a historical reconstruction, or the identification of a systematic blind spot can all be epistemic achievements.
So the interesting question is no longer whether Gender Studies has ever contributed knowledge. It has. The harder question is whether its scholarly machinery reliably produces cumulative knowledge, revises its own mistakes, and contains constraints strong enough to force change when favored frameworks stop fitting the world.
Academic activity is not knowledge production
Universities have accumulated elaborate machinery around scholarship. There are departments, journals, peer review, conferences, graduate programs, professional associations, and increasingly specialized vocabularies. None of these constitutes knowledge by itself; they are institutional technologies intended to help produce and preserve knowledge, and they can become detached from that function.
Peer review does not make a claim true, nor do a thousand citations, a PhD, a prestigious journal, or specialized terminology. These things are proxies for processes that, under favorable conditions, expose false claims and preserve better ones. When the proxies persist after the error-correcting function weakens, the appearance of scholarship can remain even as epistemic productivity declines.
Software engineering has the same distinction. Code review, issue trackers, design documents, release procedures, and architecture committees can all contribute to reliable software, but if the program never has to satisfy a specification or pass a test, the development process can become exquisitely sophisticated while the software remains wrong.
An academic field can possess every visible feature associated with serious inquiry while lacking an effective process for discriminating between successful and unsuccessful explanations. Papers cite papers, theories generate interpretations, conferences generate proceedings, and credentials generate more credentialed practitioners.
That is cargo cult academia. The term names a degree, not a verdict on a whole discipline: the pathology begins wherever the visible machinery is treated as evidence of reliability independently of whether it forces errors to be corrected.
The output test
Different disciplines produce different kinds of knowledge. Mathematics proves theorems, history reconstructs events and causal sequences from incomplete evidence, archaeology infers vanished societies from physical remains, linguistics discovers structural regularities, and philosophy can expose contradictions, establish the consequences of premises, and construct distinctions that survive criticism.
Not every legitimate academic field must run experiments or discover universal laws. But a field claiming epistemic authority owes something: sustained inquiry should leave constraints on what can reasonably be believed.
A fair output test is therefore broad. Ask what important things the field established that were not previously known or adequately understood, are publicly assessable, have survived serious adversarial scrutiny, and remain part of our best account of the subject. A concept can count if it enables distinctions that could not previously be made, resolves genuine explanatory confusion, unifies otherwise disconnected observations, or constrains subsequent reasoning in a durable way.
Under that standard, the three contributions named above plausibly qualify. The diagnosis that biomedical research treated male subjects as a default case exposed a real epistemic defect. Intersectionality identified a failure in models that treated race and sex as independent axes and therefore could not represent discrimination directed specifically at Black women. Historical work focused on gender has uncovered facts and institutional structures that older histories neglected.
These are real contributions. The output test cannot bear the whole weight of the critique.
Compare another young field
Consider complexity science, which is also young, interdisciplinary, and difficult to demarcate cleanly. It grew out of interactions among physics, mathematics, biology, and the social sciences, and its researchers frequently publish their strongest results in journals belonging to other disciplines.
Yet ask what complexity science contributed and plausible answers appear immediately. Network researchers identified small-world organization, showing that systems can combine dense local clustering with surprisingly short paths between distant nodes; subsequent work found such structures across biological, technological, and social networks and developed a substantial theory of their dynamical consequences.
Researchers also identified network motifs, small connectivity patterns occurring with strikingly different frequencies across real networks and random alternatives. Motif analysis became an empirical tool for characterizing the architecture of regulatory, neural, ecological, and engineered systems.
Work on human dynamics established that many forms of activity exhibit bursty, heavy-tailed inter-event timing rather than the Poisson timing earlier models assumed, which materially changes how diffusion and contagion have to be modeled.
Complexity science did not invent all of the mathematical and statistical tools involved in these discoveries. That is irrelevant. Its contribution was to assemble questions, representations, models, and methods into a research program that exposed previously unnoticed regularities.
The same standard should apply to Gender Studies. If feminist scholarship caused medicine to notice that male-default study designs were obscuring clinically relevant differences, that can count as an epistemic success even if cardiologists later made the measurements. A field can add knowledge by discovering that another field is asking the wrong question.
Credit should track contribution to knowledge, not departmental ownership.
That cuts both ways. The strongest empirical work associated with Gender Studies often uses methods familiar from sociology, psychology, economics, and epidemiology, which is no defect. Did the research program formulate questions other disciplines ignored, or create concepts that revealed distinctions nobody could previously state? Did it identify background assumptions that were distorting results, or motivate empirical work that would otherwise have been delayed? Where the answer is yes, the field deserves credit. Importing an established result and redescribing it in the vocabulary of oppression is not an additional discovery.
The comparison survives only if the same generosity is extended to both fields.
The revision test
A harder comparison asks not only what a field has added, but what it has abandoned under pressure.
Complexity science gives clear examples. Claims that real-world networks are generically scale-free were substantially weakened by later empirical work. Strong universality claims about urban scaling remain contested. Self-organized criticality turned out to be less ubiquitous than some early enthusiasts suggested.
Those failures count in the field’s favor because a functioning knowledge system produces discoveries and revisions. It makes claims strong enough to fail, discovers where they fail, and changes accordingly.
The corresponding test for any discipline is broader than laboratory falsification: what important claim, framework, or generalization did the field once treat as credible and later substantially abandon because evidence, argument, or counterexamples showed that it was inadequate? A healthy research tradition should have answers, because its intellectual history will contain positions that were narrowed, rejected, or replaced for reasons stronger than fashion.
In the humanities, theoretical death rarely looks like an experiment delivering a clean knockout. A historical framework may collapse when newly recovered evidence exposes its omissions. A conceptual scheme may fail because counterexamples reveal that it cannot represent cases it claimed to explain. A philosophical position may be abandoned because its commitments generate contradictions or because a rival framework explains the same material with fewer ad hoc assumptions.
Critical social-justice fields do contain examples of this kind of correction. Universalized accounts of “woman” were challenged because they often failed to represent the experiences and social positions of Black women, working-class women, and women outside Western societies. Where a framework claimed broad adequacy and counterexamples showed that claim to be false or incomplete, narrowing or abandoning the framework counts as epistemic progress.
Revision itself comes in two kinds. Sometimes a field revises a theory because the previous position has become untenable. At other times one interpretive vocabulary simply succeeds another because scholars find it more compelling, politically attractive, or philosophically fashionable.
Only the first demonstrates error correction. The revision test asks which transitions were forced.
When normative commitments become premises
The distinctive vulnerability of disciplines organized around critical social justice is not merely political homogeneity. Political homogeneity can suppress criticism, but the structural problem runs deeper because normative commitments can migrate upstream into the inferential machinery itself.
Consider claims such as these: societies contain structures of oppression; social categories can be shaped by power; persistent disparities may reflect institutional arrangements; historically marginalized groups may experience institutions differently. Each can support legitimate inquiry because each can, at least in principle, compete with alternative explanations.
Now change their logical status. Suppose oppression, power, structural inequality, and marginalization become constitutive assumptions of the framework rather than hypotheses competing with alternatives. Research no longer asks whether oppression or structural power best explains an observation; it asks how oppression or structural power manifests in the observation.
This weakens discriminatory power. If women are underrepresented in an occupation, exclusion or gendered socialization can explain it; if women are overrepresented, occupational segregation and the devaluation of feminized labor can explain that instead.
The same flexibility appears elsewhere. If members of a marginalized group speak openly, their testimony may be interpreted as standpoint knowledge; if they remain silent, the silence may be interpreted as evidence of silencing. Rejection of a social arrangement may count as resistance, while acceptance may count as internalization.
Any one of these explanations may be correct in a particular case. The methodological problem appears when the framework has enough interpretive degrees of freedom to assimilate every observation, which leaves it unable to discriminate among competing explanations.
A framework that rarely risks losing will rarely be forced to change.
Activism changes the loss function
Research always operates under distorted incentives, and no discipline possesses a perfectly neutral error function. Physics has status competition, medicine has commercial incentives, economics has ideological priors, and publication systems across academia reward novelty and positive results.
The test is whether a field contains mechanisms capable of counteracting its characteristic distortions, or instead builds those distortions into the standards of inquiry.
Some scholarship states its aims explicitly: to center marginalized voices, expose oppression, destabilize categories, challenge dominant structures, or advance social justice. These may be defensible moral and political objectives, but they are not epistemic criteria. They are also not motives imputed by critics; they are printed in mission statements, program descriptions, and methods chapters.
Aims of that kind make errors morally asymmetric. Suppose falsely concluding that discrimination occurred is treated as regrettable but cautious, while falsely concluding that discrimination did not occur is treated as perpetuating oppression. The threshold for accepting one class of explanation will fall while the threshold for accepting its alternatives rises.
This need not involve dishonesty. Researchers can sincerely believe they are being appropriately attentive to power while biasing inference in one direction, which is why the safeguards have to be institutional rather than personal.
Political efficacy and truth can coincide, but they can also diverge. A true claim can obstruct a political objective and a false one can advance it; social influence is no evidence of truth.
A knowledge-producing institution therefore needs an answer to a simple question: when the politically useful conclusion and the better-supported conclusion conflict, which one wins? A field whose methods encode one preferred answer has weakened its ability to correct itself.
Conceptual advances still need constraints
The obvious response is that some critical disciplines are not primarily empirical sciences. Their principal contributions may be conceptual, interpretive, historical, or normative, and it would be a mistake to judge them by laboratory standards.
Agreed. But conceptual inquiry still needs constraints that do not disappear whenever they threaten the preferred conclusion.
Return to intersectionality, which passes the initial epistemic test. It identified cases in which treating race and sex as independent categories obscured discrimination affecting people who occupied both, and a classificatory improvement is real when the new concept represents cases the previous scheme systematically mishandled.
What happens after the concept is introduced is a separate matter. Does it continue to sharpen distinctions, or does it expand until almost any social difference can be redescribed as an intersectional effect? Does it exclude some interpretations, or merely provide an additional vocabulary in which all observations can be narrated?
A concept earns epistemic authority by making reasoning more constrained, not merely more expressive.
The same applies to performativity, standpoint epistemology, structural oppression, and other characteristic ideas. Whether they are empirical hypotheses is beside the point. The question is whether adopting them reduces uncertainty by ruling out bad explanations, or merely increases the number of coherent stories one can tell.
Humanistic inquiry does not escape epistemic standards by being humanistic. Historians cannot decide that Napoleon won at Waterloo because the interpretation is emancipatory, philologists cannot choose a manuscript reading solely because its politics are attractive, and philosophical arguments remain subject to contradiction.
“Not science” cannot mean “nothing could show us that we are wrong.”
The constraint test
The most general diagnostic is weaker than empirical falsifiability. It asks whether inquiry contains constraints that the investigator cannot reinterpret away when they become inconvenient.
For physics, the constraint may be an experiment; for history, archival evidence; for mathematics, the proof obligation; for law, precedent and statutory text; for philosophy, contradiction and argumentative counterexample. These constraints differ radically, but they have a common structure: the preferred conclusion does not get the last word.
That suggests a third test for an academic field: what prevents its central claims from surviving indefinitely through reinterpretation?
The three tests are one requirement seen from three positions. A discipline must expose its beliefs to constraints partly independent of those beliefs. Output asks whether such constraints ever produced a new belief, revision asks whether they ever forced an old one to change, and constraint asks what stops a belief from immunizing itself.
A healthy discipline should be able to point to cases where scholars wanted a claim to survive and some constraint defeated it. The constraint need not be empirical, but it must be independent enough of the preferred framework to force an update.
Critical interpretive systems become epistemically suspect when contrary evidence does not defeat the framework but merely supplies new material for it. If every objection can itself be redescribed as evidence of privilege, power, resistance, internalization, or ideological capture, then the framework has become unusually difficult to escape.
Cargo can still arrive
The earlier Boghossian Principle argument asked whether academic institutions preserve the conditions for criticism and error correction; this one asks whether the expected products are visible. A field that discourages criticism should accumulate fewer forced revisions, because favored claims are protected from the processes that would narrow or kill them. A field whose central assumptions absorb any contrary observation should produce fewer decisive updates, because evidence has less power to discriminate among explanations.
A discipline can still produce useful insights under those conditions. It can notice neglected populations, invent serviceable concepts, recover archival evidence, or identify questions other fields failed to ask. Cargo can arrive even where the runway is ceremonial.
The institutional question is whether the process reliably steers toward better models and away from worse ones. Isolated successes do not establish that.
Perhaps these fields are misclassified
None of this implies that every activity housed within critical academic disciplines is worthless. Political philosophy is legitimate, cultural criticism can be insightful, concept engineering can improve our vocabulary, and recovering neglected voices or archival material can add historical knowledge.
The mistake is treating all of these activities as if they carried the same epistemic warrant. Universities possess unusual authority because some forms of inquiry reliably tell us things that remain true regardless of whether we would prefer them to be otherwise.
When a medical researcher finds that a treatment increases mortality or an engineer concludes that a bridge design will fail, we give those claims special weight because constraints independent of the desired conclusion discipline the inquiry. Critical scholarship often receives the same institutional markers: “research shows,” “studies demonstrate,” “scholars have established.” Those formulations imply a process capable of excluding alternatives.
If an academic practice primarily generates interpretations within a normative framework, it may still produce worthwhile work. But the relevant line does not fall between empirical and non-empirical inquiry. Political theory and intellectual history can be as tightly constrained as any experiment, and advocacy can contain excellent empirical research. The line falls between inquiry constrained independently of what its practitioners prefer and interpretation whose framework can always preserve itself. Which side a practice sits on determines what its methods warrant, not whether it deserves to exist.
AI exposes the constraint
Generative AI makes the distinction visible. Give a capable model a policy document or a historical event, supply a vocabulary of power, coloniality, and discourse, and ask for a sophisticated critical interpretation; it will produce one immediately.
AI is not itself the test. A model can produce a plausible interpretation without producing a good one, exactly as it can produce a plausible proof that is invalid, and telling the difference still requires the constraints under discussion. What cheap fluency does is make the location of those constraints easier to see.
A historian cannot make an archival document exist. A lawyer cannot invent binding precedent. A mathematician cannot stipulate that an invalid proof is valid. Where scholarship is bottlenecked by measurement, archives, proof, prediction, or logical consistency, fluent generation substitutes for only part of the work. Where the valued operation consists largely of applying a reusable interpretive vocabulary to new material, it substitutes for a great deal more.
AI does not sort science from the humanities. It sorts practices by how much of them survives cheap fluency.
Postscript
Universities routinely evaluate fields using publication counts, citations, grants, graduate enrollment, and institutional prestige. A different audit is available. What does the field know now that it did not know twenty-five years ago, and which distinctions, discoveries, or corrections arrived earlier because it existed? Which of its claims or frameworks were narrowed, rejected, or replaced under pressure, and which of those changes were forced rather than fashionable? What stops a preferred interpretation from surviving indefinitely, and what evidence, contradiction, failed prediction, or argumentative defeat would make the field relinquish a central claim? And which of its results became established enough that researchers outside the field now rely on them without adopting its theoretical commitments?
These are not hostile questions. A successful discipline should welcome them as opportunities to display accumulated knowledge, demonstrated self-correction, and epistemic discipline. A field that has existed for decades should be able to fill pages with answers. If instead they elicit methodological defenses, definitional disputes, appeals to the moral importance of the subject, explanations of why no independent constraint is appropriate, and criticism of the motives of whoever asked, the evasions themselves become informative.
“Name three discoveries” sounds like a taunt only because academia has become accustomed to judging fields by the sophistication of their discourse and the density of their institutions. Once “discovery” is read generously enough to include conceptual advances, historical reconstruction, and the correction of neglected questions, some critical disciplines can produce legitimate answers. Almost any large intellectual institution will occasionally produce something useful, which is why the concession costs so little.
Healthy inquiry leaves evidence of resistance. Some ideas survive; others are narrowed, abandoned, or replaced because the old position stopped being defensible. A field in which every framework can be redescribed into continued relevance has not demonstrated error correction; it has demonstrated interpretive resilience.
The university did not acquire epistemic authority because academics use specialized vocabulary, publish journals, and cite one another. Those practices acquired authority because sustained inquiry can force us to believe things we would not otherwise have believed and force us to stop believing things we wanted to keep.
Every discipline claiming the authority of research should occasionally have to demonstrate both, and name what has the power to tell it no.
Show us the cargo.
Show us the revisions.



