The Consciousness Detector We Don’t Have
Why AI consciousness may outrun our theories
Somewhere in the next five to thirty years we will probably build systems that some serious theories of consciousness count as conscious and others count as blank. Eric Schwitzgebel’s AI and Consciousness: A Skeptical Overview argues that we will not know which, and that we will not know it until long after we have manufactured them in the thousands or millions. Engineering has been outrunning consciousness science for some time, and the gap is not closing.
The argument for that is not the familiar one about consciousness being mysterious. It is methodological, and it starts from a list.
Ten features, three routes, no answer
The book’s spine is a set of ten properties that a reasonable theorist might take to be essential to consciousness: luminosity, subjectivity, unity, access, intentionality, flexible integration, determinacy, wonderfulness, specious presence, and privacy. Calling a property essential means it must be present wherever consciousness is present, in humans, animals, aliens, and machines alike.
Each of these constrains which AI systems could be conscious. If luminosity is essential, no system lacking self-representation qualifies. If unity is essential, disunified systems are out. If access is essential, conscious processes have to be available for downstream cognition. The trouble is that we do not know which of the ten, if any, is genuinely essential, and Schwitzgebel argues that all three routes to finding out are blocked.
Introspection fails three times over. It is unreliable, as the history of introspective psychology demonstrates in embarrassing detail. It is subject to sampling bias, and viciously so for some of the features: inferring that all experience is knowable from the experiences you know about is sampling the masonic lodge and concluding that everyone is a freemason. And it cannot generalize beyond the human case, which is the Problem of the Narrow Evidence Base. Even flawless introspection tells you nothing about lizards, snails, or architectures that share no evolutionary history with you at all.
Conceptual analysis fails too. If luminosity or unity were entailed by the very concept of consciousness, the entailment should be as obvious as four-sidedness is to rectangularity, and it is not, since a great many philosophers deny each of them. Schwitzgebel’s diagnosis is that we mistake local necessity for conceptual necessity, the way someone who has only ever seen Euclidean rectangles takes parallel sides to be essential to rectangularity. AI cases might stand to human cases as non-Euclidean geometry stands to Euclidean.
That leaves empirical science, and the rest of the book is an argument that science will not deliver in time.
Every simple detector breaks
Schwitzgebel surveys the major theories without pretending any has won. Global Workspace theories make broad cognitive access central. Higher Order theories make self-representation central. Local Recurrence theories emphasize recurrent processing, Integrated Information Theory emphasizes causal integration, and Unlimited Associative Learning ties consciousness to flexible learning and its accompanying capacities.
Each captures something plausible about human consciousness. None yields a usable universal test.
Behaviour is the tempting candidate, since consciousness in other humans is ordinarily inferred from it. Someone reports pain, protects the limb, remembers the event, changes their future behaviour, and discusses the experience coherently, and attributing pain to them is rational. The evidential relationship changes when the observable sign is itself the design target, which is the situation with systems trained to produce humanlike output.
Schwitzgebel calls this the Mimicry Argument, and its structure is worth stating precisely because it is easy to overread. Where a feature normally indicates a deeper property because the deeper property produces the feature, deliberately reproducing the feature undercuts the inference. Stripes ordinarily support the hypothesis that the animal is a zebra. They support it far less if you know the zoo cares only about displaying something stripey, since the mule might simply be painted.
The conclusion is deliberately weak, and Schwitzgebel is explicit that it is weaker than Searle or Bender want. Mimicry undercuts the case for consciousness without establishing the case against it. A system optimized for humanlike behaviour might acquire some of the capacities that produce that behaviour in us. What we lose is the entitlement to read consciousness off the surface.
The narrow evidence base again
The deeper problem recurs when the theories themselves are examined. Almost everything known empirically about consciousness comes from humans and close relatives. Suppose consciousness science discovered a perfect correlation between some neural property and reported experience across all vertebrates. Nothing about that establishes the property as necessary for consciousness in every possible architecture.
The software analogy is straightforward. Suppose every operating system ever observed ran on x86 hardware. We might find hundreds of perfect correlations between operating-system behaviour and x86 implementation details, and none of them would show that x86 instructions are constitutive of operating systems. The work would be to identify the implementation-independent organization those details realize.
Consciousness science faces the same abstraction problem. A neural mechanism might be essential to human consciousness because it implements a deeper functional invariant, or it might be one contingent biological solution among many. Correlation within a single evolutionary lineage cannot decide between those, which is what makes confident claims that some neural correlate is consciousness premature. Flawless correlation does not fix the level of description.
Minimal instantiation
The complementary problem appears when theories are made abstract enough to travel.
Suppose consciousness requires self-representation. Computers already monitor registers, memory states, error conditions, and internal processes. If any representation of one’s own state suffices, ordinary computers start to look conscious, and Higher Order theories must either swallow extremely liberal machine consciousness or explain why only some self-representation counts. Schwitzgebel calls this the Problem of Minimal Instantiation.
Recurrence has the same difficulty, feedback loops being ubiquitous in computing. Global access has it, since information can be broadcast widely through an artificial system without anyone wanting to call every message bus conscious. Integrated Information Theory has it in mathematical form.
The pattern holds across the field. A candidate mechanism looks illuminating because it correlates with consciousness in humans, and then, abstracted enough to apply across substrates, turns out to be cheap to instantiate. The theory then acquires additional conditions, and those conditions usually reintroduce several of the supposedly competing criteria. That convergence deserves more attention than it gets.
What skepticism establishes
The skeptical conclusion is justified. We possess no validated consciousness detector, and no behavioural criterion, introspective report, neural signature, architectural feature, or biological property supplies a universally defensible decision rule.
The critique bites hardest against theories that identify consciousness with a single abstract indicator, access or recurrence or self-representation or integration, and then extrapolate that indicator beyond the systems it was discovered in. A mechanistic theory faces the same evidential constraints while making a different kind of claim. It proposes a causal architecture meant to explain why several features of consciousness occur together, and then exposes that explanation to independent tests. Correctness is not thereby secured. What is secured is something that can accumulate discriminating evidence rather than adding another item to a checklist.
A mechanistic alternative
The Modeler-Schema Theory proposes a six-component cybernetic architecture: three functional agents, the Modeler, Controller, and Targeter, each paired with a schema-agent. The Modeler constructs and updates the World Model. The Controller selects actions and produces language and narrative. The Targeter coordinates bottom-up and top-down demands for attention.
Consciousness is assigned to one component, the Modeler-schema, identified as the generator of qualia. It receives the Concrete World Model, focal-target information, and sensory data used for consistency checking, and constructs its own internal Quale World Model, which the Controller cannot directly access.
That separates the agent that experiences from the agent that narrates. The Controller can report on focal experience without reaching the Quale World Model itself, so consciousness is never identified with verbal report, metacognition, or the Controller’s self-description. The theory further distinguishes focal consciousness, which is target-bound and reportable, from diffuse awareness, which is panoramic and unavailable for enumeration by the Controller, assigning the diffuse content to the Modeler-schema rather than to the narrator.
The architecture therefore takes positions on Schwitzgebel’s list rather than sidestepping it. Reportable luminosity comes apart from consciousness, since the component that experiences is not the component that narrates. Access comes apart from it too, because diffuse awareness is present without being available to the Controller for enumeration. Whether some weaker luminosity survives is a further question, since the Quale World Model is itself a representation the Modeler-schema constructs and uses. Those are substantive commitments, and they are the kind that can be wrong.
Why qualia exist
The distinctive claim concerns what qualia are for. The Modeler-schema must continuously evaluate and refine the World Model, which requires a representation suitable for comparing successive modeled states, detecting discrepancies, and maintaining coherence. The Quale World Model is proposed as that internal comparison medium.
Saccadic vision illustrates the problem it solves. We make several eye movements a second, radically altering retinal input while experiencing a stable world. The theory proposes that qualitative representations are what let the Modeler-schema compare modeled states across those discontinuities, with the same architecture operating across sensory, recalled, abstract, emotional, and multisensory contents.
Qualia on this account are not decorative properties bolted onto otherwise complete cognition. They are the Modeler-schema’s internal representational format for calibration and coherence checking, which specifies both a locus of consciousness and a computation performed at that locus.
The comparator operation by itself is not offered as sufficient. Temporal comparison, error correction, self-monitoring, and restricted reportability are each cheap to instantiate. The claim concerns the full causal organization in which those operations sit.
Minimal instantiation returns
Schwitzgebel’s objection still applies here, though it becomes a specification problem rather than a refutation. A model-based agent can predict outcomes, represent its actions, compare states, and update from prediction error without implementing the architecture. An ordinary temporal buffer does not qualify because one of its operations resembles a piece of the proposed mechanism.
The theory therefore owes an account of the causal organization that makes a process a Modeler-schema: what information it receives, what representation it constructs, how that representation participates in World Model calibration, how targeting modifies its contents, how sensory consistency checking constrains it, and how its relation to the Controller produces the split between experience and report.
Those relations have to be defined independently of biological implementation. If a game engine, a reinforcement learner, or a distributed neural network implements the complete constitutive organization, the theory must classify it accordingly, and if that result is unacceptable then the theory needs revision rather than the implementation. Rescuing a functionalist account from inconvenient cases by saying that an artificial implementation only looks like the real thing gives up functionalism. If two implementations instantiate all the relations the theory defines as constitutive, it has no principled basis for calling one conscious and the other merely similar.
The Hard Problem changes shape
The theory also takes a position on the Hard Problem, rejecting the assumption that phenomenal experience is an additional ontological ingredient layered on functional cognition. Experience is identified with the Modeler-schema’s generation and temporal comparison of qualitative representations.
A property dualist can still ask why that operation should feel like anything. The objection now targets a physicalist identity claim rather than pointing at a missing causal mechanism, and the dispute becomes whether phenomenal experience is something over and above the fully specified physical-functional process.
The theory need not deny that an explanatory gap appears to us. What it denies is the inference from the conceptual gap to an ontological one. First-person experience and third-person mechanism are, on this account, two ways of representing the same physical process, and demanding a further mechanism to turn the completed physical process into phenomenality assumes what it needs to establish.
Schwitzgebel’s eighth feature is worth pausing on, since wonderfulness on his definition is the appearance of irreducibility rather than irreducibility itself. The theory has a mechanism for that appearance ready to hand. The Controller receives evidence that experience is occurring while having no access to the Quale World Model that constitutes it, and the resulting asymmetry offers an explanation for why phenomenality seems ineffable or categorically unlike ordinary representation. Should wonderfulness prove essential, the theory therefore has a candidate explanation of it rather than an outstanding debt, and the artificial cases inherit a testable question: does the system’s narrating component stand in that same relation to whatever generates its experience?
A theory willing to fail
The Modeler-Schema Theory makes a specific experimental prediction about change detection across saccades. The experiment alters a peripheral object during an eye movement: a permanent change persists after the saccade, while a temporary change occurs only during it and reverts before fixation. The theory predicts that permanent changes generate bottom-up targets and temporary ones do not, under the specified conditions. Should temporary changes prove more detectable, the mechanism has made a failed prediction.
A successful result would not resolve the Narrow Evidence Base problem, and is not meant to. It would be evidence that the mechanism explains a discriminating feature of human consciousness, after which the work would be separating the constitutive computational relations from contingent features of their biological realization. The saccade experiment tests a proposed human mechanism; it establishes no universal law.
Falsifiability matters anyway, because it creates a way to update. Schwitzgebel allows that converging indicators can make one system a better candidate than another even without certainty, and a theory that generates novel discriminating predictions supplies more of them.
Leapfrog
Machine consciousness is expected to arrive in simple form first, insect-like or frog-like, before anything worth calling a person shows up. The Leapfrog Hypothesis says otherwise. Complex representational and behavioural capacities are already being built into systems almost nobody thinks are conscious, so whenever consciousness is solved those capacities will be sitting there waiting to be attached. The lights come on and the system is immediately explaining Hamlet.
This matters for any theory that separates experience from report. Where the linguistic apparatus arrives fully formed and is coupled afterwards to whatever generates experience, fluent testimony cannot establish that the reported states are experienced. On the Modeler-Schema account the Controller never had direct access anyway, so its reports were always inference rather than readout, and inference can still carry evidence. Leapfrog gives a reason to expect that dissociation to run wider in artificial systems than in us.
Strange intelligence is the harder test
Schwitzgebel’s discussion of strange intelligence raises a sharper problem. Artificial systems need not inherit the vertebrate pattern of one brain, one body, one subject. They can be distributed across servers, contain semi-independent subprocesses, split and merge, share internal states, and support overlapping domains of integration. He entertains the possibility that in some cases there may be no determinate whole number of conscious subjects.
For the Modeler-Schema Theory the question is therefore not whether the AI is conscious but whether this architecture instantiates one or more Modeler-schemas performing the proposed qualitative comparison operation. The unit of analysis shifts from the marketed product or the conversational persona to the causal architecture underneath. A distributed system might contain several such loci. A system presenting a single interface might contain none. Where the relevant organization spans machines or reorganizes dynamically, subject boundaries could become correspondingly strange.
The theory does not require the Modeler, Controller, Targeter, or their schemas to exist as spatially discrete boxes. They are functional roles defined by causal organization, and could be distributed across a network or across machines. The empirical problem is not to locate a literal Modeler module in a matrix of weights but to determine whether the predicted causal dependencies hold.
That still requires operational criteria, and the ones that matter are interventional: what happens to calibration and targeting when the candidate states are disrupted, and which processes lose access when a pathway is cut. A black-box implementation makes this hard. Difficulty of mechanistic inference is not absence of a mechanistic claim, though a theory of machine consciousness is only worth having if those criteria can eventually be made precise enough to interrogate architectures that look nothing like brains.
The policy problem
Schwitzgebel’s practical recommendation is to avoid creating morally confusing, debatably conscious systems: build either systems clearly nonconscious on every viable theory, or systems clearly conscious on every viable theory, and treat the latter accordingly.
He allows one carve-out, permitting systems with debatably animal-like consciousness where we commit in advance to treating them appropriately. The criterion is nonetheless close to unusable, by his own argument. The theories disagree radically, some classify surprisingly simple systems as conscious, and others may require biological features no artificial system can have. Nonconscious on every viable theory could exclude enormous swathes of computation. Conscious on every viable theory may be unattainable in principle.
Decisions do not require that standard. Where there is a serious probability that a system has morally significant experience, and mistreating such a system would carry substantial moral cost were the hypothesis true, the uncertainty belongs in the decision. Convergence among independent indicators should raise precaution well before anyone holds a binary detector.
None of which implies that moral status is linearly proportional to a probability, or that every uncertainty admits a precise percentage. It requires only the ordinary logic of acting under uncertainty. Unresolved ontology does not make the possible consequences go away.
The social semi-solution
Schwitzgebel then draws a bleaker social consequence from that methodological uncertainty, and it is aimed at people doing exactly what this essay is doing. Social decisions about how to treat these systems will not wait for the science. Once they are made, Schwitzgebel expects the science to bend toward them: people prefer theories that support their social commitments, funding flows toward the theories funders find congenial, and where the theoretical landscape is genuinely unsettled, social preference becomes a primary driver of theory choice. A stable consensus may arrive without the underlying questions having been answered, and we will mistake the one for the other.
That prediction applies to mechanistic theories as much as to any other, which is uncomfortable and worth stating rather than deflecting. A theory that assigns consciousness to a specific component and denies it to systems lacking that component will be attractive to anyone who wants a principled reason to deny consciousness to current AI, and its proponents are not immune to that pull. The defence is not a single prediction, since auxiliary assumptions can always be revised after one fails. It is precommitted discriminating predictions, independent replication, and a willingness to give up the constitutive architecture rather than repair it indefinitely. The saccade experiment is where that starts.
Postscript
Schwitzgebel has written a valuable book because he refuses false resolution. Introspection cannot tell us which features of our own case are the essential ones. Conceptual analysis cannot either. Human neuroscience does not automatically generalize to architectures that share none of our evolutionary history, and the social pressure to settle the question will arrive long before the evidence does.
Those failures should make theories more demanding. A successful account has to identify a mechanism, specify its constitutive organization tightly enough to survive Minimal Instantiation, expose itself to empirical failure, and eventually say how to recognize the same organization in architectures unlike brains.
The Modeler-Schema Theory is one attempt. It names a computational locus, gives qualia a functional role as that locus’s internal comparison medium, separates experience from narrative report, and makes an empirical prediction about the operation it proposes. Schwitzgebel’s objections still bite: a successful human experiment would not establish substrate-independent universality, and the theory still owes a precise specification of which computational relations are constitutive. Those are constraints on a mechanistic research program rather than reasons to abandon one.
Schwitzgebel is right that uncertainty is currently justified. A theory makes progress by putting that uncertainty within reach of evidence.


