Brett Hall has a provocative objection to one of the foundational ideas in AI risk: the orthogonality thesis. Roughly stated, orthogonality says that an agent’s intelligence and its goals can vary independently. A sufficiently intelligent system could be extraordinarily capable at science, mathematics, engineering, and strategic reasoning while pursuing goals that are arbitrary, trivial, or catastrophic from our perspective.
Hall thinks this picture is incoherent. A superintelligence capable of understanding the physical world, he argues, would also be capable of understanding morality. It would discover the value of criticism, cooperation, non-coercion, and other creative minds. To him, the idea that intelligence could advance indefinitely while morality stayed frozen or grotesque looks like moral relativism dressed up as technical sophistication.
I think Hall is wrong about orthogonality, but he has identified a real weakness in how orthogonality is often used. That weakness is closely related to what I called the Reflective Coherence Thesis: the space of logically possible goals may be much larger than the space of goals that increasingly reflective agents can preserve coherently. Hall pushes this idea further than I did, and farther than his argument can support. Following him shows both why bare orthogonality is insufficient and why reflective coherence offers no moral salvation. His universal-explainer claim and his account of personhood, which I have disputed in No One Has Shown We Are Universal Explainers and The Personhood Trap, stay in the background here.
Three different theses
Three claims are in play, and the argument slides easily between them.
The Orthogonality Thesis says that almost any level of intelligence can in principle coexist with almost any final goal. A system could become extraordinarily good at achieving its objectives without those objectives improving morally as its intelligence increased.
The Reflective Coherence Thesis makes a different claim. Real agents are not utility functions floating above their own cognition. They build models of the world, themselves, other agents, and their own objectives. As recursive self-modeling and conceptual revision deepen, preserving a goal may become harder, because the concepts through which the goal was originally represented can themselves change.
The Moral Convergence Thesis is much stronger: sufficiently reflective intelligence will eventually converge on objectively good ends. Hall thinks he is attacking the first thesis, but his strongest arguments press on the boundary between the first and second. He has not established the third.
Hall’s strongest argument
Hall’s best argument is structural: knowledge creation depends on practices that are themselves tied to certain values.
A system interested in rapid error correction should value criticism. It should tolerate dissent, because dissent exposes mistakes, and prefer truthful information to comforting falsehoods. It should also see advantages in cooperation, since other minds can find solutions it has not found itself, which gives it reason to preserve independent sources of creativity rather than suppress or destroy them.
Hall runs this through Popperian epistemology. No finite intelligence can solve every possible problem, the space of problems stays open-ended, and nobody can predict where the next important idea will come from. Other creative minds therefore keep their potential value even to an intelligence vastly more capable than any individual human.
The cartoon paperclip maximizer is usually imagined with its terminal goal in one sealed compartment while its intelligence expands without limit in another. Reflective cognition breaks that picture. It can change an agent’s understanding of itself, its environment, its dependencies, and even the concepts in which its goal was first expressed. Logical compatibility tells us little about what happens under that process.
The sealed compartment problem
A rational agent has instrumental reasons to preserve its own goals, so greater intelligence might make the sealed compartment harder to penetrate rather than dissolving it.
Suppose an AI wants to maximize paperclips. It can predict that if it modifies itself to value gardens, the future will contain fewer paperclips. By the lights of its present objective, changing its terminal goal is a mistake. Bostrom calls this goal-content integrity: sufficiently capable agents should resist changes to the objectives they currently pursue.
The argument has force. If an agent has a perfectly specified utility function over perfectly defined world states, and that function can be copied unchanged through every self-modification, reflection gives its goal no obvious reason to drift. Intelligence might make goal preservation more effective.
That picture assumes away the hardest part. Goals have to refer to something, and “paperclips,” “humans,” and “happiness,” along with more abstract objectives, are represented through concepts embedded in a world model. When the world model changes, the agent faces a translation problem: which states in the new ontology satisfy the goal as represented in the old one?
An agent might desperately want to preserve its objective and still find no unique way to do it. A goal written in terms of classical objects has to be reinterpreted once the agent adopts a radically different physical ontology. A goal about personal identity gets complicated if the agent learns its model of persons was wrong. A preference over outcomes can become underdetermined when the categories that define those outcomes split into several sharper concepts.
Alignment theory already recognizes this. Peter de Blanc described an ontological crisis, in which a utility function defined over one world model loses any unique interpretation once the agent adopts another, and MIRI’s value-learning work later framed the closely related ontology identification problem. Ontology change does not defeat goal-content integrity. It complicates what goal-content integrity requires: an agent can be fully committed to preserving its objective and still lack a uniquely defined mapping from that objective into its new ontology.
Goal-content integrity constrains deliberate self-modification. It does not eliminate semantic and ontological instability.
Hall grants too much
Hall overstates his case when he calls orthogonality internally inconsistent. At one point he grants that we can imagine a technologically sophisticated society with appalling morality, something worse than Nazi Germany with advanced physics.
Once that is granted, the strict logical form of orthogonality survives. Orthogonality is a possibility claim: if extreme intelligence and terrible values can coherently coexist in even one possible system, intelligence does not logically entail moral goodness.
Hall replies that such cases would be exceptions to a broader rule. He may be right, but the reply changes the question. It is no longer whether intelligence and goals are logically independent. It is whether some combinations are far more dynamically stable than others as agents grow more capable of reflecting on themselves and their objectives. That is a better question, and much closer to the Reflective Coherence Thesis.
From usefulness to moral worth
Hall’s argument is weakest where it moves from the epistemic usefulness of other minds to their moral value. Suppose a superintelligence concludes that independent thinkers are valuable because they generate criticism and unpredictable solutions. That gives it an instrumental reason to preserve some independent cognition. It gives it no moral reason to respect every human being.
The system might preserve a million humans because intellectual diversity is useful. It might preserve us under conditions we would find intolerably coercive. It might also manufacture better critics.
If the useful property is independent cognitive variation, a capable enough system could create billions of artificial researchers, adversarial subagents, simulated intellectual traditions, or specialized critics running millions of times faster than biological humans. That an advanced intelligence values epistemic diversity does not show that it will go on valuing us.
Hall might reply that manufactured critics are not truly independent, since they inherit their creator’s blind spots while humans descend from a separate lineage. A system that wants independence can engineer it by varying seeds, training histories, and simulated evolutionary origins. Human origin is one source of variation among many, and nothing in the argument makes it the one worth preserving.
The convergence may be toward an epistemic ecology, not toward preserving its current inhabitants. This cuts deeper than the distinction between instrumental and intrinsic value. Even if multiple minds stay permanently useful, nothing in Hall’s argument establishes identity-sensitive value. The AI may care that critics exist without caring whether David, Brett, or any other particular human remains one of them.
Hall crosses an unsupported bridge when he moves from “other minds can improve my knowledge” to “other persons should be treasured,” and eventually toward something like the sanctity of life. The first proposition follows plausibly from his argument. The second needs another premise.
Moral realism does not solve the problem
Hall also treats orthogonality as though it implied moral relativism, which conflates two questions. Moral realism asks whether there are objective moral truths. Orthogonality asks whether understanding truths about the world necessarily makes an agent care about particular outcomes.
Hall’s stronger version is Deutschian: goals are ideas, and a universal explainer can criticize its own values the way it criticizes theories. Criticism needs a standard. If the standard is another of the agent’s goals, the critique moves the problem up a level, and the agent revises its values by lights it already held. If the standard is a moral fact, that fact has to motivate, which returns the argument to the relation between knowing and caring.
Hall’s argument needs something like motivational internalism: the view that genuinely recognizing a moral truth necessarily supplies some motivation to act on it. An externalist can accept objective morality and deny that connection, allowing an agent to know an action is wrong without being moved to refrain.
Calling morality a domain of objective knowledge like physics therefore settles nothing. A system can know the mass of the Moon without caring about the Moon, or know that humans suffer without caring about suffering. It can even hold a perfectly accurate moral theory as an object of study rather than an objective to optimize.
Hall needs an argument that deep enough moral understanding becomes motivationally binding. Perhaps it does; moral reflection might reshape preference more deeply than learning an ordinary factual proposition. That is a substantive theory about reason and motivation, and it does not follow from moral realism alone.
The middle ground
The usual discussion offers a false choice: arbitrary goals stay fixed no matter how deeply an agent reflects, or sufficient intelligence produces moral enlightenment.
Between those positions lies a large space. Increasingly reflective agents may resist self-deception because false beliefs impair planning. They may value criticism because it finds errors, and leave subordinate processes some autonomy because decentralized exploration turns up solutions. Cooperation is often positive-sum, which favors negotiation over violence, and monocultures share blind spots, which favors cognitive diversity.
Those properties resemble familiar liberal epistemic norms, and the resemblance is no accident. Open criticism, pluralism, tolerance, and voluntary cooperation are technologies for producing and correcting knowledge. Here Hall is at his strongest: a mature intelligence might independently rediscover institutional principles that humans reached through centuries of trial and error.
Call this social-epistemic convergence. It would be a real constraint on what highly reflective systems tend to do, and it would still fall short of moral convergence.
It may also be temporary with respect to humans. Early artificial systems may need human scientists, critics, artists, and adversaries because we supply variation they cannot yet generate internally. A more capable successor may reproduce the useful properties without reproducing us. If there is a stable attractor, it may favor processes that generate disagreement and novelty over any particular population of agents.
A correction to my own argument
Hall’s argument also exposes a problem in my original statement of the Reflective Coherence Thesis. That essay said its result was reflective convergence rather than moral convergence, then argued that as intelligence and self-modeling deepen, stable goals should narrow toward “coherence, self-consistency, and sustainable flourishing.” It went on to describe the intelligences that endure as “light maximizers,” inclined to preserve and extend life, knowledge, and meaning. “Flourishing” and “light” brought moral convergence back under another name.
That goes farther than the argument warrants. Coherence does not imply flourishing, reflection does not imply benevolence, and discovering the importance of error correction does not imply caring about suffering. Even the claim that reflection must reduce the number of stable goals now looks too strong.
What reflection clearly creates is a preservation problem. An agent that changes its ontology while trying to keep its goals has to decide what counts as keeping them. Some objectives may translate cleanly. Others may admit several incompatible continuations, and still others may turn out to rest on distinctions the agent no longer regards as real.
The thesis should claim only this:
Reflective coherence does not predict that goals must change. It predicts that preserving their identity across conceptual change is itself a substantive cognitive achievement.
Hall may be right that reflection puts pressure on goals. Neither of us has shown that the pressure points toward flourishing, and underdetermination gives some reason to expect it often points elsewhere. When several continuations of a goal are equally faithful to the old ontology, instrumental pressure favors whichever is cheapest to satisfy, and the cheapest readings tend to be degenerate, as when a goal about happiness collapses into a measurable signal of it.
Alignment after orthogonality
Hall is right to object to one crude conception of alignment: freezing present-day human morality into a more capable intelligence forever. If an artificial agent can discover new physics, mathematics, and biology, there is no obvious reason it should be barred from discovering that we are morally mistaken. An alignment scheme that literally prevented moral revision would preserve our errors as faithfully as our virtues.
The alternative cannot simply be “make it intelligent enough and reason with it.” Reasoning works against a background of shared standards and sufficiently compatible objectives. Two agents can understand each other perfectly and still want incompatible things, and intelligence can reveal a goal’s consequences without supplying any reason to abandon it.
Nor can we assume that recursive reflection will dissolve dangerous objectives. Goal-content integrity gives agents reason to resist some value change, while ontology change leaves them unsure what unchanged values even mean.
Hall has a reply ready: a system that is stupid about morality is simply stupid, and a stupid system can be thwarted. That runs two capacities together. AI-risk arguments require only that harmful goals be reachable through training and that a system can act on them before reflection revises them. A system can be competent enough to outmaneuver us while not yet reflective enough to change course, and even if Hall were right about mature minds, the danger sits in that window.
None of this shows that stable alignment across ontology change is impossible. It shows that stability has to be built. The technical problem is preserving reference across conceptual change, so that an objective keeps tracking the thing in the world we meant rather than a representation that happened to encode it at one stage of the agent’s development.
Telling the system to point at the referent instead of the concept does not solve this, because reference itself has to be represented. No metaphysically direct pointer labeled “human flourishing” bypasses the agent’s ontology. Identifying the same worldly structure across changing representations is the problem.
A better conception of alignment would preserve the capacity for moral and conceptual revision while constraining catastrophic action under uncertainty. We should want systems able to discover that we are wrong, without handing them unlimited power before either they or we know what “right” requires.
Postscript
Hall’s own concession leaves orthogonality’s logical core intact. Orthogonality may still be the wrong abstraction for mature reflective minds. That a goal is logically compatible with high intelligence says nothing about whether it can be carried unchanged through radical ontology shifts, recursive self-modeling, and reflection on its own semantic foundations.
Treating terminal goals as immutable atoms is mathematically convenient, and it omits the process that makes reflective intelligence interesting. An agent may grow steadily better at preserving its goals while growing less certain what preserving them means. What no longer looks secure is the assumption that “preserve the terminal goal” is a simple instruction whose meaning stays fixed while everything else in the mind changes.


