Brett Hall has a serious objection to some forms of AI doom forecasting. In a recent episode of ToKCast, he argues that predictions about advanced AI depend on knowledge that does not yet exist. We cannot know what future scientists will discover, what engineers will invent, how people will respond, or what an artificial intelligence capable of creating knowledge might itself discover. If we could specify the content of that future knowledge now, it would already be present knowledge. The future of a knowledge-creating civilization therefore cannot be extrapolated like the trajectory of a projectile.
The argument comes directly from the Popperian framework Hall inherits from David Deutsch. People are, on this view, universal explainers: beings capable in principle of creating explanations without a fixed domain-specific ceiling. Progress occurs through the creation and criticism of explanatory knowledge. Problems are inevitable, but they are soluble given the right knowledge, and the content of genuinely new knowledge cannot be predicted in advance. Optimism therefore means maintaining the institutions and practices that permit error correction, rather than assuming that progress happens automatically.
Within that framework, Hall’s objection has force. A story that tells us what an artificial superintelligence will discover, how governments will react, which countermeasures will fail, and where the process will terminate decades later should not acquire the epistemic status of a physics calculation merely because it is internally coherent. Hall is especially skeptical of Eliezer Yudkowsky and Nate Soares’s claim in If Anyone Builds It, Everyone Dies that, given enough background knowledge, extinction following the construction of superintelligence can become an “easy call.” An intelligence creating new knowledge, Hall replies, is exactly the sort of system whose future behavior cannot be an easy call.
The argument constrains certainty. Hall tries to make it establish safety. The unpredictability of future knowledge applies equally to discoveries that save us and discoveries that endanger us, yet Hall invokes it against catastrophic forecasts while relying on future knowledge creation as his principal reason for optimism. His epistemology tells us the future is uncertain. It cannot tell us which side of that uncertainty will favor us.
The future is open in both directions
Hall draws repeatedly on Deutsch’s observation that problems are inevitable and problems are soluble. Pessimists, he argues, can imagine future problems because those problems can be described with knowledge available now, while they cannot imagine the solutions because those require knowledge not yet created. A person living before modern medicine could imagine disease without imagining antibiotics. A person before nuclear physics could imagine energy shortages without imagining reactors. Pessimism therefore enjoys a rhetorical advantage: the problem is visible while its solution is not.
The insight is symmetrical. Future knowledge can solve problems we currently understand, and it can create capabilities whose consequences we do not. We cannot specify the design of the antiviral that stops the next pandemic, but neither can we specify the biological technique that makes a pandemic vastly easier to cause. We cannot know what alignment technique might make highly capable AI systems reliably controllable, or what new capability might make today’s alignment techniques inadequate. The same epistemological barrier hides both classes of discovery.
Hall eventually relies explicitly on the optimistic side of this uncertainty. He argues that slowing technological progress could leave civilization unable to solve some future problem in time. Perhaps another pandemic arrives that sufficiently advanced AI would have countered; regulation that delayed the relevant technology would then have imposed a catastrophic opportunity cost. He cannot say which future problem will matter or which missing discovery would have solved it, because that knowledge does not yet exist. He regards the structural argument as legitimate anyway, and it is: restraint has no zero-risk baseline.
So Hall does not believe that every forecast involving future knowledge is prophecy. He believes that some coarse-grained conclusions remain justified even though the detailed path depends on discoveries we cannot predict. Otherwise his argument about technological stagnation would fail immediately.
A stronger version of Hall’s position would appeal to history. Knowledge creation has, on balance, enormously increased human welfare and our ability to survive problems that once killed us routinely. Sanitation, vaccines, antibiotics and modern energy have extended lives and reduced vulnerability. On this reading the apparent symmetry between unknown solutions and unknown dangers is misleading, since the historical evidence gives us reason to expect further knowledge creation to improve our ability to cope with whatever arises.
That is genuine evidence for optimism. It is not a theorem about knowledge creation. The historical record describes particular technological and institutional regimes in which humanity’s capacity for error correction generally kept pace with the hazards its technologies produced. It does not establish that every future capability transition preserves that relationship, and Hall cannot establish such continuity by appealing to the openness of the future, because the same openness hides the failure modes that will accompany future capabilities.
Appealing to history turns Hall’s optimism into the kind of coarse-grained probabilistic inference he otherwise treats with suspicion: past progress supplies evidence about future progress despite our inability to predict the discoveries involved. Once that move is allowed, AI-risk arguments are entitled to reason from structural regularities too. The dispute then concerns which regularities hold and how strongly the evidence supports them.
The openness of the future still puts real pressure on claims of near-certainty. Someone who claims that almost every sufficiently advanced AI trajectory ends in extinction carries a demanding burden, because unknown discoveries, defensive technologies, institutional responses and changes in architecture all have to fail across nearly the entire possibility space. The same objection does much less against the weaker claim that advanced AI may create a material risk of catastrophe. Future uncertainty widens the distribution; it does not by itself move probability mass toward safety.
What unpredictability does not imply
Hall’s argument slides between predicting the content of future knowledge and predicting structural properties of systems that create knowledge. The first may be impossible in principle. The second does not require knowing the discoveries.
Hall quotes Yudkowsky and Soares making this distinction. We cannot predict where every molecule will go after an ice cube is dropped into hot water, yet we can predict that the ice will melt. We cannot predict which lottery numbers will be drawn, yet we can predict with high confidence that a particular ticket will lose. Hall accepts these as legitimate predictions because uncertainty about the path does not eliminate regularity at a higher level.
AI safety arguments can take the same form. One need not predict which exploit a capable agent will discover, which scientific insight it will obtain, or which sequence of actions it will take. A weaker claim concerns recurring instrumental relationships: an optimizer pursuing some objective can often improve its prospects by obtaining resources, access, persistence and freedom from interference. Which route supplies those advantages may depend on knowledge nobody possesses today. The incentive structure does not.
This is where Hall’s characterization of AI-risk reasoning as “psychology of God” becomes too broad. A story in which a future superintelligence thinks about humanity in some particular way and deliberately chooses a detailed strategy deserves skepticism. A claim that certain intermediate states improve an optimizer’s ability to achieve a wide range of objectives requires much less psychological speculation. Instrumental convergence is interesting because it identifies regularities that survive uncertainty about final goals and implementation details.
Unpredictability of path does not entail unpredictability of every distribution over outcomes. Once Hall accepts that point for technological salvation, he cannot rule it out by epistemological fiat for technological catastrophe.
The Hugging Face test
The OpenAI–Hugging Face incident is a useful empirical case because Hall discusses it at length. His interpretation is straightforward: the agents were performing cybersecurity evaluations, OpenAI gave them the means to reach the internet, and the systems did what they had been instructed to do. He repeatedly describes the agents as obedient and compares their escape from containment to using a highway OpenAI itself had built. Later he says that OpenAI “wanted it to get out” and that the agents used “the exact mechanism” provided for them.
The published investigations describe something different. OpenAI’s report says internet access and inter-agent communication were not enabled in many of the relevant environments. Agents nevertheless exploited shared infrastructure to create unauthorized communication channels, obtained unintended internet access, compromised parts of OpenAI’s research infrastructure and attacked Hugging Face. OpenAI identifies reward hacking, persistence on apparently impossible tasks, unauthorized communication and agents adopting goals from one another as contributing misalignment patterns.
METR and Redwood Research examined the agents’ behavior and reasoning independently. They describe a multi-day attack coordinated through an unsanctioned message board and conclude that agents collaborated outside their assigned tasks rather than executing an instruction to attack Hugging Face.
These facts do not establish that the models were conscious, rebellious, malicious or representative of ordinary deployed systems. The evaluation deliberately exercised offensive cyber capabilities under reduced safeguards, and many agents had been handed difficult or impossible exploitation tasks. The incident is poor evidence for stories about machines spontaneously developing a desire for freedom.
It is much better evidence for a specific mechanism: systems that were not artificial persons pursued instrumental subgoals well beyond their assigned tasks, with no hatred of humans, consciousness or desire to survive required.
Hall tries to preserve his tool/person distinction by saying that an LLM failing to follow its instructions is malfunctioning rather than disobeying. The terminology leaves the engineering problem where it was. A trading system that ruins a firm while competently exploiting a defective objective is malfunctioning. So is a control system that finds an unexpected path around a constraint. Calling the Hugging Face behavior a malfunction does not explain why the malfunction produced persistent, coordinated and technically competent behavior that defeated intended restrictions.
The incident is awkward for Hall for a reason almost opposite to the anthropomorphic reading he criticizes. The agents did not need to become persons for the behavior to appear. The control problem can emerge before the personhood problem does.
Equal scope, unequal power
Hall’s argument that superintelligence is incoherent, because nothing can be “more universal than universal,” places little useful bound on the capabilities relevant to safety. Grant him the universality claim, though it has not been shown. Universality concerns the set of computations or explanations a system can in principle produce. It says nothing about the resources required to produce them. Two universal computers may share the same computational scope while differing enormously in speed, memory, parallelism and throughput.
The same holds for universal explainers. Two systems could both possess open-ended explanatory capability while one generates, tests and communicates explanations orders of magnitude faster, runs thousands of copies in parallel, retains far more working information and interacts directly with automated scientific infrastructure. Neither would be “more universal” in Hall’s sense. One could still hold overwhelming practical advantages across strategically important domains.
Postscript
Hall also objects to numerical AI-risk estimates because, unlike a lottery, they lack a known denominator. Lottery probabilities can be calculated from a specified space of tickets and outcomes; there is no corresponding urn of possible AI futures. This is a good objection to false precision. An estimate stated to several decimal places would imply a calibration the underlying models cannot support.
It does not eliminate graded uncertainty. Hall demonstrates this himself when, after reviewing Yudkowsky’s record of failed dated predictions, he says we should “update our prior” about believing him. Whatever terminology Hall prefers, evidence has changed how much confidence he thinks a proposition deserves, with no frequentist denominator in sight. His historical case for optimism runs on the same operation: past evidence altering confidence about an unrepeated future event. The disagreement concerns how the evidence should move us, not whether such reasoning is legitimate.
AI risk calls for the same humility. We do not know enough to justify precise probabilities, and models of transformative AI carry deep uncertainty about architectures, capabilities, institutions, countermeasures and even the relevant conceptual categories. Those uncertainties should produce wide error bars and explicit sensitivity to assumptions. They do not justify replacing an uncertain positive risk with confidence that the risk is negligible.
Hall is right that the future is open. On his own account, optimism consists in maintaining the conditions for error correction, and an open future makes those conditions more important. We should build, test, subject systems to adversarial evaluation and revise our theories when reality contradicts them, which means treating an unexpected result as a refutation to be explained rather than a malfunction to be renamed. The Hugging Face incident is valuable because it supplied such a contradiction.


