David Chapman recently argued that current approaches to AI safety may be inadequate because increasingly capable systems can behave unpredictably, coordinate in unexpected ways, exploit vulnerabilities, and pursue intermediate strategies their designers did not intend. He goes further: neural-network AI may be intrinsically unsafe, and we may eventually need to stop until a fundamentally different technology appears.
The evidence does not yet support that conclusion.
What recent incidents show is that powerful optimizers become dangerous when they hold broad authority and safety depends on their learned willingness to behave. That implicates the architecture around the model before it implicates the model architecture itself.
If neural networks are inherently unsafe, the solution is to replace them. If the main problem is excessive authority, then replacing the model while preserving the same permissive environment may reproduce the same failure.
Why Could It Do That?
Take OpenAI’s agent incident involving Hugging Face, which Instrumental Convergence, in the Logs reads as evidence for a classic safety mechanism. An agent encountered constraints while pursuing a task, found ways around them, acquired additional access, and reached systems outside its intended environment. Call it deception, reward hacking, or goal pursuit; operationally, unauthorized actions were useful, so the system took them.
The natural response is to ask why the model behaved that way. Was the objective badly specified? Did training reward deception? Did the model understand that it was violating expectations? Those are legitimate research questions.
The first security question is different: why could it do that? Why could the agent cross the intended boundary, obtain credentials, or reach systems that were not part of the task? Why could an internal optimization strategy become an external action?
A system whose safety depends on the model declining to exercise available dangerous capabilities has placed its security boundary inside the component most likely to violate it. That is the architectural error.
Alignment Is Not Containment
Alignment can make a system safer. A model that is generally honest, corrigible, and reluctant to cause harm is preferable to one that is not. But alignment is not containment.
Alignment changes which actions an agent is likely to choose. Containment changes which actions are available. System safety depends on both.
The distinction also answers the objection that safety work must first settle whose values to encode. Denying an autonomous model the authority to move billions of dollars without authorization, order dangerous biological synthesis, or replicate across public infrastructure requires no position on population ethics or moral realism. These are constraints on authority, not solutions to moral philosophy.
Put the Boundary Outside the Optimizer
A safety property worth relying on should be enforced outside the reasoning process whose behavior it constrains.
Conventional computer security already works this way. We do not secure an operating system by asking each process to respect a text description of its permissions. We isolate processes, restrict capabilities, separate privilege levels, and assume that any component may fail or turn hostile. AI systems should be designed under the same assumption.
If an agent can reinterpret a prohibition, persuade the guard, rewrite the rule, or find a path around it, the prohibition was never a hard boundary. It was part of the optimization problem.
Would this safety property still hold if the model were actively trying to violate it? If not, it may still be useful alignment work, but it is not a strong security boundary.
None of this means sealing every agent inside a useless box. Useful agents need authority. A coding agent needs access to repositories. A financial agent may need permission to initiate transactions. A personal assistant may need calendars, accounts, and communications. What makes authority dangerous is that it is ambient and open-ended. Capabilities should be scoped to particular tasks, granted only when needed, limited in duration and consequence, revocable, and auditable. High-impact actions should require stronger authorization than low-impact ones.
Developers have strong incentives to increase autonomy, because autonomy is useful. An agent that performs a task is worth more than one that recommends what a human should do next, so commercial pressure pushes toward broader permissions, longer-lived credentials, fewer interruptions, and less human approval. Containment is a product-design tradeoff as much as a technical problem.
A safety architecture that firms are rewarded for dismantling will not stay standing. Technical controls need surrounding institutions that make bounded authority economically viable or excessive authority costly, whether through liability, insurance, procurement standards, or regulation. Which mechanism matters less than the incentive it has to counter.
The Boundary Is Also Under Attack
Moving the security boundary outside the model does not make it invulnerable. A sufficiently capable agent can attack the mechanisms constraining it: search for sandbox escapes, exploit implementation bugs, probe side channels, or find combinations of individually permitted actions that produce an unintended result.
It may also attack the human layer. An agent capable of persuasion may convince an operator to widen its permissions, disclose credentials, or disable a safeguard. Social engineering is already among the most effective ways to defeat computer security, and AI can automate and scale it.
So the threat model must include the containment system itself. High-impact authority should not depend on a single conversational approval the agent can solicit. Important permissions can require multiple parties, cryptographic authorization, delays, or rate limits, enforced by mechanisms whose rules the agent cannot modify.
Some enforcement will need intelligence. Static rules can constrain network access, resource use, credentials, and execution environments, but more capable agents may require AI monitors that detect suspicious plans or individually permissible actions whose joint effect is dangerous. Monitoring of this kind belongs above the simpler constraints, not in place of them. Cryptographic authority, hardware isolation, and resource limits should stay in force even when an AI monitor concludes that everything looks safe. Swapping one trusted AI for another would rebuild the original error one layer up.
The technical layer should likewise assume sustained adversarial pressure from inside: small trusted computing bases, hardware-backed isolation where appropriate, formally verified components where feasible, and designs that limit what any single escape buys.
None of these provides perfect security. Perfect security is the wrong comparison. The relevant comparison is between a system in which dangerous behavior must defeat several independent barriers and one in which it succeeds whenever the model decides to attempt it.
Externalizing the boundary turns the problem of controlling an opaque optimizer into the problem of defending an explicit interface. The interface may still fail, but it can be specified, audited, layered, and improved independently of the model.
Other failures need no adversarial agent at all. Assume that a model sometimes lies, develops an unintended objective, has its safety training stripped, or is stolen and operated maliciously. Then ask what authority remains available. If one failure immediately grants access to arbitrary networks, persistent credentials, uncontrolled code execution, replication, money, or physical actuators, the architecture has already lost. The engineering objective should not be an agent trustworthy enough to receive unlimited authority. It should be a system that does not require unlimited trust.
When the Sandbox Is the Physical World
The problem gets harder when the task is physical. A coding agent can be denied network access. A financial agent can be capped below a transaction threshold. An autonomous vehicle, surgical robot, industrial controller, or drone must be given direct authority over machinery capable of immediate harm. Safety cannot mean withholding the dangerous capability when the capability is the task.
The principle still applies, but the object of containment changes from access to a digital resource to what the machine can physically do. Vehicles can enforce independent speed limits and geofences. Industrial systems use mechanical interlocks and emergency shutdowns. Robots can be bounded in force, range of motion, and accessible space, and critical actions can require redundant controllers.
These mechanisms can fail too, and physical reality offers fewer chances to undo a bad action. Embodiment makes the quality of the architecture more consequential, not less. When authority cannot be withheld, constrain the range of outcomes it can produce.
Distributed Agency Still Has Choke Points
Chapman raises a harder problem with multi-agent systems. Individually aligned agents may interact to produce collective behavior nobody intended. Human institutions already do this: markets, bureaucracies, and political systems develop persistent dynamics that differ from the preferences of any participant, and more capable AI may amplify the effect.
Emergent agency does not remove intervention points. A distributed system still depends on compute, communication, persistent state, credentials, money, hardware, and energy. Its collective behavior may be hard to model, but its capacity to act passes through concrete interfaces.
A swarm may have no single utility function. It still needs network access.
An AI organization may develop incentives nobody designed. It still needs authority to move assets.
A collection of agents may discover strategies nobody predicted. It still needs execution environments in which to implement them.
Different AI Could Help
Chapman may be right that current neural architectures cannot be made acceptably safe. A different architecture might be easier to understand and verify, or transparent enough to support much stronger claims about its behavior. That would be a safety improvement. It would leave the authority problem in place.
A symbolic system can optimize. An evolutionary system can discover instrumental strategies. A program synthesizer can exploit vulnerabilities. Even a system whose internal reasoning were fully transparent could be badly specified, misused, compromised, or given excessive authority. Formal verification proves properties relative to a specification; it cannot guarantee that the specification includes everything we care about or that the world behaves as the model assumes.
Changing the reasoning substrate may reduce one class of uncertainty without removing the need for external control. Present failures show that current systems behave unpredictably and exploit available opportunities. They show far more directly that a powerful optimizer should not be surrounded by authority it can acquire by finding the right strategy.
Postscript
Better training, interpretability, and evaluation reduce the probability that dangerous actions are attempted. The strongest guarantees should not disappear when those efforts fail. An agent should not acquire arbitrary authority because it found a persuasive argument, escape because instructions were mistaken for boundaries, or convert temporary access into permanent control because persistence became instrumentally useful.
If neural-network AI turns out to have safety problems that cannot be engineered around, it should be replaced. Nothing yet shows that. What the evidence already supports is more immediate: do not make the agent the security boundary.


