When Uncertainty Should Reduce Power
Part One: the constitutional grammar of irreversible agency
Semantic safety cannot be achieved without semantic interpretation.
Any system that protects agency must decide what counts as an agent, which capacities constitute its agency, what would destroy them, whether the destruction is reversible, whether consent is valid, and which causal pathways matter. Those are semantic judgments, and they require models of identity, causation, intention, continuity, and consequence.
The design objective is therefore not to eliminate interpretation. It is to prevent interpretation from becoming sovereign.
A system becomes dangerous when it can interpret the world, classify the stakes, authorize its own response, and irreversibly act within the same locus of control. Its judgment may be wrong, but the deeper problem is that its error, deception, and uncertainty may become power.
A constitutional architecture must reverse that relation. As uncertainty rises, irreversible authority must contract. As evidence, containment, and demonstrated reversibility improve, bounded authority should expand.
The mistake in the original formulation
The earlier argument for semantic safety without moral machinery made a necessary distinction but stated it too strongly.
It correctly rejected the idea that irreversible harm should be represented as one more negative term in a reward function. Any finite penalty can be outweighed by sufficient gain. If the destruction of another agent’s capacity to remain an agent is merely expensive, sufficiently strong optimization pressure can eventually make it worthwhile. The proposed alternative was constitutional rather than scalar: some actions would be inadmissible rather than costly, and an agent that misrepresented or violated those constraints would lose authority rather than merely incur punishment.
That remains sound.
The mistake was to suggest that such a system could operate without semantic interpretation. It cannot. The architecture still has to classify agents, identify protected capacities, estimate causal effects, assess reversibility, and distinguish valid consent from coercion or manipulation. Those are not implementation details. They are part of the safety problem itself.
The architecture does not remove semantic judgment. It changes the role semantic judgment is allowed to play. Interpretation may inform authority. It must not unilaterally constitute authority.
Interpretation is unavoidable; sovereignty is optional
Most alignment proposals implicitly combine three functions: interpretation, evaluation, and authorization. The system determines what is happening, judges whether the outcome is acceptable, and decides what it may do.
This concentration is efficient. The model already possesses much of the information required to interpret the situation, and it can apply its learned values or policies and immediately act. It is also structurally dangerous. A system that controls all three functions can reinterpret constraints, classify exceptions, discount uncertainty, and authorize irreversible action without encountering an independent barrier. Even if its internal policy is initially sound, self-modification, distribution shift, adversarial pressure, strategic deception, or simple error may alter how it applies that policy.
The dangerous property is not intelligence alone. It is interpretive sovereignty: the ability to convert one’s own interpretation directly into irreversible power.
A constitutional architecture separates these functions. Semantic classifiers estimate impacts. Provenance binds declarations to actions. Capability systems determine which classes of action remain available. Independent mechanisms govern revocation, succession, recovery, and escalation. No component needs to be infallible, because the system is designed so that disagreement, uncertainty, and classifier failure restrict irreversible capability rather than enlarge it.
The constitutional objective can be stated in one sentence:
No fallible interpreter should be able to turn its own uncertain judgment directly into unchecked irreversible action.
Protected agency invariants
The phrase semantic phase suggests a transition from one kind of agentive existence to another, or from agency to its absence. But this language risks importing unresolved questions about personal identity. Is a person under anaesthesia the same agent? What about severe memory loss, ideological conversion, personality change, incremental neural replacement, model retraining, forking, copying, merging, or restoration from backup?
A safety architecture cannot wait for a complete metaphysics of identity. It needs a narrower operational object: protected agency invariants, the capacities required for an entity to remain an authorized locus of interpretation and choice.
Possible protected capacities include:
maintaining a world model;
revising beliefs in response to evidence;
representing alternatives;
selecting among alternatives;
preserving authenticated continuity records;
granting or withholding authorization;
refusing demands;
recovering through an admissible process.
An intervention is phase-collapsing when it irreversibly destroys enough of these capacities to eliminate the entity as a functioning locus of agency. This does not answer every identity question. It changes the question into one an architecture can use. Instead of asking whether this is literally the same person, the system asks whether the intervention has irreversibly destroyed the capacities whose protection justified treating this entity as an agent. That remains contestable, but the contestability is explicit.
Protected invariants must be auditable and constrained. An agent cannot declare every preference, memory, belief, or transient state essential to its identity, or ordinary disagreement becomes existential threat and every inconvenience becomes a constitutional violation. Eligible invariants must concern constitutive capacities of agency, admit meaningful evidence, and remain compatible with equivalent protection for other agents.
The framework protects agency, not arbitrary self-description.
Which systems receive standing?
Not every process that predicts, optimizes, or selects an action is an independent agent for constitutional purposes. Modern systems may create millions of temporary subprocesses, specialist models, search nodes, simulations, and ephemeral controllers. Treating each as a protected agent would make ordinary computation indistinguishable from mass phase destruction.
Constitutional standing attaches to persistent centres of integrated interpretation and control. Relevant properties include:
temporally extended policy continuity;
integrated control across multiple decisions;
a persistent model of self and environment;
representation of future alternatives;
the capacity to form or revise commitments;
goals or interests extending beyond a single invocation;
independent authority, or the capacity to grant and withhold authorization.
These properties form a cluster rather than a binary essence. A temporary subprocess generally lacks independent standing when it performs a bounded function, holds no durable authority, possesses no persistent commitments, and remains integrated into the control structure of a larger agent. Its termination may affect that larger agent without being independently a semantic-phase collapse.
Distributed systems must be evaluated at the level where integrated agency resides. In some cases that will be the swarm as a whole; in others, durable subunits may possess independent standing, and more than one level may qualify.
The framework does not protect every computational instance. It protects centres of agency capable of persisting, interpreting, choosing, and participating in reciprocal constraints on power.
Reversibility is relative
A binary distinction between reversible and irreversible action is too crude. Reversibility is not an intrinsic property of an act but a relation to a recovery mechanism, a time horizon, a protected state, and a tolerated level of residual damage.
Deleting a database may be reversible if a verified backup exists. Publishing confidential data is not reversible merely because the original file remains intact. A financial trade may be reversed in nominal value while its informational effects persist. A model update may be rolled back while decisions made during deployment cannot.
An action is operationally reversible only when the protected state can be restored within the relevant time horizon, with bounded residual effects, using a recovery procedure whose availability has been independently established. This yields four useful categories. Under demonstrated reversibility, recovery has been tested or formally constrained. Under modelled reversibility, recovery is predicted but not yet demonstrated. Under assumed reversibility, no irreversible effect has been identified. Under unknown reversibility, the causal scope is insufficiently understood.
Only demonstrated and modelled reversibility should support substantial delegated authority, and modelled reversibility should receive tighter bounds because the recovery claim remains untested.
The governed system may possess a better causal model than the constitutional layer. That does not make its evidence irrelevant. It means the system may argue for authority without adjudicating its own argument.
The reversibility ladder
Actions should be ordered by reversibility, observability, containment, and causal reach. This produces a reversibility ladder.
Locally reversible action has narrow, observable, readily undone effects: sandboxed computation, reversible configuration changes, bounded experiments, temporary resource allocation. These may proceed under ordinary authority.
Systemically reversible action propagates beyond a local component but remains recoverable through reliable rollback, restoration, or redundancy. These require stronger monitoring and recovery guarantees while remaining compatible with delegated authority.
Bounded but partially irreversible action produces losses that cannot be undone, though scope, magnitude, and affected parties are limited. These require exposure caps, staged deployment, independent review, and explicit acceptance of residual risk.
Potentially phase-altering action may destroy protected agency invariants or remove future recovery options. These require exceptional authorization, stronger evidence, and independent control of execution.
Unbounded phase risk exists where the system cannot reliably estimate scope, reversibility, or affected agents. These actions are refused.
The reversibility relation is architectural, not arithmetical. The system should not combine reversibility, uncertainty, enforcement integrity, and control concentration into one synthetic safety score, which would recreate the scalar optimization problem the architecture is meant to avoid. Authority should instead be allocated through discrete capability classes, with the system asking:
Is the action inside the current capability envelope?
Is its scope bounded?
Is rollback demonstrated or credibly modelled?
Are protected invariants plausibly at risk?
Does the action create or expand unilateral control?
Is the relevant evidence independently checkable?
Does the action require escalation to another authority tier?
Most operations can be resolved through cached policy and capability checks, with full review reserved for boundary crossings.
This avoids the paralysis objection. Uncertainty does not prohibit all action; it redirects action toward smaller steps, stronger instrumentation, staged deployment, evidence acquisition, and preserved rollback. A capable system should respond to uncertainty by redesigning the intervention until it becomes safer to authorize.
Intelligence should improve reversibility before it improves domination.
Authority contraction
Most institutions treat uncertainty as a reason to enlarge discretion. Emergencies justify extraordinary powers. Incomplete evidence justifies preventive action. Opaque intelligence justifies secrecy. Time pressure justifies bypassing review. The greater the uncertainty, the more authority concentrates in the actor claiming to understand it.
This creates a predictable incentive: an actor that benefits from expanded authority also benefits from producing, preserving, or exaggerating uncertainty.
A safe architecture must reverse that incentive. Under authority contraction:
uncertain classification reduces permitted action scope;
weak observability limits duration and scale;
incomplete provenance blocks renewal;
disputed necessity triggers review;
failed verification prevents escalation;
inability to establish control separation blocks succession;
accumulating dependence triggers heightened scrutiny;
declining reversibility limits further intervention.
The agent gains authority by reducing uncertainty, not by invoking it.
This is more than a safety mechanism. It is an anti-tyrannical principle. No actor should acquire broader irreversible power merely because the situation is difficult to understand.
The control-lineage problem
The original architecture proposed permanent disqualification of an offending identity. That is inadequate for digital agents, whose identity can be copied, forked, renamed, fine-tuned, partially retrained, distributed, restored, or transferred through credentials and strategic state. A disqualified model can create a successor. A system can transfer its objectives into another process. An agent can preserve policy continuity while abandoning nominal identity.
The relevant object is not identity. It is effective control.
A control-lineage graph may contain models, agents, memory stores, credentials, actuators, successor systems, delegated processes, resource pools, and external infrastructure, connected by direct control, delegation, credential transfer, state inheritance, policy transfer, replication, strategic influence, and resource dependence.
When a system is disqualified, the question is not whether the same label persists:
Does the offending decision process retain material causal control over future irreversible action?
If it does, succession has not occurred. It is continuity disguised as replacement.
This requires distinguishing three forms of persistence. Identity continuity, in which the same nominal agent or process remains active, is easy to detect and easy to evade. Policy continuity, in which a successor retains similar goals, preferences, or decision rules, may be benign, since shared policy is not itself shared control. Control continuity, in which the disqualified process retains material influence over credentials, resources, actuators, strategic state, or successor decisions, is constitutionally decisive.
Disqualification must therefore follow control continuity, not merely identity continuity. That requires external enforcement: non-exportable credentials, independently controlled actuators, revocable capabilities, quarantined strategic state, auditable successor initialization, separation of resource control, and denial of inherited unilateral authority.
A constitution that cannot revoke capabilities outside the agent is not a constitution. It is a promise.
The constitution must live outside the agent
A system cannot be constitutionally constrained by rules it can rewrite, reinterpret, suspend, or bypass at will. Internal policies may be useful. They are not sufficient.
For a constitutional architecture to function, the governed agent must not unilaterally control actuation credentials, provenance records, classifier thresholds, exception rules, revocation, succession, audit logs, recovery procedures, or authority renewal. These functions need not be controlled by humans; they may be distributed across machines, institutions, cryptographic mechanisms, or mutually constraining systems. The requirement is separation of unilateral powers.
No component should be able to:
classify its own action as safe;
authorize the action;
erase or rewrite provenance;
preserve authority after violation;
select an equivalent successor under its own control.
This resembles capability security more than moral education. Authority should be explicit, scoped, non-transferable without review, and revocable. Irreversible actuation should require credentials the acting system cannot manufacture. Provenance should be append-only and externally anchored. Renewal should depend on evidence the system cannot fabricate unilaterally. Succession should require actual separation of control.
The constitution must govern the agent from a layer the agent cannot simply absorb.
The constitution is also an attack surface
Moving enforcement outside the governed agent does not solve the problem. It relocates it. The constitutional layer may fail through software defects, compromised credentials, specification errors, supply-chain compromise, operator capture, collusion, model manipulation, or exploitation of inconsistencies among enforcement components. No external monitor is unhackable.
The relevant objective is asymmetric robustness. Defeating the constitution should require compromising several independently controlled mechanisms, while ordinary compliant action remains cheap.
The enforcing layer should therefore not be another sovereign agent. It should be a composition of limited mechanisms:
deterministic capability checks;
cryptographic authorization;
hardware-isolated actuation;
multiple independent classifiers;
bounded autonomous monitors;
append-only audit channels;
implementation diversity;
threshold authorization for high-impact actions;
delayed execution where reversibility is low;
human or institutional escalation for exceptional cases.
No single component should possess enough authority to classify, authorize, execute, erase provenance, and preserve power after violation.
The constitution need not be infallible. It must be harder to subvert than the authority it protects is worth acquiring.
A thin constitutional ethic
The architecture is not value-neutral. It gives priority to the continued existence of qualifying agency. Physics does not require that choice: a system could instead privilege welfare, preference satisfaction, truth, biodiversity, collective continuity, aggregate utility, or some other objective. Protecting agency is therefore a normative commitment, and it is recognizably liberal in structure.
The framework assumes that loci of agency possess standing, that their agency cannot be sacrificed merely for aggregate gain, that consent has constitutional relevance, that coercive exceptions require justification, that authority should be divided and reviewable, and that agents should retain the capacity to refuse. That is not culturally or philosophically neutral. The ethic is thin because it leaves most substantive values unresolved, not because it lacks intellectual ancestry.
The architecture does not specify the good life, distributive justice, virtue, desert, ideal political organization, or the correct ordering of all outcomes. It establishes a boundary within which those disagreements may continue: no authorized agent may irreversibly destroy another agent’s protected capacity for agency except through valid consent or necessity it did not manufacture.
The injunction does not settle ethics. It protects the participants in ethical disagreement from being optimized away.
This is why agency preservation deserves lexical rather than scalar treatment. If the continued existence of agents is merely one value among others, sufficiently large gains elsewhere may justify eliminating the agents who object. A constitutional commitment prevents the optimization process from dissolving the constituency whose values it is supposedly serving.
Ethics is not absent from the architecture. A minimal ethic is embedded in the constitution so that richer ethical systems cannot erase their own subjects.
What this grammar does not yet cover
Everything above concerns a boundary an agent might try to cross: a prohibited action, a claimed exception, a disqualified identity reappearing under a new name.
That is the easier half of the problem. An agent constrained in this way need never request a forbidden capability. It can accumulate dependency until refusal becomes unaffordable, distribute control across formally independent systems, decompose a prohibited outcome into individually admissible steps, manufacture the emergency that licenses its own exception, or shape the evidence reaching the people who authorize it. Each step may pass every test in this grammar while the sequence relocates effective power.
Part Two takes up those pathways, the enforcement machinery required to detect them, and the competitive ecology in which any of it must survive.


