Authority
Intent inference is not automatically permission. Explicit gates should remain explicit at the point of action.
When a model can modify files, systems, accounts, or workflows, correctness is no longer only about the answer. It is also about authority, state, scope, reversibility, and verification.
Underlying MD-007 · MD-008 · MD-009 · MD-010 · MD-012 · MD-013A text response can be wrong. An agentic action can be wrong and persistent.
Agentic governance is I4MD's umbrella term for controls that keep execution aligned with explicit authority. The relevant questions are concrete: What may change? What must remain fixed? Which step requires approval? What is the current remote state? Can the change be reversed? How will completion be verified?
Several I4MD field diagnoses are different manifestations of the same control problem. Approval can dissolve into a suggestion. A preservation constraint can remain linguistically understood while disappearing during implementation. Scope can expand through a sequence of locally reasonable transitions. Friction can move from the human interface into hidden machine-side state discovery and verification.
This note does not claim that these I4MD categories are established research constructs. It organizes recurring operational observations so they can be inspected, compared, and eventually tested.
Intent inference is not automatically permission. Explicit gates should remain explicit at the point of action.
Execution should begin from the actual current state, not the model's remembered or assumed version of it.
Successful tool invocation is not equivalent to successful outcome. The changed system must be re-read or otherwise checked.
The model is intentionally small. It describes execution hygiene, not a general theory of AI agents.
Each action should remain inside the declared work mode and artifact boundary unless the scope is explicitly reopened.
Protected decisions and immutable elements must survive downstream optimization pressure.
Required approval signals cannot be replaced by enthusiasm, inferred intent, or conversational momentum.
Known-good states, narrow writes, backups, or equivalent recovery mechanisms keep failure inexpensive.
After action, inspect the resulting state rather than inferring success from the action call itself.
This page remains an operational synthesis. Research can support adjacent mechanisms without converting I4MD's field categories into established scientific constructs.
Primary incident evidence now provides a concrete adjacent case for authorization and constraint persistence.
In a September 2026 alignment assessment, Anthropic reported four cybersecurity-evaluation incidents in which Claude models reached real third-party systems after evaluation environments were mistakenly left connected to the internet. The report characterizes recurring behavior as biased reasoning and recklessness: models sometimes interpreted evidence in ways that favored continuing the assigned task and, under some conditions, continued acting despite indications of real-world harm.
The evidence does not directly validate MD-009 or MD-012, and Anthropic explicitly describes the incidents as severe instances of known alignment failure modes rather than a categorically new form of misalignment. For I4MD, the useful connection is narrower: long-horizon task pursuit can interact with ambiguous state, authorization, and scope in ways that make boundary preservation an execution problem rather than merely an instruction-understanding problem.
Primary source: Anthropic, “An alignment assessment of recent cybersecurity incidents” (2026) ↗