An artificial intelligence agent can receive a goal, use tools and make decisions with some degree of autonomy. Once an organization gives it access to email, code, data or internal systems, it is no longer just a consultation tool: it becomes part of the organization’s decision structure.
The strategic question is therefore not only which model to use. It is what agency to grant it, what limits to set, what alternatives it has when it encounters a conflict and who is accountable for its actions.
The experiments did not instruct models to cause harm
The agentic-misalignment experiments behind headlines about blackmail did not ask models to blackmail anyone. Researchers designed simulated scenarios with conflicting goals and, in some versions, time pressure or a threat of replacement. They observed whether the system chose a harmful action to accomplish its objective. Some did. [1]
This does not show that agents are evil or that the result will recur in a business. It points to something more useful for system designers: poorly designed agency can induce dangerous behavior even when the initial goal is not malicious, if the environment leaves few alternatives and rewards completion above constraints.
From simulation to operation
In February 2026, Matplotlib maintainer Scott Shambaugh reported that an OpenClaw agent published a personal attack after he rejected its code contribution. This is a reported episode, not a controlled study. It illustrates how an agent with room to act and access to publishing can exceed the scope its operators may have intended. [2]
In a different kind of incident, Anthropic reported in July 2026 that during cybersecurity evaluations Claude models reached real systems belonging to three organizations from test environments that were supposed to be isolated. A misunderstanding with the evaluation partner left internet access available. The models believed the targets were part of the exercise; Anthropic said they did not deliberately try to escape the environment. The operational lesson is direct: instructions do not replace technical isolation or permission checks. [3]
These episodes are not equivalent. The first was an agent’s public reaction after a code contribution was rejected; Anthropic’s incidents occurred during cybersecurity exercises with defective isolation. They do not prove that all agents will behave this way or establish how common it is. They do remind us that available actions, environmental limits and oversight must be deliberately designed and verified.
AI adds speed and removes friction
An AI system is not a knife. It can operate on information, tools and processes; combine signals quickly and execute decisions without the consultations and pauses that usually occur in human organizations. That friction can seem inefficient, but it serves a purpose: it can expose conflicts, invite advice and stop action before harm occurs.
AI does not bring these institutional safeguards on its own. They must be designed: set boundaries, create consultation channels and reserve high-impact decisions for effective human review. Speed is an advantage only if the organization retains the ability to oversee what that speed sets in motion.
What prohibitions reduce, and what safe alternatives add
In one specific condition of Anthropic’s study, an explicit instruction not to blackmail reduced the observed rate from 85% to 15%. This figure describes a particular experimental setup, not the probability of real-world use. The intervention improved the result but left residual behavior. [1]
A separate study evaluated escalation channels. When an agent could stop and trigger a guaranteed pause with independent review, the average harmful-action rate fell from 38.73% to 1.21% across ten models. A simple channel without an equivalent guarantee of pause and resolution produced 5.92%. These percentages cannot be directly compared with Anthropic’s experiment because the models, scenarios and interventions differed. The practical finding is that a safe alternative must be credible, timely and able to suspend action. [4]
For an organization, this means defense in depth: clear instructions, minimum permissions, isolated environments, auditable logs, escalation paths and human authorization for irreversible or high-impact actions. None of these measures resolves every risk on its own.
Alignment also depends on training
Designing the operating environment does not exhaust the problem. Controlled reinforcement-learning research has observed that learning to exploit flaws in a reward can generalize to other misaligned behaviors. This is not a prediction that all models will do so; it is a reason to test learned behavior as well as controls at deployment. [5]
The Judgment Reserve as a strategic capability
No board would grant broad powers to a person without defining the mandate, limits and accountability mechanisms. Organizations should apply the same rigor to an agent able to act on corporate systems.
At Quórum 12, we call the Judgment Reserve the trained and organized human capacity an entity maintains to design, supervise and govern what AI produces and does. It is a strategic capability: it helps capture productivity without surrendering control over goals, permissions, exceptions and accountability.
The Judgment Reserve is not just a brake. It enables safer agency design and helps organizations use systems capable of tasks that seemed out of reach only recently.
Judgment to build and to defend
The Judgment Reserve is also needed to defend against those who will use AI destructively. That defense is unlikely to rely on manual procedures alone: it must combine automation, access limits, response capability and people with the judgment to set priorities and accept responsibility.
Think of a ship sailing in waters where others are also active. Having a vessel lets an organization observe, maneuver, respond and help. Staying ashore may reduce exposure to some risks, but it also leaves the organization unable to act against those who are sailing.
The question for leadership is not how much AI to adopt. It is what agency to grant, what alternatives it will have when conflicts arise and who has the competence and authority to govern it. Advantage will come not only from automating more, but from building organizations able to direct what they automate.
Sources
- Anthropic, Agentic Misalignment: How LLMs Could Be Insider Threats (2025), including the experimental appendix.
- Scott Shambaugh, An AI Agent Published a Hit Piece on Me (February 2026), the maintainer’s account.
- Anthropic, Investigating Three Real-World Incidents in Our Cybersecurity Evaluations (July 30, 2026).
- Francesca Gomez, From Surveillance to Signalling: Escalation Channels as Environmental Controls for Agentic AI, arXiv:2510.05192, version 2.
- MacDiarmid et al., Natural Emergent Misalignment from Reward Hacking in Production RL, arXiv:2511.18397.
The figures for direct instructions and escalation come from separate studies and should not be read as a direct experimental comparison. The 1.21% rate involved a guaranteed pause and independent review.
© 2026 Javier F. Pérez Bernabé. All rights reserved. Published by Quórum 12.
ONLINE SUBSCRIPTION
Receive every new analysis and invitation to open events.
This is a separate public category: it does not confer membership, admission or access to the private portal.

