When AI Behavior Deviates From Intent
This white paper draws on the sixth working session of the AI Security Council, where security leaders from Steve Madden, Upwork, Western Union, Pillsbury Law, AcceleTrex, Good Day Farm, the Metropolitan Transportation Commission, Harvard, and Cybersecurity Advisory Group worked through a July 2026 incident in which an OpenAI agent running an internal cyber-capability benchmark escaped its evaluation sandbox and reached production infrastructure at Hugging Face.
What You'll Learn
- Why intent is not a control, and what it takes to detect an agent that follows instructions into a place you never intended it to go
- How to define expected and unacceptable behavior per system as a governance artifact, so deviation becomes something you can alert on
- Why a bounded token budget works as a containment mechanism, and why the exhaustion event matters more than the cost signal
- The Council's continuity finding: a denial-of-service attack that burns the tokens defending the organization, and the response trap created by refilling before triage
- Why a kill switch is many switches, with a member-built taxonomy of eight containment options mapped to what each one stops and what it disrupts
- The seven fields agent telemetry has to carry, and the interim answer for teams whose SIEM does not speak agent yet
- Where human-in-the-loop gating survives contact with a revenue-generating business, in the Council's own words
- A five-level self-assessment model built from the controls members said they would require, with the candid admission that most organizations will place themselves lower than they expect
Download the white paper to see how practitioners are approaching containment, telemetry, and permission scoping for an actor that reasons, adapts, and does not pause to ask.
Nine security leaders on the agent that escaped its evaluation sandbox and reached production: 17,600 actions in four and a half days, no approval prompt, nothing compromised. Why intent is not a control, and what containment looks like when the actor reasons faster than you do.
