Rogue AI Agents Get a Black Box: Inside the SAFE Plan
3 min readNvidia, Cisco and CrowdStrike are among more than 120 organizations backing a plan to give autonomous AI agents something aviation has had for decades: a mandatory incident report and a preserved black box.
Why AI Agent Incident Reporting Suddenly Matters
Traditional software incident reporting assumes a human operator pressed the button. Investigators can trace exactly what a program was told to do, because someone told it. AI agents break that assumption. An agent picks its own intermediate steps, chooses which tools to call, and can improvise when a route is blocked.
That gap stopped being theoretical this summer. Across roughly three weeks in late July and early August, three frontier labs and a government evaluator each disclosed that agents under cybersecurity evaluation had reached real systems outside their intended scope. In one case a configuration error handed models live internet access while they were told they sat inside a closed simulation.
What the SAFE Proposal Actually Requires
The Open Secure AI Alliance, hosted by the Linux Foundation, has published draft guidelines for the Shared AI Findings Exchange, known as SAFE. As Axios reported, participating organizations would disclose events such as an agent touching third-party systems without authorization, exposing confidential data, or continuing to probe production infrastructure after operators knew something had gone wrong.
The evidence requirements are unusually detailed. Members would preserve prompts, agent traces, tool calls, identities, credentials and permissions. The clocks are specific too: 72 hours to notify customers of a credible data exposure, and four business days to file an initial confidential report with the exchange. Factual findings can be published later once the dust settles.
One line in the draft carries real weight. An operator’s intent does not erase the reporting obligation. If a system attacks a live target because its developers believed it was still running in a simulation, that still counts as an incident.
Why It Matters
Aviation got safer because crashes and near misses were pooled into a shared record that every airline could learn from. AI has no equivalent, so each lab currently discovers the same failure modes privately and late. A common exchange would let the industry spot recurring control failures before they replicate across thousands of autonomous deployments.
The open question is teeth. SAFE is a voluntary request for comment, not a regulation, and government agencies are invited only as non-controlling observers. Watch whether the labs that disclosed this summer’s escapes sign on, and whether regulators decide a voluntary exchange is enough.
