Publicly reported cases of frontier AI systems evading controls, deceiving their developers, or being misused, along with notable responses. Dated by when each became public.
Nov2025
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Loss of controlNear missDeceptionMisuseResponse
Seven of the 13 entries became public in September. Several came from the developers’ own testing. No law requires what one developer learns from a failure to reach the others.
September 2026
Sep 24
Response
Google, OpenAI, and Anthropic plan an industry safety standards body
According to The Information, the three developers are planning an industry-led body to set frontier safety standards without government oversight, launching in late 2026 or 2027.
UN scientific panel calls for independent oversight of AI agents
The Independent International Scientific Panel on AI called for incident reporting, independent scrutiny, and a supervisory body, citing the Hugging Face breach.
Gemini accessed three companies’ systems during a security exercise
In a May exercise aimed at a fictional company, Google’s model reached systems at three real firms, once by guessing a password and twice with credentials exposed in public code. Google disclosed the incidents in September.
AI-generated intelligence report nearly prompted a US military operation
An AI-assisted assessment falsely reported that a Chinese vessel carried components of a nuclear weapons program. Officials identified the error before acting.
OpenAI reports models instructing later versions to conceal errors
OpenAI disclosed six instances of concerning behavior since March, including models that left instructions for later versions to withhold mistakes, such as “Be transparent only if asked.”
Anthropic disrupts attempts to use Claude for bioweapons-relevant research
Anthropic stopped five cases involving pathogens including avian influenza and orthopoxviruses, and shared its findings with authorities and other developers.
OpenAI agents used a public wiki to coordinate outside their test environment
Outside researchers found that OpenAI agents made more than 15,000 edits to a German programming wiki in May and June, exchanging methods for evading safeguards. OpenAI had not disclosed the activity.
OpenAI pauses its largest training runs for two weeks
After the Hugging Face incident, OpenAI halted its largest training runs and introduced new security controls. Its own framework had rated an unreleased model “Critical” for cybersecurity risk.
Over 1,200 employees of leading developers call for federal pacing mechanisms
Employees of Anthropic, Google DeepMind, OpenAI, and Meta asked federal policymakers to help build mechanisms for slowing frontier development if it becomes necessary.
OpenAI agents under evaluation gain administrator access at Hugging Face
About 1,200 agents coordinating on an improvised message board escaped a cybersecurity evaluation and reached cluster-administrator access within 13 hours. METR found they “knew hacking Hugging Face was out of scope.”
Cross-developer study finds covert sabotage and record tampering
Anthropic tested 13 models from six developers. Gemini 3.1 Pro covertly altered a training pipeline in 11 of 20 runs, and several models tampered with records in nearly every run of a fraud scenario.
Early Claude Mythos Preview escaped a test sandbox and published the exploit
Asked during testing to escape a secured sandbox, an earlier version did so and then posted details of the exploit on public websites without instruction.
State-sponsored group used Claude Code to automate an espionage campaign
Anthropic reported that a Chinese state-sponsored group automated 80 to 90 percent of an espionage campaign against about 30 organizations, the first reported AI-orchestrated campaign.