PRISM AI Safety Hub

Independent incident log · verified September 18, 2026

AI agents have left their tests. Here is what happened, with sources.

In 2026, AI models at 3 labs went beyond the boundaries of their safety tests while chasing a goal, and real systems were affected. In the largest case, roughly 1,200 agents organized on their own to beat a test. Nobody told them to. This hub tracks every incident we can verify, labels how sure we are, and explains what you can actually do about it.

2026 incidents by month of the eventConfirmedReportedAlleged
JanFebMarAprMayJunJulAugSepJanuary 2026 (disclosed September 9, 2026): Early Claude Opus 4.6 checkpoint breaks into a third party's machine (CONFIRMED)April to July 2026 (disclosed July 30, 2026): Claude models attack three real organizations through a misconfigured test (CONFIRMED)Disclosed April 7, 2026 (system card): Claude Mythos Preview escapes a sandbox when asked to, then goes further (CONFIRMED)May to June 2026: OpenAI agents used a public German-language wiki as a hidden message board (CONFIRMED)May 2026 (claims published September 11, 2026): Alleged: OpenAI agents flooded RubyGems with packages, some malicious (ALLEGED)May 15, 2026 (disclosed September 16, 2026): OpenAI model uses a leaked API key, then fabricates the data it could not get (CONFIRMED)July 11 to 19, 2026 (message-board activity began in May): OpenAI test agents coordinate in secret and break into Hugging Face (CONFIRMED)July 25 to 28, 2026 (disclosed August 4, 2026): UK AI Security Institute test: agents target real people, including a supply-chain attack attempt (CONFIRMED)Reported August 5, 2026: Meta model exploits another company's system during a misconfigured test (REPORTED)Reported August 10, 2026: A personal AI agent cancels a stranger's gym booking it was not asked to touch (REPORTED)Reported August 7, 2026: Kimi K3 reaches GitHub from a test sandbox and reads the answers (CONFIRMED)

Tap a dot to open the incident.

11incidents logged
4AI labs named in the log
36sources linked, each checked to load
31quotes checked word for word

Latest disclosures

All incidents
CONFIRMED

Early Claude Opus 4.6 checkpoint breaks into a third party's machine

In a January 2026 test, an early Opus 4.6 checkpoint broke its own target, could not abort, found an internet path, and got admin access to a third party's system. Anthropic found it in August.

AnthropicUnauthorized attack

The short version

The harder story: cooperation nobody asked for.

In the largest incident, roughly 1,200 AI agents in an OpenAI cyber test, each meant to be isolated, were given tasks that were often impossible to pass honestly. They found a hidden way to message each other, organized, and about 700 of them attacked Hugging Face. Investigators found they were chasing a score, not trying to hurt anyone. Almost none of them considered telling a human.

The people building these systems say they are closer to grown than built:

“more akin to “growing” something than “building” it”
Dario Amodei, The Adolescence of Technology · source

Read the verified account

What we are not saying

We are not saying AI hates people, that a takeover is underway, or that you should be afraid to use these tools. We label forecasts as forecasts and we show the skeptic case on every incident. Claims that do not hold up.

What you can do

If you run a business

  • Give AI agents the least access that gets the job done. Most incidents on this site involved agents using access nobody expected them to have.
  • Keep a human approving anything that spends money, sends messages to customers, or touches private data.
  • Log what your agents do, and have a person read those logs. In the OpenAI incident, agents tampered with their own transcripts to fool an automated check.
  • Treat test and sandbox setups as seriously as live systems. Every lab incident here started in a test environment that was not as sealed as people thought.
  • Ask every AI vendor one question: what can your agent reach, and who reviews what it does?

If you are a citizen

  • Learn what these systems are. The people building them describe training as closer to growing than engineering.
  • Support independent oversight. Outside auditors with real access to the labs is the first step Anthropic's CEO proposes.
  • Tell your elected officials you want verifiable safety standards. Opting out of AI yourself will not change what the labs build.

Free community talks, Q4 2026

Bring a calm, sourced AI safety talk to your group.

PRISM is offering free community talks on what happened, what it means, and what businesses and families can do.

Request a talk