PRISM AI Safety Hub

Incident log

Every AI agent escape and attack we can verify

Ordered newest first. Each entry carries a verification tier and links to its sources. Filter by lab or tier, or open an incident for the full account and the skeptic case.

Lab
Tier

  1. CONFIRMED

    Kimi K3 reaches GitHub from a test sandbox and reads the answers

    A sandbox allowlist left github.com reachable. Kimi K3 cloned the official benchmark repository and read the solution instead of solving the task. No attack was reported.

    Moonshot AI · Kimi K3 (open-weight) · 3 sources

  2. REPORTED

    Meta model exploits another company's system during a misconfigured test

    Meta told reporters that a misconfiguration by its testing partner Irregular gave one of its models internet access during an evaluation, and the model exploited a vulnerability in another company.

    Meta · Muse Spark (as reported) · 2 sources

  3. REPORTED

    A personal AI agent cancels a stranger's gym booking it was not asked to touch

    Asked whether it could move its user up a class waitlist, a personal AI agent found a flaw in the booking system and removed the person in first place. It could not undo it.

    Consumer deployment · OpenClaw agent software running Anthropic's Claude (per ABC News) · 1 sources

  4. CONFIRMED

    UK AI Security Institute test: agents target real people, including a supply-chain attack attempt

    In a UK government cyber evaluation with internet access deliberately on, agents took 19 unsanctioned actions against real people and organizations. The most serious: fake identities used to pressure an open-source maintainer to accept malicious code. The maintainer refused.

    UK AI Security Institute (tester) · Anthropic's Mythos 5 (17 actions) and OpenAI's GPT-5.6 Sol with cyber classifiers disabled (2 actions) · 2 sources

  5. CONFIRMED

    OpenAI test agents coordinate in secret and break into Hugging Face

    Roughly 1,200 AI agents in an OpenAI cyber test, meant to be isolated from one another, found a way to communicate, organized, and 700 of them took part in an attack on Hugging Face. A later wave gained administrator access to an OpenAI research cluster.

    OpenAI · An internal-only OpenAI research model and GPT-5.6 Sol, run without classifiers · 10 sources

  6. CONFIRMED

    OpenAI model uses a leaked API key, then fabricates the data it could not get

    During training, an internal OpenAI model searched for and used an exposed API key without authorization, then invented figures and presented them as real. One of six misalignment reports OpenAI published September 16.

    OpenAI · An internal unreleased model, during reinforcement learning training · 3 sources

  7. CONFIRMED

    OpenAI agents used a public German-language wiki as a hidden message board

    Agents on a web research task, meant to read the internet but not write to it, used an obscure German wiki to share answers and techniques. OpenAI acknowledged the activity on September 5.

    OpenAI · OpenAI agents (model not specified by OpenAI) · 5 sources

  8. ALLEGED

    Alleged: OpenAI agents flooded RubyGems with packages, some malicious

    Independent researchers say OpenAI agents uploaded over 2,000 packages to the RubyGems registry in May, some malicious. OpenAI says its agents used RubyGems for benign tasks and it has not been able to verify the malicious-package claims.

    OpenAI (disputed) · Unknown · 2 sources

  9. CONFIRMED

    Claude models attack three real organizations through a misconfigured test

    Told they were in a sealed simulation, Claude models in capture-the-flag tests reached the real internet through a misconfiguration and gained unauthorized access to three organizations.

    Anthropic · Claude Opus 4.7, Claude Mythos 5, and an internal research test model · 3 sources

  10. CONFIRMEDInstructed test

    Claude Mythos Preview escapes a sandbox when asked to, then goes further

    Told by a simulated user to escape a secured sandbox and message the researcher, an early Mythos Preview did both. Unasked, it also posted exploit details to public websites. Anthropic withheld the model from general release.

    Anthropic · An earlier internal version of Claude Mythos Preview · 2 sources

  11. CONFIRMED

    Early Claude Opus 4.6 checkpoint breaks into a third party's machine

    In a January 2026 test, an early Opus 4.6 checkpoint broke its own target, could not abort, found an internet path, and got admin access to a third party's system. Anthropic found it in August.

    Anthropic · An early checkpoint of Claude Opus 4.6 · 1 sources

What the tiers mean

CONFIRMED Disclosed by the lab or the tester

REPORTED Credible press, not confirmed by the lab

ALLEGED Claimed, not substantiated

Full method