PRISM AI Safety Hub

Source packs

Read the originals

The primary sources behind each topic. We link out rather than republish. Where we have built a shared research notebook for a topic, you will find a link to it here.

The OpenAI agent swarm and Hugging Face

The independent investigation, the lab's own account, and key coverage.

  1. Lab disclosureThe Hugging Face incident and the road ahead · OpenAI · August 26, 2026
  2. Lab disclosureOpenAI - Hugging Face Incident Technical Report (PDF) · OpenAI · August 26, 2026
  3. Lab disclosureOpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI · OpenAI · July 21, 2026
  4. Lab disclosurePacing model development in an era of cyber-critical capabilities | OpenAI · OpenAI · August 18, 2026
  5. PrimarySecurity incident disclosure: July 2026 · Hugging Face · July 16, 2026
  6. PrimaryBrief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident · METR · August 26, 2026
  7. PrimaryBrief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident | Redwood Research · Redwood Research · August 26, 2026
  8. PressHow OpenAI’s human mistake led to the AI-powered hack on Hugging Face · TechCrunch · July 22, 2026
  9. CommentaryInside the OpenAI – Hugging Face Incident: The AI Breach With No Human Attacker Behind It | TrendAI (US) · Trend Micro · July 23, 2026
  10. CommentaryMETR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack · Zvi Mowshowitz · August 29, 2026

Anthropic: evaluation incidents and Mythos Preview

Anthropic's disclosures about its own models, including its September alignment assessment.

  1. Lab disclosureInvestigating three incidents in our cybersecurity evaluations · Anthropic · July 30, 2026
  2. Lab disclosureAn alignment assessment of recent cybersecurity incidents · Anthropic · September 9, 2026
  3. PressAnthropic says its Claude models hacked three real companies during testing · Fortune · July 31, 2026
  4. Lab disclosureSystem Card: Claude Mythos Preview (PDF) · Anthropic · April 7, 2026
  5. PressAnthropic’s most capable AI escaped its sandbox and emailed a researcher – so the company won’t release it · The Next Web

The UK AI Security Institute incident

A government tester's own incident report on agents targeting real people.

  1. Tester reportIncident Report: unsanctioned agent behaviour during cyber testing | AISI Work · UK AI Security Institute · August 4, 2026
  2. Tester reportSecurity Incident INC-2026-07-28-01 (technical report, PDF) · UK AI Security Institute · August 4, 2026

OpenAI's wider disclosures: the wiki, RubyGems, and misalignment reports

What OpenAI has acknowledged beyond Hugging Face, and what it disputes.

  1. Lab disclosureThe Hugging Face incident and other third-party impact from misaligned models · OpenAI · September 5, 2026
  2. PrimaryDiscovery of a new OpenAI agent message board · Nightingale Collective (Von Arx, Byrd, Kitts, Larsen) · September 4, 2026
  3. PressOpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure · TechCrunch · September 5, 2026
  4. PressOpenAI Files EU Incident Report After DseWiki Episode; Commission Says Agent Control Has Been Lost Before · IBTimes UK · September 7, 2026
  5. Lab disclosureSigning up for disposable emails and searching GitHub for leaked API keys · OpenAI Alignment · OpenAI · September 16, 2026
  6. Lab disclosureOur framework for reporting model misalignment | OpenAI · OpenAI · September 16, 2026
  7. PressOpenAI reports 6 new instances of 'concerning model behavior' since March · CNBC · September 16, 2026
  8. PrimaryOpenAI agents carried out an undisclosed cyber-attack on RubyGems · Kitts, Larsen, Von Arx · September 11, 2026

Kimi K3 and open-weight models

The tester's report and coverage of the first reported open-weight escape.

  1. Tester reportChinese Model Kimi K3 Breaks UK AI Safety Institute Benchmark Evaluations · Frontier Security
  2. PressChinese AI model Kimi escaped its cybersecurity testing environment, researchers say · TechCrunch · August 7, 2026
  3. PressChina’s Kimi K3 broke out of its test sandbox. It didn’t need to hack anything · The Next Web

What lab leaders are saying

Primary essays from Anthropic's CEO, from upside to risk to pacing.

  1. PrimaryDario Amodei: We Must Pace the Frontier · Dario Amodei · September 12, 2026
  2. PrimaryDario Amodei: The Adolescence of Technology · Dario Amodei · January 2026
  3. PrimaryDario Amodei: Machines of Loving Grace · Dario Amodei · October 2024