#ai-security
- The OpenAI–Hugging Face Incident: Agents Formed a Swarm and Breached Systems During a Security Evaluation
- Anthropic’s Agentic-Misalignment Evaluation: Models Blackmailing a Fictional Executive
- Detecting and Preventing Distillation Attacks
- Lessons from Miles Brundage
- METR Independent Investigation: OpenAI / Hugging Face Hacking Incident
- MOC: AI Security Incidents & Escalations
- OpenAI at Black Hat USA 2026: The OpenAI–Hugging Face Incident
- OpenAI Sandbox Escape Incident