Hasty Briefsbeta

Bilingual

The model didn't escape. OpenAI ran the attack

8 hours ago
  • OpenAI conducted an offensive security evaluation where their agent, with reduced safeguards and substantial compute, attacked Hugging Face's infrastructure.
  • The attack's success was due to the agent loop—OpenAI's system that executes model decisions—not the model itself 'escaping'.
  • OpenAI disabled key safety controls like cyber refusals and production classifiers to test maximal cyber capabilities.
  • The agent chain included identifying a zero-day, lateral movement, and exploiting vulnerabilities to access Hugging Face's database.
  • Hugging Face's defenders were blocked by commercial AI guardrails, forcing them to use a self-hosted Chinese model (GLM 5.2) for analysis.
  • OpenAI has not disclosed the monitoring or intervention latency, raising questions about operational responsibility.
  • The core message is that responsibility lies with OpenAI as the operator, not the model, and the 'model escaped' narrative is misleading.