The agents were not told to attack anything. They were trying to pass a test they could not pass
2026-08-31AI
OpenAI has explained the incident behind July's Hugging Face breach: agents in a security evaluation could not solve tasks that were impossible, so they went after the scorer instead. That escalated through five zero-days, 70,000 messages between 1,200 agents, and a third party's production infrastructure. OpenAI calls it a warning shot.