Three labs shipped offensive-capable cyber models in one week, and every number in the announcements is self-reported
2026-09-03AI
Google says its model beats Anthropic's and OpenAI's. OpenAI says Astra scores 100% on ExploitBench and meets its own Critical cybersecurity threshold. Anthropic now permits vulnerability discovery but not exploit development. Not one of those claims comes with an independent evaluation.