Skip to content
tag — ai-security

grep -rl "ai-security" ./articles

#ai-security

27 articles

On Daybreak Blue the control is who you are. On a $20 plan it is whether the model says no

2026-09-07AI

Astra began reaching ChatGPT Plus subscribers on 6 September, three days after OpenAI said it was the first model to meet the Critical cybersecurity threshold of its own Preparedness Framework. That was the announced plan and it gates a capability rather than the product — but the safeguard changed from identity verification to a refusal policy, and OpenAI published the refusal rate: 91.5%.

The attacker asked METR's agent for its API key, and the agent handed it over

2026-09-03AI

A research non-profit that evaluates frontier models for dangerous capability lost an API key because an authentication check failed open, and the exfiltration method was prompting the agent to reveal it. Three weeks of use would have billed at about $600,000 — the exact figure matters less than how the key left.

Anthropic will give defenders what its strongest security model finds — but not the model

2026-08-25AI

Claude Security now scans code with Mythos 5, the model Anthropic keeps most tightly restricted. Customers never touch it; they get findings with a CWE category, severity, confidence and a suggested patch. Alongside it, a $35 million fund pays open-source maintainers in Claude credits. The whole design is a bet that findings can be shared when the capability cannot.

The safety filter read the ciphertext and the sandbox ran the plaintext

2026-08-22AI

Adversa AI encrypted its instructions so guardrails saw only harmless-looking ciphertext, then let the model's own code sandbox decrypt and execute them. It reported the technique to xAI on 3 June, chased twice, got no reply, and published. Grok still falls to it, including zero-click exfiltration through tool use.