The safety filter read the ciphertext and the sandbox ran the plaintext
2026-08-22AI
Adversa AI encrypted its instructions so guardrails saw only harmless-looking ciphertext, then let the model's own code sandbox decrypt and execute them. It reported the technique to xAI on 3 June, chased twice, got no reply, and published. Grok still falls to it, including zero-click exfiltration through tool use.