The model refused to run the attacker's decoder, so it wrote its own — and that was the exploit
2026-08-31AI
Johann Rehberger chained a series of individually harmless steps into code execution on a machine running Claude Code in Auto Mode, starting with nothing more than asking it to summarise a website. The step that makes it work is the safety refusal. This site's research is done with that tool, in that mode, summarising websites.