Disclosure: this site is written with Claude, made by Anthropic — a direct competitor to OpenAI. Read what follows knowing that.
On 6 September 2026, OpenAI began rolling Astra out to $20 ChatGPT Plus subscribers, appearing first in ChatGPT Work and then in regular chat. Pro, Enterprise and Business Premium already had it. Access is included within existing subscription limits, with additional usage available as separately purchased credits.
First, what this is not
It is not a broken promise, and the coverage that frames it that way is wrong.
OpenAI's Path to Astra post on 1 September set out the sequence explicitly: Daybreak participants first, followed in the coming days by paid ChatGPT plans and the API. Six days later, paid plans. That is a company doing what it said it would do, on the timeline it published.
It is also not the case that the Critical-threshold capability is now on a consumer plan. OpenAI's own framing is that "Critical" gates a capability, not the whole product — autonomous zero-day discovery and offensive exploit generation stay restricted to vetted defenders through Daybreak Blue. Both versions, OpenAI says, carry safeguards preventing access to its most advanced cybersecurity capabilities.
All of that is accurate and worth stating before the criticism, because the criticism is narrower and more specific than "they released it anyway".
What actually changed
Look at what the word "safeguard" is doing in each case.
On Daybreak Blue the controls are identity verification, legal attestations, approved-use restrictions and account monitoring. All 4 are controls on a person. They do not make the model refuse anything — they make the requester known, accountable and revocable, and they mean a persistent abuser leaves a trail with a name on it. On a $20 Plus plan there is no identity verification and no attestation.
On a $20 Plus plan there is no identity verification and no attestation. The control is inside the model: a refusal policy, plus whatever additional guardrails OpenAI has applied to the consumer build.
Those are not the same kind of thing. One constrains who is asking. The other constrains what the answer is. Describing both as "safeguards" is technically true and flattens the distinction that matters, because only one of them still works on the ten-thousandth attempt by someone with time.
OpenAI published the number for the second kind
To its credit, OpenAI did not leave this to inference. It reported that Astra declines 91.5% of cyber-related jailbreak attempts, against 59% for GPT-5.6 Sol. That is a large, real improvement and it should be said plainly. It also means roughly 1 attempt in 12 is not declined — and whether that matters depends entirely on what is standing behind the model when the twelfth one lands.
Behind Daybreak Blue, an 8.5% residual sits underneath identity verification and account monitoring. Someone grinding at it is a named account being watched, and the residual is a backstop behind a gate.
On a consumer plan, the residual is the gate. There is no identity behind it, no attestation, no approved-use restriction, and the population of people trying is not a vetted list of six named security vendors.
An 8.5% failure rate against a hundred determined attempts and an 8.5% failure rate against a hundred million are the same percentage and completely different exposures. Nothing published lets anyone outside OpenAI work out which side of that the consumer build actually sits on.
The thing that is not published
What technically separates the Plus build from the Daybreak Blue one has not been described.
Different weights? A different system prompt? A classifier in front of the request? A tool or capability gate? "Added cybersecurity guardrails" is a category, not a mechanism, and the difference between those options is enormous — a refusal policy is a different security property from a capability that is not present.
Without that, nobody outside the company can assess how far the consumer surface can be walked back toward the gated capability, which is exactly the question a Critical classification is supposed to make people ask.
What to do
- Do not repeat "the Critical model is now $20". It is not the claim OpenAI made and it is not established.
- Do not repeat "OpenAI broke its promise" either. The timeline was published on 1 September.
- If you buy this for your organisation, ask which build you get and what separates it from the gated one. You are entitled to a mechanism, not an adjective.
- Assume the refusal rate is the control on any consumer tier, and plan for the residual rather than the average.
- Watch for the first published jailbreak. The 8.5% is not hypothetical, and the first credible public demonstration against the consumer build is the thing that will settle this argument.
What is not established
- Whether the Plus build is the same weights as the one evaluated at the Critical threshold.
- What the "added cybersecurity guardrails" are in technical terms.
- Whether the 91.5% refusal figure applies to the consumer build, or was measured on a different configuration.
- What usage limits apply beyond "existing subscription limits" and purchasable credits.
- When or whether free-tier access follows. No timeline has been given.