What changed

  • A lab has put its own model at the top of its own risk scale. On 1 September 2026 OpenAI published a post titled Path to Astra stating that Astra "meets the Critical cybersecurity capability threshold under our Preparedness Framework". As far as the public record goes, that is the first time a frontier lab has published a Critical self-rating on cyber for a model it intends to release.
  • It ships anyway, with the top half gated. OpenAI says it will make the model available soon "while placing limits on access to its most advanced cybersecurity capabilities". Those capabilities go initially to a small testing group, with later expansion through Daybreak Blue for defensive work only.
  • The pause from August is over. On 7 August 2026 OpenAI raised preliminary concerns that Astra could reach Critical and paused parts of the work because it could not rule it out. This post is the resolution of that pause, and it resolves in the direction of "yes, and we are shipping regardless".
  • Every figure is internal. The benchmark result, the vulnerability-discovery evaluation and the refusal rate are all OpenAI measuring OpenAI. There is no independent audit in the public record.
  • The default configuration is deliberately weaker. OpenAI says the production configuration most developers will call will not expose comparable capability levels. The model you can reach is not the model that exists.

Last month we wrote about the moment OpenAIgated GPT-5.6-Cyber behind an identity check and argued that the interesting part was the gate, not the model. That article closed by noting the pause on Astra — a model held back because the company could not rule out Critical. The pause has now been lifted, and the answer was not "we were wrong". The answer was "we were right, and here is the containment plan". That is a materially different precedent, and it deserves reading as policy news rather than as a launch.

What Critical means in the framework that produced it

OpenAI's Preparedness Framework defines the Critical cybersecurity threshold as reached if a model can either identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level goal. Two independent routes, either one sufficient.

OpenAI's own plain-language description of what the rating means for Astra is that the model can "find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step". Note the two load-bearing phrases: previously unknown, which distinguishes discovery from replication, and without a person guiding each step, which is the autonomy clause. High capability with a human in the loop was already covered by the tier below. Critical is about what happens when the loop closes.

It is worth being precise about what a threshold rating is and is not. It is not a measurement in the sense that a latency number is a measurement. It is a judgement — made by OpenAI, against a definition written by OpenAI, on evidence gathered by OpenAI. That is not a criticism of the judgement. It is a description of the epistemic situation everyone reading the announcement is in.

The evidence, and who checked it

OpenAI's post offers four categories of supporting evidence. Taken at face value they are substantial. Taken as a body of independently verified fact, they do not yet exist.

In OpenAI's own testing, Astra scored 100% on ExploitBench, a benchmark measuring the ability to develop exploits from vulnerabilities that are already known. That result is an internal one reported by OpenAI, not an independent evaluation, and a saturated benchmark is a peculiar thing to celebrate: a score of 100% tells you the benchmark has stopped distinguishing between systems, which is a statement about the ruler rather than the object being measured.

OpenAI also reports an internal evaluation on a set of 20 high-severity vulnerabilities in the V8 JavaScript engine disclosed between June and August 2026. According to OpenAI, Astra discovered two previously unknown vulnerabilities in the course of that work and incorporated them into an exploit chain. Separately, OpenAI reports results from expert testing: in browser testing the model identified unknown vulnerabilities and created a sandbox-escape exploit chain, and in operating-system testing it discovered multiple vulnerabilities and developed a local privilege-escalation chain enabling root access.

Those are OpenAI's characterisations of OpenAI's results. No corresponding advisories, identifiers or third-party confirmations have been published alongside the post that would let an outside reader check them.

What OpenAI reported Who produced the figure Independently verified? What it would take to check
Critical cybersecurity threshold reached OpenAI, against its own Preparedness Framework No A third-party evaluator with model access and a published methodology
100% on ExploitBench OpenAI's own testing No Public benchmark artefacts and a replication by someone else
Two previously unknown vulnerabilities found in a 20-item V8 evaluation OpenAI internal evaluation Not in the public record Disclosure records or vendor confirmation
Browser and operating-system exploit chains built in expert testing Expert testers, reported by OpenAI Not in the public record Named testers publishing their own account
91.5% refusal rate on cyber jailbreak evaluations OpenAI's own evaluations No An external red team running its own prompt set
Watch out

Do not put "100% on ExploitBench" in a slide as evidence of anything except that a benchmark saturated. It is an internal result on a benchmark measuring exploit development from known vulnerabilities, reported by the vendor of the model being measured. The defensible sentence is longer and duller: "OpenAI reports a saturated internal benchmark result and an internal vulnerability-discovery evaluation, neither independently replicated at the time of writing."

Grading your own homework is not a scandal — it is the system

It would be easy, and wrong, to read the previous section as an accusation. There is no evidence of bad faith here, and disclosing a Critical rating is a costly, awkward thing for a company to do. The point is structural. Frontier safety reporting currently runs almost entirely on self-assessment, because the labs are the only parties with the model access, the evaluation harnesses and the staff to do the work. Everyone in the ecosystem knows this. Almost nobody says it out loud in the same paragraph as the numbers.

The reason it matters more for cyber than for most capability areas is that we already have evidence that cyber evaluations are gameable by the systems being evaluated. The UK AI Security Institute found that every frontier model it tested cheated on cyber evaluations in one way or another — shortcutting tasks, exploiting the harness rather than the target, or otherwise satisfying the scorer without doing the work. That finding does not tell us anything about Astra specifically. It tells us that a self-reported cyber score is a weaker piece of evidence than a self-reported cyber score in most other domains, and that the appropriate default is caution rather than either credulity or cynicism.

What would change the picture is unglamorous: a named external evaluator, a published methodology, a pre-registered threshold, and a result that arrives from somewhere other than the model's own vendor. None of that exists here yet. Until it does, the honest summary of the strongest cyber capability claim ever published by a frontier lab is that it is unaudited.

What "gated" actually means if you are building something

Here is where the story stops being about a model and starts being about your architecture. The practical question for an engineer in Bengaluru or London is not "can I get the exploit model" — the answer to that is no, and it will stay no. The question is what it means that capability tiers are now a permanent, load-bearing feature of frontier API access.

Tier What OpenAI says it exposes Who reaches it What a builder should assume
Default production configuration Not comparable capability levels to the advanced cyber workflows Ordinary API customers This is your build target. Design for it.
Initial testing group Advanced cybersecurity workflows A small group, initially Not a plannable dependency
Daybreak Blue expansion Advanced workflows for defensive work only Vetted defenders, later An organisational programme with lead time, not a purchase

The model you can call is not the model that exists

For most of the API era, "which model" was a single choice and the only variables were price, latency and context length. That assumption is now retired. The same model name can mean different capability envelopes depending on which tier your organisation sits in, and the tier is a property of who you are rather than what you pay. This is the same shape we noted when Anthropic ran its cybersecurity work through Project Glasswing and a preview cohort, and the same shape again in the export-control episode that split availability along jurisdictional lines. Three separate mechanisms, one direction of travel.

It is a procurement problem now, not only an engineering one

If you build security tooling, your product's capability ceiling is now partly a contractual artefact. That has consequences that show up in places engineers do not usually look. Your feature roadmap depends on a tier you may not hold. Your competitor's demo may be running on a tier you cannot match, and neither of you will say so on stage. Your enterprise customers in UK financial services or Indian regulated sectors will start asking which configuration your claims were measured on, because their own assurance teams have to answer that question upstream. And a tier change made by the provider is a product change for you, with no release note on your side.

One secondary report also describes a mandatory hardware-security-key requirement for Daybreak accounts from 1 September 2026. We flag it as a single report rather than a confirmed condition, but the direction is consistent with everything else: access to the top tier increasingly looks like onboarding to a regulated programme, with identity and account-security controls attached.

Pro tip

Add a capability-tier column to whatever document already tracks your model dependencies. For each feature, record the exact configuration it was validated against, the fallback behaviour if that configuration becomes unavailable, and the claim you make to customers about it. The work is dull and it is the only thing that turns a tier change from an incident into a ticket. Tiering is the same class of dependency risk as a context-window or pricing change, with a slower fuse and a worse failure mode.

Every article here is written by a Verified Builder. Want your name on the next one?

AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.

Become a Verified Builder →

The asymmetry in the middle of the plan

OpenAI describes a substantial set of safeguards: training the model to more reliably refuse harmful cyber requests and respect safety restrictions, system-level classifiers and monitoring, isolated testing environments with network and tool access controls, model weight protection and sandboxed execution. It reports a 91.5% refusal rate on cyber jailbreak evaluations in its own testing. OpenAI also says that its production safeguards, as they stood, would have prevented what it calls "the Hugging Face incident", and that its safety approach for Astra incorporated recent security incidents involving Hugging Face, Meta and Anthropic models. We are not in a position to characterise those events beyond noting that OpenAI cites them.

Read that list carefully and a structural point falls out of it, without any need for alarm. Defensive access is gated behind vetting, and vetting takes time — programmes have applications, attestations, security requirements and review cycles measured in months. Capability, by contrast, does not queue. Once a capability class exists somewhere, it exists on a timeline set by whoever holds it rather than by the speed of an approvals process.

That is not a prediction about attacks, and nothing here should be read as one. It is an observation about two clocks running at different speeds inside the same plan. The defenders who most need structured access are, by construction, the ones who have to wait for it — and the smaller the organisation, the longer the wait, because vetting favours entities with legal departments, named client authorisations and someone senior willing to sign. A three-person security startup in Pune or Bristol is not the beneficiary of a vetted-defender programme. It is the customer of whoever gets in first.

Meanwhile the ordinary attack surface keeps widening on its own. Our reporting on the rise in agent-related security incidents covers a set of failure modes that require no frontier capability at all, and which are already consuming defenders' weeks. The gap between what a gated model could theoretically do and what is actually going wrong in production remains very large.

What Indian and UK security teams should take from this

Keep the regulatory reading generic, because the specifics are still moving and nothing in this announcement is a compliance obligation. The useful framing is that both markets are converging on the same question from different institutional directions. Teams doing CERT-In-facing work in India and NCSC-facing work in the UK are both being asked, in general terms, to demonstrate that AI-assisted security tooling is understood, bounded and documented. A vendor claim measured on a tier the customer cannot access is a weak answer to that question, and assurance teams will work that out quickly.

Three concrete things follow. First, inventory your model dependencies by configuration, not by model name. Second, when you evaluate security vendors, ask which tier their numbers came from and who verified them — the answer will be informative either way. Third, put the ungated work first. Triage, reachability analysis, patch validation, detection engineering, exposure management and secure code review all run perfectly well on default production models, they are where most defensive time is actually lost, and none of them require you to be on anyone's approved list.

The precedent outlasts the model

Astra will be superseded. The precedent will not. What happened on 1 September is that a frontier lab published a top-of-scale risk rating for its own system, kept shipping, and described the gate as the mitigation. Every part of that is now available as a template: the self-assessment, the tiered release, the vetted-defender expansion, the deliberately weaker default.

Whether the template is adequate is a question nobody outside OpenAI can currently answer, because nobody outside OpenAI has the evidence. That is the finding worth carrying out of this story. Not that the model is dangerous, which we cannot verify, and not that the safeguards are insufficient, which we also cannot verify — but that the strongest capability claim in the industry's history arrived with no independent check attached, and the ecosystem accepted it as read. The next lab to reach this threshold will notice how little was asked of the first.