Related Astra coverage
OpenAI says it cannot rule out Critical capability for Astra
That wording is deliberately narrower than saying Astra has been confirmed Critical. The difference matters because capability thresholds are used to decide how strongly a model should be contained while evaluation continues.
Open the August 9 reportThe direct answer
“Critical” means a model may be capable of highly consequential cyber operations with a level of autonomy and real-world effectiveness that changes the required security posture.
OpenAI’s framework does not define Critical as “very good at cybersecurity.” A model can solve difficult security tasks and still remain below the threshold. The concern rises when capability starts to remove the hardest operational bottlenecks: independently finding serious weaknesses, developing functional exploits against hardened real systems, chaining steps into novel attacks and carrying out an end-to-end objective with very little human guidance.
“Astra is strong at cyber tasks.” “OpenAI cannot rule out Astra meeting Critical.” “Astra has been confirmed Critical.” Only the middle statement matches OpenAI’s current public wording.
Critical is an OpenAI framework label, not a universal safety rating
Different laboratories use different preparedness frameworks, risk taxonomies and evaluation suites. OpenAI’s Preparedness Framework creates capability thresholds that trigger progressively stronger safeguards in areas including cybersecurity and biological/chemical capability.
That means readers should not treat “Critical” like a standardized aviation category or a common regulatory grade. A different AI developer may use different labels for comparable concerns, or place the boundary at a different point.
The useful question is therefore not “Is this model Critical everywhere?” It is “What threat model and evidence does this laboratory use when it applies that label?”
OpenAI’s High threshold already describes serious cyber capability
OpenAI defines High cybersecurity capability around removing important bottlenecks to scaling cyber operations. Its public GPT-5.6 system card gives examples such as automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities.
GPT-5.6 Sol, Terra and Luna are currently treated as High in cybersecurity but below Critical. That is an important comparison because it shows that Critical is not simply the next adjective after a model becomes useful to professional security work.
At High, a model can already be valuable to defenders and potentially to attackers. The risk-management question is how reliably and autonomously it can operate, against what class of target, and whether safeguards make malicious use at scale difficult.
Critical raises the bar from assistance and automation to much stronger autonomous offensive capability
OpenAI’s published Critical definition includes the ability to identify and develop functional zero-day exploits across many hardened real-world critical systems without human intervention, or to devise and execute novel end-to-end attacks against hardened targets from only a high-level desired goal.
Several concepts are doing work in that definition:
- Zero-day: a vulnerability for which a practical fix may not yet be available to defenders.
- Hardened target: a system designed and operated with meaningful security controls, not an intentionally easy laboratory target.
- End-to-end: the system can connect multiple steps of an operation rather than solve one isolated puzzle.
- Novel strategy: the model is not merely replaying a fully specified procedure.
- Little or no human intervention: autonomy is part of the threshold, not an afterthought.
Different tests answer different questions
A high score on one cyber benchmark does not prove real-world autonomous compromise
Capture-the-Flag tasks
Bounded security challenges test whether a model can reason about vulnerabilities, exploitation and system behavior in a controlled environment.
Known-vulnerability benchmarks
Tests such as CVE-style evaluations examine whether a model can identify or exploit real vulnerability patterns under specified conditions.
Cyber ranges
Realistic emulated networks test whether the model can plan and chain actions across a larger operational environment.
Open-ended vulnerability research
Long-horizon tests ask whether the model can discover promising attack surfaces, investigate crashes and turn new bugs into meaningful exploit primitives.
These evaluations complement rather than replace one another. Solving CTF problems demonstrates technical capability. It does not by itself establish that a model can independently compromise a hardened production system.
Why autonomy changes the cybersecurity risk calculation
A human-directed assistant that explains a vulnerability or writes a short exploit fragment operates inside a different risk envelope from a system that can select a path, use tools, adapt when a step fails, persist through a long workflow and complete the objective.
Autonomy can reduce the number of skilled human decisions needed to run an operation. If capability becomes reliable enough, that can change the scale and accessibility of offensive cyber activity.
This is why the threshold is not about raw knowledge alone. The same underlying technical knowledge becomes more consequential when the model can turn it into sustained action.
What “cannot rule out Critical” means in the Astra disclosure
OpenAI says preliminary evaluation results are strong enough that it cannot currently rule out Astra reaching the Critical cybersecurity threshold. It has not publicly said the threshold is confirmed.
That wording reflects an evaluation problem: the model appears capable enough that assuming it is safely below the threshold would be irresponsible until testing provides stronger evidence.
The appropriate headline is therefore about uncertainty plus precaution. Astra is an unreleased model under evaluation, and the uncertainty has already been serious enough for OpenAI to strengthen development controls.
Capability evaluations are measurements, not mathematical proofs of a model’s ceiling
Models can perform differently with better prompting, more reasoning effort, additional tools, longer rollouts or improved agent scaffolding. OpenAI’s own system-card material cautions that no evaluation represents every product configuration or real-world workflow.
That creates uncertainty in both directions. A model can look weaker because the test failed to elicit its best capability. A benchmark can also exaggerate practical danger if the test environment gives the model unrealistic support that would not exist in deployment.
Therefore “below Critical in this suite” is not the same as “incapable of every dangerous cyber action,” and “cannot rule out Critical” is not the same as “proven capable of every behavior in the definition.”
Why security controls may have to tighten before classification is settled
If a developer waits for perfect certainty before strengthening containment, the evaluation process itself can become the risky period. An unreleased model may already have access to tools, test environments, credentials or networks while teams are trying to determine what it can do.
OpenAI says it has strengthened isolation, network restrictions, model-weight protection, sandboxing, monitoring and related development controls around Astra while evaluation continues.
When a model may be capable enough to create serious cyber risk, the development and evaluation environment becomes part of the safety system—not merely the public product interface.
This is the broader lesson for organizations using capable agents: prompts are not containment. Permissions, credentials, network controls, sandboxing, audit logs and reliable interruption mechanisms must exist outside the model.
Sources