The number that matters is one: the first model OpenAI has ever been unable to clear of its highest cyber-risk tier. In a blog post Friday, the company said internal evaluations of Astra, its upcoming model, showed strong enough performance in agentic coding and cybersecurity that it cannot rule out the Critical level of its Preparedness Framework, the safety rating system it has run since 2023. Critical, by the framework's own text, means a tool-enabled model that can find and develop working zero-day exploits of any severity in many hardened, critical real-world systems without human intervention, or that can devise and execute novel end-to-end cyberattack strategies against hardened targets given only a roughly defined goal. Every OpenAI model before this one, GPT-5.6 Sol included, has rated High at most.
What triggered it is a capability reading, not an incident. OpenAI says the past days' evaluations returned 'significant advancements' in agentic coding and cybersecurity, and it says plainly that Astra, an upcoming model, was not involved in exploiting Hugging Face. The framework's text for Critical says development stops until safety standards meeting that tier are specified, and OpenAI's actions stay just short of that line: it is pausing the internal Astra activities that do not yet meet the stricter guardrails, isolating test environments, restricting network and tool access, hardening and encrypting the model weights, and running a universal monitor over every agentic use of Astra, training and evaluation included, with Chain-of-Thought monitors that interrupt high-risk activity. Government agencies and select AI safety organizations are being brought in to test the model, and third-party test partners are getting recommended controls for high-risk evaluations.
Sam Altman confirmed the practical consequence on X: the launch will take 'a little bit longer for a safe rollout, but hopefully not too long.' In the same post he aimed at Anthropic, which ships its strongest model, Mythos, only to selected partners and governments: 'We don't think it's a good strategy to make capable models available only to a few powerful parties.' OpenAI researcher Noam Brown, in his own posts, urged taking the Hugging Face incident seriously, comparing it to the 2017 story in which Facebook's AI was said to have invented its own language, an account that turned out to be overblown. This one, he wrote, is not that, and he tied the concern to test-time compute, arguing current models can be pushed much further before any plateau.
The weighing that matters is between two true sentences. The first: this is a potential, not a classification. OpenAI says it cannot rule out Critical; it has not rated Astra Critical, and if the tier is never reached, the company will have made a lot of attention out of a model it then ships anyway. The Decoder expects the fear-marketing accusation to follow, and the accusation has a prediction attached: if this is publicity, the pause quietly ends when the cycle moves on. The second: a framework is only real when it costs something, and this week it cost a launch window and a public admission, on camera, that the company's next model might be too capable to release on schedule. Watch the pattern from here: whether OpenAI publishes the final rating when it exists, and whether Anthropic and Google now face the question of rating their own frontier models in public, with the same framework language or without it.
