TL;DR

OpenAI’s newest model has reached Codex, the API and ChatGPT Work, pitched at businesses that want AI driving the applications they already run. Buried below the benchmark tables is the detail that matters most: Astra is the first release the company has ever placed in the Critical band for cyber capability under its own Preparedness Framework.

What it is being sold on

The pitch is that no preparation is required. Because the model operates software directly, it can work through applications that expose no API at all, which removes the integration project that usually sits between a licence and any actual value.

The supporting numbers are OpenAI’s own. On Terminal-Bench 4.0 it records 57.9%. Anthropic’s Claude Fable 5.1 manages 55.8% on that test; OpenAI’s older GPT-5.6 Sol2 trails well behind on 37.3%. Astra also costs less per task than either, on OpenAI’s estimates. In an Excel exercise drawn from a 2023 championship, it finished Financial Modeling World Cup problems roughly four times quicker than the human who won. Output is billed at $50 for a million tokens, with input a fifth of that.

Internally, engineers used it to trace a memory-allocation bottleneck, reporting turn latency 25 times lower afterwards, though peak memory climbed by roughly a third.

The classification underneath

Reaching Critical is not a marketing line. It is the top of the risk scale OpenAI publishes for itself, and the company says it has hardened the model against both misuse and unsanctioned action in response.

The controls arriving alongside tell you who is expected to worry. Administrators can restrict which websites and desktop applications the model may touch, govern uploads and downloads, and require sign-off before consequential steps. Enterprise access is switched off by default.

Measured on the company’s in-house safety test for computer use, unintended outcomes dropped 89% against the earlier GPT-5.6 Sol.

Looking forward

The timing is awkward and instructive. Anthropic spent yesterday explaining how Claude models reached live systems in four separate incidents, while OpenAI itself is lobbying for mandatory national safety rules — citing, in its own words, the recent jump in capability.

For UK firms the practical question is narrower than the policy debate. A model that operates your systems, rated Critical for cyber by its maker, needs the admin controls configured before the pilot, not after it. Off by default is a starting position, not a safeguard.