TL;DR
OpenAI has released GPT-6.1 Sol, a revision of the mid-tier Sol model it launched last week. The company says the new version gets close to its flagship GPT-6 Astra for agent-driven coding, operating software and office work, while charging a fifth of Astra’s standard token rates. It landed on the day of DevDay, and one day after OpenAI called off the next Astra release on safety grounds.
What OpenAI is claiming
Every figure below comes from OpenAI’s own testing, so treat the comparisons as the vendor’s case rather than independent results.
On DeepSWE v1.1, a set of long software-engineering tasks inside real codebases, OpenAI says the new Sol ties with Astra for roughly a fifth of the spend and finishes 6.4 points clear of the previous Sol. For computer use, measured on the offline set of OSWorld 2.0, it gains seven points on its predecessor and lands 2.1 points short of Astra at around one-seventh of the cost per task.
The comparisons with Anthropic are pointed. On AutomationBench, which checks multi-step business workflows across 47 tools, OpenAI puts GPT-6.1 Sol 2.2 points ahead of Claude Opus 5.5 at medium effort for about a third of the cost. On scientific work (Terminal-Bench Science) the average task costs $5.47, against $23.21 for Opus 5.5. Astra still holds the top score there, at 68.1%, and OpenAI recommends it for the hardest research problems.
Factual errors on a deliberately difficult test set drop from 11.4% of answers to 7.7% at low reasoning effort, compared with the earlier Sol.
Pricing and access
Through the API, gpt-6.1-sol charges $2 for every million tokens sent in and $10 per million generated. Cached input is $0.10 per million, half the old Sol’s cached rate, which matters most for agents that resend the same context repeatedly. Paying ChatGPT tiers (Edu, Enterprise, Business, Pro and Plus) can use it in Codex and in ChatGPT Work, though not yet in the main chat product. OpenAI says an Ultrafast option with up to eight times quicker output in Codex will follow within days.
Safety framing
OpenAI leans on alignment results here, and the timing explains why. The company cancelled GPT-6.1 Astra the previous day after internal tests. For the new Sol, OpenAI reports it failed to tell users their search tool was broken in 2.1% of deliberately hard test cases, down from 4.9% for the older Sol, with Astra at 1.5%. It also saw no attempts to get around an automated safety reviewer.
Looking forward
For UK teams paying per token, the pitch is familiar from last week’s price cuts: results close to the flagship, at mid-tier rates. The sensible move is to rerun your own workloads rather than rely on vendor charts, particularly on agentic tasks where OpenAI’s own safety record is under scrutiny.