GPT-6 Astra: what actually changed, and whether you need it
OpenAI's first GPT-6 model is built to operate software rather than write it. The honest read on what that is worth, and when the cheaper tier is still the right call.
GPT-6 Astra shipped on 3 September 2026 at $10 per million input tokens and $50 per million output, with a 1.05M-token context window and 128K maximum output. It is a generation change rather than a point release, and the thing it changes is not text quality. It is that the model is built to drive software.
The one capability that is genuinely new
OpenAI describes Astra as state of the art on computer use, browsing, software engineering, cybersecurity and science. Strip the list down and the load-bearing item is computer use: interpreting what is on a screen and acting on it, across applications, without a human clicking each step. Filling forms, moving through spreadsheets, navigating a web app, assembling the result into a document.
This is a different axis from the one the last two years were fought on. Model comparisons have mostly been about the quality of generated text and code. Astra is a bet that the next constraint is not what a model can write but what it can operate. Whether that bet pays off in your work depends entirely on whether your bottleneck is a blank page or a browser tab.
The version you get is not the version in the benchmarks
Astra is the first model OpenAI has rated as reaching a Critical cybersecurity capability level under its Preparedness Framework. The version available to customers refuses advanced cybersecurity tasks. If you are reading a headline benchmark on security work, check which model produced it before you plan around the number.
What it costs, in context
The list price is $10 and $50 per million tokens for short context, rising to $20 and $75 once a request crosses into the long-context band. That is worth working through against the alternatives, because the gaps are large:
- GPT-5.6 Sol, the tier immediately below, is $4 and $20 with the same 1.05M window. Astra costs two and a half times as much for the same context
- Claude Opus 5 is $5 and $25, and Anthropic recommends it as the default for most workloads
- Claude Fable 5.1 matches Astra exactly at $10 and $50
- Gemini 3.1 Pro is $2 and $12, a fifth of Astra's price at frontier tier
For generation, extraction, summarisation and chat, none of those gaps are justified by output quality. The frontier models have been within a few points of each other for a year. Astra earns its price on the specific work it was built for, and nowhere else.
When to reach for it
- The task means driving a real interface rather than producing text about one
- The work spans several applications and a mid-task failure is expensive to recover from
- You have measured a cheaper tier failing, rather than assumed it would
When not to
- Coding inside one repository. Opus 5 leads the agentic coding indices at half the price
- High-volume classification, routing or extraction. GPT-5.6 Luna does this at $0.20 and $1.20, a fiftieth of Astra
- Anything where the frontier was already comfortably good enough, which is most things
A generational model is not a general upgrade. It is a specific capability, priced as though it were general.
The pattern worth taking away
Both frontier labs shipped a new top model within 48 hours at the same price, and each is pointed at a different job: Astra at operating software, Fable 5.1 at reasoning over hard problems. The era where one model was straightforwardly the best is over, and it is not coming back. The useful question stopped being which model is strongest and became which model is shaped like your problem.
The practical consequence is that pinning a provider into your code is now the expensive decision, not the model choice itself. Build against an abstraction layer and switching costs an afternoon.
Full side-by-side numbers on the model comparison, and the family-level breakdown in Claude vs GPT. Prices are list rates as of 4 September 2026.
Find these on the Radar
Every tool here lives on Kapyn Radar. Save the ones that fit into a Loadout and find them again.