
Significantly lowers the cost of AI coding agents and improves reliability for analyzing very long business contracts or technical docs.
What is GPT-6 Astra and what does it do?
GPT-6 Astra is OpenAI’s new flagship model for high-end coding, security work, and long document analysis. It began rolling out on September 3 to a limited set of organizations, and OpenAI says all ChatGPT Plus, Pro, Business, and Enterprise users get it over the coming days, along with API and AWS access.
The API label is gpt-6-astra, and OpenAI’s own pricing documentation lists it at $10 per million input tokens and $50 per million output tokens. That’s the same headline rate as Anthropic’s Claude Fable 5 and 5.1, and it’s a 2.5x jump from GPT-5.6 Sol’s $4 per million input tokens.
One safety detail matters for anyone buying security tooling: OpenAI designated Astra the first model to meet its Critical cybersecurity capability threshold under the Preparedness Framework. The company delayed parts of Astra’s development and release while it strengthened protections against cyber misuse and unauthorized model actions.
Astra is a specialized agentic model priced for volume, not another general chat upgrade.
How accurate is GPT-6 Astra?
The security benchmarks are the core of this launch. Simon Willison’s analysis of the launch numbers confirms Astra scores 100% on ExploitBench, where GPT-5.6 Sol scored 78.5%, and 42.4% on ExploitGym against Sol’s 30.3%.
On SRE-Bench binary reverse engineering, Astra hits 99.2% within 4 attempts, up from Sol’s 68.7%. Long context is the other headline: OpenAI’s eight-needle benchmark returned 100% accuracy at 256K to 512K tokens and 96.3% from 512K to 1 million tokens, with a 128K token max output cap on the API.
The honest asterisk sits on ARC-AGI 3. Astra’s 99.9% score came from OpenAI’s custom Provider Adapter harness at a measured cost of $19,000, while the default harness scored 62.7% for $26,000, so the harness behind a benchmark number matters as much as the model itself.
The security and long-context numbers hold up, as long as you read the benchmark conditions first.
GPT-6 Astra vs Claude Fable 5: which is better for a small business?
Artificial Analysis benchmarked GPT-6 Astra against the field on both of its indexes. On the Intelligence Index, Astra sits beside GPT-5.6 Sol at 61, which is 5 points behind Claude Fable 5.1, and it also trails Meta’s newly released Muse Spark 1.3.
Agentic coding is where the math flips. At max effort on the Coding Agent Index, Astra scores 2 points above GPT-5.6 Sol for about the same cost per task, and it finishes tasks for less than half of Claude Fable 5’s cost at the same score.
DataCamp’s benchmark roundup adds a third data point: Claude Opus 5 scored 70% on ExploitBench against Astra’s 100%.
For a business buying completed tasks rather than raw intelligence, that spread is the entire purchasing decision. Same score, less than half the invoice.
Fable 5.1 wins on raw intelligence, Astra wins on cost per completed task.
Who is GPT-6 Astra actually for?
This model targets teams running autonomous coding agents, security audits, and bulk document review, not general conversation. If your bill scales with tasks completed rather than questions asked, the daily AI pricing signals we track point the same direction this launch does: capability held flat while the per-task cost got cut in half.
The fit test is concrete. Long contracts, large codebases, and vulnerability research map to the 96.3% reliability at up to 1 million tokens, while a 200-word email draft doesn’t need a Critical-capability security model behind it.
Astra is for workloads where reliability at volume converts directly into money.
The printing press technician starts a 1,000-page catalog run on the heavy offset press. He checks the first 10 pages, confirms the colors are spot on, and lets the machine rip through the rest of the stack.
Hours later he finds a registration shift on page 400 that makes every page after it unreadable, and the job was reported as running smooth the whole time. That’s the failure mode of a model that loses the thread mid-document, and it’s the one Astra’s 100% accuracy at 256K to 512K tokens is built to close.
The fine print on page 300 of a contract is where a hallucinated clause turns into a signed problem.
Should you switch to GPT-6 Astra?
If you run coding agents or document audits, test gpt-6-astra against your current model this week and measure cost per completed task, not cost per token. The standard rate is $10 per million input tokens with cached input at $1, and long-context input runs $20 per million.
Keep Fable 5.1 for reasoning-heavy work where its 5 point Intelligence Index lead pays for itself, and route the high-volume agentic queue to whichever model finishes the same score for less. Right now that model is Astra.
Move the agent workload now, keep the premium general model for the work that actually needs it.
Source: simonwillison.net