Skip to content
AI NewsGPT-6 AstraOpenAIClaude Fable 5.1AI BenchmarksComputer UseAI Pricing

GPT-6 Astra Review: Benchmarks, Price and the AGI Claim

GPT-6 Astra is here: 100% ExploitBench, $10/$50 per 1M tokens, 1.05M context. Benchmarks vs Claude Fable 5.1, pricing math, and the AGI claim examined.

By Soufiane B.14 min read
OpenAI GPT-6 Astra benchmark charts against Claude Fable 5.1 and GPT-5.6 Sol across computer use, coding, and cybersecurity evaluations.

TL;DR

What actually shipped:

OpenAI started the staged rollout of GPT-6 Astra on September 3, 2026: 1.05M context, $10/$50 per 1M tokens, first to enterprises in the Trusted Access Program, wider API and ChatGPT access in the coming days.

Why it matters:

Astra leads on computer use (92.7% ScreenSpot-Pro), coding agents, and cybersecurity (100% ExploitBench, first model past OpenAI's Critical threshold), while trailing Fable 5.1 on aggregate intelligence indices.

The honest caveat:

Launch benchmarks are OpenAI-reported at max effort with gaps in rival coverage, access is gated so nothing here is independently hands-tested yet, and price per task is unproven.

GPT-6 Astra Is Here: OpenAI's Computer-Use Flagship, Priced Like One

In brief: GPT-6 Astra is OpenAI's flagship AI model family for end-to-end computer work, launched September 3, 2026. It scores 92.7% on ScreenSpot-Pro and a perfect 100% on ExploitBench at $10 per million input tokens, though access is still gated to enterprise trusted testers.

OpenAI fumbled the entrance and stuck the landing anyway. Users found the official GPT-6 Astra blog post sitting on the company website, watched it get pulled, then watched OpenAI share a link that did not work for hours. By the end of September 3, the post was live, the API docs were up, and company president Greg Brockman was calling Astra a "generational leap" that could one day be seen as the arrival of AGI.

Strip away the launch theater and you get something genuinely interesting: the first frontier model explicitly built to operate software rather than advise about it. Forms, spreadsheets, browsers, codebases, documents. Brockman's pitch is that Astra users "won't ever have to click around a mouse or type on a keyboard ever again." That is marketing. The benchmark table underneath it deserves a closer look, because parts of it are remarkable and parts of it are doing quiet rhetorical work. This review separates the two.

One disclosure up front, in keeping with how we do things: access is gated to Trusted Access enterprises on day one, so nothing below is hands-tested by us yet. Every number is attributed to its source, vendor-reported figures are flagged as such, and independent Artificial Analysis data gets its own section.


The Bottom Line Up Front

Question Answer
What OpenAI's GPT-6 Astra: flagship model for reasoning, coding, computer use, research, and documents
When September 3, 2026 release date; staged rollout with enterprises first, API and ChatGPT tiers in coming days
Who Agentic coding teams, automation engineers, security teams, enterprises doing document-heavy work
Why it matters Best measured computer-use and cyber scores ever, with token efficiency that changes the cost math, at roughly double Sol's promo input price, which demands scrutiny
Honest caveat Launch numbers are OpenAI-reported at max effort with rival gaps; price per task is unproven

Specs and Pricing: the Full Table

The headline figures first, all from the official API docs published at launch:

Spec GPT-6 Astra
Context window 1,050,000 tokens
Max output 128,000 tokens
Knowledge cutoff April 30, 2026
Input modalities Text, image
Output modalities Text
Reasoning effort low, medium, high, xhigh, max
Input price (per 1M) $10.00
Cached input (per 1M) $1.00
Cache writes (per 1M) $12.50
Output price (per 1M) $50.00
Fine-tuning Not supported at launch

Three pricing details matter more than the headline:

  1. The long-prompt surcharge. Prompts over 272K input tokens are billed at double input and cache rates and 1.5x output for the full request. The 1.05M context window is real, but filling it is priced to hurt. Budget your RAG pipelines accordingly, and price scenarios with our API cost calculator.
  2. Cache writes cost 1.25x. At $12.50 per million, heavy prompt-caching strategies need rechecking against the old math.
  3. No tiers (yet). Unlike GPT-5.6's Luna/Terra/Sol split, the lineup is just Astra and Astra Pro. Simpler to reason about, with fewer cheap on-ramps.

At double Sol's $5 promotional input price (about 1.67 times its $30 output price, per GPT-5.6 Sol's pricing), Astra matches Claude Fable 5.1 dollar for dollar while Meta's Muse Spark sits at a fraction of either. Sticker price is only half the story, though, which brings us to efficiency.


Computer Use: the Category Astra Was Built to Win

This is where OpenAI leaned hardest, and the numbers justify the posture:

Benchmark Astra GPT-5.6 Sol Claude Fable 5 Claude Opus 5
ScreenSpot-Pro (no tools) 92.7% 76.9% 87.3% n/a
OSWorld 2.0 72.6% 65.7% n/a 70.2%
Agents' Last Exam 59.3% 53.6% 48.7% (Fable 5.1) 55.5%
Avg minutes per OSWorld task ~40 ~75 n/a n/a

Two things stand out. First, the ScreenSpot-Pro jump (76.9% to 92.7% in one generation) is the kind of discontinuity that usually signals a real capability unlock rather than benchmark tuning. Second, OpenAI cut average task time nearly in half, which matters more than it sounds: an agent that finishes in 40 minutes instead of 75 fails less often, because every extra step is another chance to wander off.

The asterisk file is real, though. Anthropic separately claims 77.9% for Fable 5.1 on OSWorld, but on a different release of the test, so OpenAI left Fable 5.1 out of its table and the two numbers cannot be compared. Note it, do not average it. And OpenAI's Mind2Web claim (1.9x faster task completion) describes the new Codex harness plus Astra together, not the model in isolation.


Coding: Efficient, Not Quite King

For developers, the independent data is more useful than the launch table. Artificial Analysis tested Astra in Codex on day one:

  • Coding Agent Index: 67, effectively tied with Claude Opus 5, Fable 5, and Muse Spark 1.3. Fable 5.1 in Claude Code leads at 70.
  • Token efficiency is the story: Astra uses roughly one third the tokens of GPT-5.6 Sol at max effort and one fifth of Claude Opus 5 at xhigh. Every effort level sits on the cost-efficiency frontier.
  • DeepSWE v1.1 (OpenAI-reported): 74.1%, a clear step up for OpenAI's own models.
  • Terminal-Bench Science 0.1: 64.6% vs 52.6% (Fable 5.1) and 29.0% (Opus 5).

Translation: Astra is not the highest-scoring coding model, but it may be the cheapest per completed task among frontier models, at less than half the per-task cost of Fable 5 for equal scores. If your bottleneck is agent reliability per dollar rather than raw benchmark crowns, that is arguably the more important number. Our best AI coding tools breakdown will be updated once we can hands-test Astra in a real harness.


Cybersecurity: the 100% That Got Everyone's Attention

OpenAI framed this as the most significant part of the release, and for once the framing understates nothing:

Benchmark Astra Nearest rival
ExploitBench 100.0% Sol 78.5%, Opus 5 70.0%
SRE-Bench 88.0% Opus 5 12.5%

Astra is the first model to cross OpenAI's own "Critical" capability threshold under its Preparedness Framework, which is precisely why rollout starts inside an application-based cybersecurity program instead of a public API flood. A perfect ExploitBench plus an 88-to-12.5 SRE-Bench gap is not incremental; it is the kind of result that rearranges red-team roadmaps.

The honest caveat cuts both ways here. These are OpenAI-reported numbers with no Fable scores in the table at all, and "first past Critical" is OpenAI grading its own homework against its own framework. But the direction is unambiguous, and the gated rollout is the correct response to it. If you run a security team, the takeaway is not "buy Astra" (you cannot yet). It is that the exploit-generation overhang the industry warned about for 2027 may have arrived early, and defensive tooling budgets should assume so.


Long Context, Math, and Professional Work

  • MRCR v2 retrieval: 100% at 256K-512K (Sol: 91.5%), 96.3% at 512K-1M (Sol: 73.8%). Genuinely strong, though neither Claude nor Gemini numbers appear on this test, so treat it as Astra-vs-Sol only.
  • FrontierMath Tier 4 (v2): 97.6% vs 87.8% (Fable 5.1/5) and 73.2% (Opus 5). Combined with the ten Lean-certified math results from August, Astra's research-reasoning story is the best-evidenced part of the launch.
  • ARC-AGI-3: 98.6% (OpenAI, high-effort settings) vs Opus 5 at 30.2%. Independent evaluations on standard harnesses report 62.7%, so read the vendor figure as a ceiling, not a baseline. No Fable scores published. Staggering on paper, thin on comparables.
  • AutomationBench 41.4% (Sol 18.1%, Fable 5.1 31.4%) and BenchCAD 95.9% support the "built for real workplace tasks" pitch.
  • The messy middle: on the Artificial Analysis Intelligence Index v4.1.1 aggregate, Fable 5.1 posts 65.7, Astra 61.2, Sol 60.9. On HealthBench Professional, Astra leads at 63.4%, while Fable 5.1 (56.6%) scores below its own predecessor, which Anthropic attributes to safety classifiers intervening on health queries. Aggregates hide as much as they reveal; pick the benchmark that matches your workload.

The Independent Check: What Artificial Analysis Found on Day One

Vendor tables deserve an independent second opinion, and Artificial Analysis delivered one within hours:

  1. Intelligence Index: 61, equal to GPT-5.6 Sol, 5 points behind Fable 5.1, trailing the new Muse Spark 1.3. No leap on general knowledge work.
  2. ~10% fewer output tokens than Sol for equal performance, a new efficiency frontier, but Artificial Analysis calculates 75% more cost per task than its predecessor at max effort once the higher list prices are applied.
  3. Hallucinations halved: 92% to 51% on AA-Omniscience. The least flashy chart in the launch and possibly the most important one for production use.
  4. Mixed agentic knowledge work: roughly +80 points in AA-Briefcase, similar regression in GDPval-AA v2.

Net: a more efficient, more truthful model that costs more per task. Whether that trade wins depends entirely on your retry rate today, which is why "it depends" is the only honest headline until independent reproductions land.


Head-to-Head: Astra vs Fable 5.1 vs GPT-5.6 Sol

Metric GPT-6 Astra Claude Fable 5.1 GPT-5.6 Sol
Input / output per 1M $10 / $50 $10 / $50 $5 / $30 (promo)
Context window 1.05M 1M (128K max output) 1M
AA Intelligence Index 61 65.7 61
Coding Agent Index 67 70 ~62
ScreenSpot-Pro 92.7% n/a (87.3% Fable 5) 76.9%
ExploitBench 100% n/a 78.5%
Hallucination rate (AA) 51% (halved) n/a 92%
Token efficiency 1/3 of Sol Baseline Baseline
Availability Trusted Access first Generally available Generally available
Best for Computer-use agents, coding agents, security work General knowledge, chat, research Budget frontier work, speed (750 tok/s)

For the full live leaderboard across every model mentioned here, see the AI Model Intelligence Index and our AI model releases timeline. For the closest head-to-head we have hands-tested, read ChatGPT vs Claude.


The AGI Claim, Examined Without Hype

Brockman called Astra a "generational leap" and suggested it could one day be seen as AGI's arrival. Here is the sober version:

  • What supports excitement: 100K+ Stargate GPUs in pre-training, earlier models supervising training, 0% out-of-boundary behavior on impossible tasks (Sol: 48.2%), ten Lean-certified math results, near-perfect scores on several genuinely hard benchmarks.
  • What does not support AGI: benchmark scores are scores on defined tasks, not open-ended general intelligence. Fable 5.1 still beats Astra on aggregate intelligence. The model cannot be independently tested yet. And OpenAI itself is gating access behind a cybersecurity program, which is not what you do with a tame product.
  • The right mental model: the best computer-operating model ever measured, from the lab that also says it needs new safeguards to ship it. Treat it as a capability discontinuity in agency, not the arrival of a new mind.

Who Should Actually Care (and What to Do)

  • Agentic coding teams: start planning a pilot. One third the tokens of Sol with equal-or-better agent scores changes unit economics. Benchmark on your own repos before committing; harness effects (see Mind2Web) are real.
  • Security teams: assume exploit-generation capability just moved forward a year. Audit defensive tooling and disclosure pipelines now, regardless of when you get API access.
  • Enterprises in document workflows: AutomationBench and BenchCAD say this is your model class, but wait for the API tiers and test prompt-caching math against the $12.50 writes and the 272K surcharge.
  • Everyone else: Fable 5.1 and Sol remain available, cheaper per task, and independently tested. There is no reason to contort your stack around a gated model today.

Compare Every Model Side by Side

For full benchmark scores, live pricing, and context windows across every model mentioned here: Compare AI models on Renovate QR.


Explore all AI models and their live pricing in our tools directory

For more AI model reviews, hardware breakdowns, and benchmark analyses, visit our blog.

Compare these models side by side on Renovate QR

Sources & References


Published 2026-09-03. Last verified: 2026-09-03 against OpenAI API docs, the official launch benchmarks, Artificial Analysis day-one testing, and The New Stack reporting. We update this article as independent reproductions land.

Frequently Asked Questions

What is GPT-6 Astra?

GPT-6 Astra is OpenAI's flagship AI model family launched September 3, 2026, built for end-to-end work: complex reasoning, coding, computer use, research, and document creation. It offers a 1.05M token context window, 128K max output tokens, reasoning effort settings from low to max, and API pricing of $10 per million input tokens and $50 per million output tokens.

How much does GPT-6 Astra cost?

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, with cached input at $1.00 and cache writes at $12.50. Prompts over 272K tokens cost double on input and 1.5x on output. That is double Sol's $5 promotional input price (about 1.67 times its $30 output price) and matches Claude Fable 5.1, though Astra uses substantially fewer tokens per task.

Is GPT-6 Astra better than Claude Fable 5.1?

It depends on the workload. Astra leads on computer use, coding-agent efficiency, math, and cybersecurity benchmarks, while Claude Fable 5.1 leads on aggregate intelligence scores (65.7 vs 61.2 on the Artificial Analysis Intelligence Index). For autonomous computer work, Astra is the stronger pick; for general knowledge work, Fable 5.1 still holds the crown.

When can I use GPT-6 Astra?

Rollout started September 3, 2026 with enterprises in OpenAI's Trusted Access Program. Plus, Pro, Business, and Enterprise plans plus API access (including an Astra Pro tier and Zero Data Retention for eligible API customers) follow in the coming days.

Is GPT-6 Astra AGI?

No responsible reading supports that yet. Greg Brockman called it a generational leap and mused it could one day be seen as AGI's arrival, but the evidence is benchmark scores on defined tasks, not open-ended general intelligence. The honest headline is the best computer-using model ever measured, not a new species of mind.

Should developers switch to GPT-6 Astra now?

If you run computer-use agents or long agentic coding workflows, start planning a pilot: the token efficiency gains are real and independently confirmed. If you need general chat, research summaries, or the lowest cost per task, benchmark it against Fable 5.1 and GPT-5.6 Sol on your own workload first, and wait for independent reproductions of the launch numbers.

Renovate QR Newsletter

Stay sharp on AI & tech

Our best reviews, comparisons, and guides — delivered weekly. No noise, no spam.

No spam. Unsubscribe anytime.

Published

Related Articles