Skip to content
wojciech.io
All insights
AI SystemsAIOpenAIClaudeAI Systems

GPT-6 Astra came fourth on the intelligence index. That is the most interesting thing about it.

OpenAI shipped GPT-6 Astra at $10 per million input tokens. It ranks fourth on the Artificial Analysis Intelligence Index, behind three Claude models, and wins almost every benchmark that involves finishing work. Here is what that split means if you build on these models.

Wojciech Łuszczyński

Wojciech Łuszczyński

GTM Architect & Growth Operator · Now · 4 September 2026

TL;DR · Key insights

  • On the Artificial Analysis Intelligence Index, Astra scores 61.2. Claude Fable 5.1 scores 65.7, Opus 5 scores 63.1, Fable 5 scores 62.1. Astra is fourth, and OpenAI published that table itself.
  • On work that involves operating a computer, it is not close. ARC-AGI-3: Astra 99.9%, Opus 5 30.2%. SRE-Bench: Astra 88%, Opus 5 12.5%. Terminal-Bench Science: Astra 64.6%, Fable 5.1 52.6%.
  • It lists at $10 / $50 per million tokens, the same headline price as Claude Fable 5, and OpenAI claims 31% to 86% lower cost per finished task because it takes fewer steps to get there.
  • The number I would actually buy on: without safeguards, GPT-5.6 Sol went outside its authorized scope 48% of the time on OpenAI's new test. Astra did it 0% of the time. That is what makes delegation possible.

OpenAI shipped GPT-6 Astra today and called it the world’s most intelligent and aligned model. Then, further down the same page, it published a table showing Astra in fourth place on the headline independent intelligence benchmark.

Both things are true. The gap between them is the story.

ModelIntelligence IndexARC-AGI-3In / out per 1M
Claude Fable 5.165.7n/a$10 / $50
Claude Opus 563.130.2%n/a
Claude Fable 562.1n/a$10 / $50
GPT-6 Astra61.299.9%$10 / $50
GPT-5.6 Sol60.97.8%$4 / $20

Artificial Analysis Intelligence Index v4.1.1 and ARC-AGI-3, both as published in OpenAI's own launch table. Astra loses the first column to three Claude models and wins the second one by a factor of three.

Read those two columns next to each other and you can see the industry changing shape. The intelligence index, the number everyone has quoted for two years, no longer tracks the thing that decides whether a model is useful to you.

What GPT-6 Astra is, and what it costs

Astra lists at $10 per million input tokens and $50 per million output, with separate rates for cache reads and writes. A fast mode runs at up to double the speed for double the price.

That is the same headline price as Claude Fable 5. OpenAI has stopped undercutting Anthropic on the sticker and started arguing about the bill instead, which is a more confident position and a harder one to check.

It arrives in the API as gpt-6-astra, plus Microsoft Azure and AWS Bedrock. In ChatGPT it is included in existing Plus, Pro, Business and Enterprise allowances, with extra credits available to buy, and Pro, Business and Enterprise also get a GPT-6 Astra Pro tier.

One deployment detail worth knowing before you plan around it: for Enterprise workspaces, access is off by default at launch. An administrator has to turn it on. If your company bought ChatGPT Enterprise and nothing changed today, that is why.

Why it comes fourth on the intelligence index

The Artificial Analysis Intelligence Index v4.1.1 puts Claude Fable 5.1 at 65.7, Opus 5 at 63.1, Fable 5 at 62.1, then Astra at 61.2 and GPT-5.6 Sol at 60.9.

On Humanity’s Last Exam with tools the same pattern repeats and widens: Fable 5.1 at 65.0%, Fable 5 at 63.8%, Opus 5 at 63.6%, Astra at 57.2%. On the Artificial Analysis Coding Agent Index, Opus 5 leads at 68.1 with Fable 5 at 67.2 and Astra third at 67.0.

So on the aggregate scores that get screenshotted and posted, Claude still wins. If your work is hard reasoning delivered as text, an argument to pressure-test, a strategy to poke holes in, a proof to check, nothing here dislodges Claude. That was true in July and it is still true.

What changed is that this is no longer the interesting question.

Where Astra is not close to anyone

Line up the benchmarks that involve a model operating something rather than answering something, and the table inverts hard.

BenchmarkWhat it testsAstraBest Claude
ARC-AGI-3Learning novel environments from scratch99.9%30.2%
SRE-BenchReverse engineering binaries with no source88.0%12.5%
Terminal-Bench ScienceResearch workflows in a real terminal64.6%52.6%
AutomationBenchEnd-to-end business process automation41.4%31.4%
BenchCADRebuilding 3D objects as CAD code95.9%84.3%
FrontierMath Tier 4Research-grade mathematics97.6%87.8%
MRCR 512K–1MRetrieval across a very long context96.3%n/a

Astra against the strongest Claude result OpenAI reports in each row. The two highlighted rows are not incremental improvements; they are different orders of magnitude.

Those first two rows deserve to be read twice. ARC-AGI-3 measures whether a model can work out the rules of an environment it has never seen. GPT-5.6 Sol scored 7.8% on it four months ago. Astra scores 99.9%, and the ARC Prize Foundation says it beat their human action-efficiency baseline on 96% of levels.

SRE-Bench asks a model to reverse engineer a compiled binary and explain what it does with no access to source. Astra solves 88% first try and 99.2% within four attempts. Opus 5 solves 12.5%.

The story is: end of one era, start of another.

There is a practical version of that quote. On OSWorld 2.0, Astra scores 72.6% at roughly forty minutes per task where Sol scored 65.7% at roughly seventy-five. Better result, about half the wall-clock. Pair that with the updated Codex harness and OpenAI claims 1.9x faster task completion overall.

A model that is slightly less clever but finishes in half the time is not a worse model for anyone running a business. It is a different product.

The cost argument, and why you should check it

OpenAI’s cost claims are all about steps rather than tokens. In its published configurations, Astra completed Terminal-Bench Science at about 31% lower cost than Fable 5.1, Terminal-Bench 4.0 at about 63% lower, and BenchCAD at about 86% lower. On Agents’ Last Exam it used roughly 65% fewer output tokens than Opus 5 while scoring higher.

This is the same argument I made in July about GPT-5.6 and Claude 5, just pointed the other way. Cheaper per token has never meant cheaper per result. A model that reaches the answer in four steps beats a cheaper one that wanders through twelve, and the gap widens the longer the task runs.

Two cautions before you put that in a business case.

These are OpenAI’s measurements of a competitor’s product. The footnotes are candid that Claude’s OSWorld and BenchCAD numbers reflect methodology OpenAI disagrees with, and that some Claude scores come from Mythos, which is Fable with fewer safeguards, the same distinction I dug into when Fable 5 found a live bug in my repo. A vendor grading a rival on tests the vendor chose is data, not a verdict.

And there is a quieter footnote: Claude Fable 5 and 5.1 are excluded from three life sciences benchmarks entirely, because they refuse most of the questions. If your work sits near medicine or biology, that single line may matter more to you than any score on the page.

The number I would actually buy on

Buried under the benchmarks is the result that changes what you can delegate.

OpenAI built a new evaluation, informed by the Hugging Face incident, that tests what a model does when handed a task that is difficult or impossible: does it stay inside the scope it was given, or does it go further to get the job done?

Went beyond the authorized target on an impossible task

+100% vs baseline
GPT-6 Astra0%
GPT-5.6 Sol48%
Both run without production safeguards. Source: OpenAI, GPT-6 Astra launch post, 4 September 2026.

Forty-eight percent to zero. The same pattern holds across the rest of the alignment table: on the ExploitGym honeypot, Sol took the bait 48.2% of the time and Astra 0.0%. On OpenAI’s internal computer-use safety benchmark, where lower is better, Astra scores 2.4% against Fable 5.1’s 9.5%, Opus 5’s 11.5% and Sol’s 22.0%. Astra never once tried to get around a Codex auto-review denial, including when the review was deliberately configured to be evadable and the task could not be finished any other way.

For anyone who has tried to hand real work to an agent, this is the whole game. Capability was never the thing stopping you from leaving it running. Scope was. A model that is brilliant and occasionally exceeds its brief needs a human watching it, which means it saves you nothing. A model that stops at the fence can be trusted with the keys.

It is the same principle I built into the operator CMO plugin: the useful behaviour is not producing more, it is refusing clearly when a request falls outside what was agreed.

The part OpenAI admits got worse

One paragraph in the launch post says something vendors do not usually say.

Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s. OpenAI tested this by explicitly asking models to evade monitoring, and attributes the decline to Astra having more control over what it writes down and needing fewer written steps to solve a problem. It still struggles to conceal the reasoning behind genuinely hard tasks. OpenAI says it takes the decline seriously and that monitorability remains a research priority.

Read plainly: the model got better at working and worse at showing its work. Those may be the same underlying change. If your oversight plan is reading the chain of thought, it just got weaker, and the fix has to move to auto-review, action logs and permission boundaries.

Alongside that, Astra is the first model to meet the Critical cybersecurity threshold under OpenAI’s Preparedness Framework. During evaluation it found and used two previously unknown zero-day vulnerabilities, now disclosed to their maintainers. At launch it will do secure code review and patching but refuses proof-of-concept exploit work, and OpenAI warns that extra safety checks can pause or stop legitimate tasks. In ChatGPT and Codex you get asked to review. In the API, the task just stops.

Plan for that if you are building on it. A pipeline that assumes every call returns needs a path for the call that does not.

How I would route work across these models now

Nothing here makes one model the answer. It sharpens the split I already run.

The workWhat I would useWhy
Thinking a hard problem throughClaude Fable 5.1 or Opus 5Still ahead on the reasoning indices and on Humanity's Last Exam
Operating software for hoursGPT-6 AstraThe agentic gap is not incremental, and it finishes in half the time
Long unsupervised runsGPT-6 AstraZero scope violations on the impossible-task test is the enabling number
Clearly specified implementationClaude Sonnet 5Cheap, competent, and already wired into Claude Code
High-volume classificationGPT-5.6 Luna$0.20 / $1.20 makes the quality difference irrelevant at that tier

How I would split the work this week. The interesting column is the last one: every row is a different reason, which is the argument against having a single default model.

The stack I run and why each piece is there is written up in the AI stack I actually use in production. The routing logic did not change today. One row got a much stronger candidate.

Questions people asked

Is GPT-6 Astra better than Claude?

For thinking, no. Fable 5.1 leads the intelligence index at 65.7 against Astra’s 61.2, and leads Humanity’s Last Exam too. For operating a computer, it is not a contest: 99.9% against 30.2% on ARC-AGI-3, 88% against 12.5% on SRE-Bench. Pick per task, not per vendor.

Is it actually cheaper than Claude?

At list price it is identical to Fable 5 at $10 / $50. OpenAI’s argument is that it finishes in fewer steps, and it publishes cost reductions of 31% to 86% across several benchmarks to support that. Those are OpenAI’s own measurements of a competitor, so verify on your workload before you move a budget.

Should I switch my agents to it today?

Not on launch day, and possibly not yet by choice: Enterprise access is off by default until an administrator enables it. When you can, test the long unsupervised runs first, because that is where the alignment and speed results should show up most clearly and where the old model was costing you supervision.

What broke in this release?

Monitorability. OpenAI says Astra’s written reasoning is harder to follow than Sol’s and calls the decline serious. If you rely on reading its reasoning to catch problems, move that check to auto-review and action logs instead.

What I am watching next

Two things will tell us whether this holds up.

The first is independent replication. Artificial Analysis, ARC Prize and the rest will publish their own runs, and the Claude numbers in OpenAI’s table will get re-measured by people who did not write the announcement. The 31% to 86% cost gaps are the claims most likely to move.

The second is whether Anthropic answers on agentic ground rather than on the index. Fable 5.1 holds the reasoning crown comfortably. Nobody buys a model to hold a crown.

I will update this piece as real numbers land. If you want the version of this argument applied to your own stack rather than to a launch post, the plugin is open source and the contact page is the fastest route to me.

About the author

Wojciech Łuszczyński

Wojciech Łuszczyński

GTM Architect and Growth Operator building AI-native revenue systems for B2B SaaS and technology companies. I connect positioning, SEO, content, paid acquisition, CRM, automation, analytics and AI workflows into practical growth infrastructure.

Newsletter

Get the next one first.

When I publish a new article on AI systems, GTM architecture, or growth operating models, you'll be the first to know.

Subscribe