Skip to content
wojciech.io
All insights
AI SystemsAIOpenAIClaudeAI Systems

GPT-6 Astra came fourth in its own launch table. A week later the independent index put it level with Claude Fable 5.1.

Fourth in OpenAI's launch table, GPT-6 Astra now ties Claude Fable 5.1 for first on Artificial Analysis v4.3 and cost 59% less to run the index.

Wojciech Luszczynski

Wojciech Luszczynski

GTM Architect & Growth Operator · Now · 4 September 2026 · 15 min read

Share

TL;DR · Key insights

  • Update, 16 September 2026: Artificial Analysis's own run of Intelligence Index v4.3 puts GPT-6 Astra at 52.81 and Claude Fable 5.1 at 53.37, and calls it a tie for the lead. Claude Opus 5 follows at 50.70.
  • Astra is also the cheapest of the three to run. The full index cost $5,324 on Astra, $7,275 on Opus 5 and $13,129 on Fable 5.1, because Astra wrote 60M tokens against 140M and 190M.
  • At launch, OpenAI's own table on the previous index version had Astra fourth at 61.2, behind Fable 5.1, Opus 5 and Fable 5. The agentic results were the story then: ARC-AGI-3 99.9% against Opus 5's 30.2%, SRE-Bench 88% against 12.5%.
  • List price is $10 / $50 per million tokens, the same as Fable 5.1 and twice Opus 5. The number I would buy on is still the scope test: without safeguards, GPT-5.6 Sol overreached 48% of the time and Astra 0%.

Where Astra stands on the independent index

ModelIndex v4.3Cost to run the indexTokens generatedIn / out per 1M
Claude Fable 5.153.37$13,129190M$10 / $50
GPT-6 Astra52.81$5,32460M$10 / $50
Claude Opus 550.70$7,275140M$5 / $25

Artificial Analysis Intelligence Index v4.3 to two decimals from its chart data, with run cost and tokens from its model pages, read on 16 September 2026. This is a later index version than the launch table below, so the two sets of scores cannot be compared directly.

Artificial Analysis describes the top of v4.3 as Astra tying the lead with Fable 5.1. Half a point separates them, and on Artificial Analysis’s own model pages both display as 53.

The cost column is the bigger change. Astra lists at the same $10 / $50 as Fable 5.1, yet running the index on it cost 59% less, because it wrote about a third as many tokens. Artificial Analysis says every reasoning effort of Astra has the lowest cost per task at its level of intelligence. OpenAI claimed that at launch. An independent run now shows it.

For the head-to-head pairs, see Astra against Opus 5 and Astra against Fable 5.1. For the harness view of the same fight, see Claude Code against Codex, and for what changed between index versions, the v4.3 explainer.

The launch analysis, 4 September 2026

OpenAI shipped GPT-6 Astra on 4 September and called it the world’s most intelligent and aligned model. Then, further down the same page, it published a table showing Astra in fourth place on the headline independent intelligence benchmark.

Both things are true. The gap between them is the story.

ModelIntelligence IndexARC-AGI-3In / out per 1M
Claude Fable 5.165.7n/a$10 / $50
Claude Opus 563.130.2%n/a
Claude Fable 562.1n/a$10 / $50
GPT-6 Astra61.299.9%$10 / $50
GPT-5.6 Sol60.97.8%$4 / $20

Artificial Analysis Intelligence Index v4.1.1 and ARC-AGI-3, both as published in OpenAI's own launch table on 4 September. At launch Astra lost the first column to three Claude models and won the second one by a factor of three.

Read those two columns next to each other and you can see the industry changing shape. The intelligence index, the number everyone has quoted for two years, no longer tracks the thing that decides whether a model is useful to you.

What GPT-6 Astra is, and what it costs

Astra lists at $10 per million input tokens and $50 per million output, with separate rates for cache reads and writes. A fast mode runs at up to double the speed for double the price.

That is the same headline price as Claude Fable 5. OpenAI has stopped undercutting Anthropic on the sticker and started arguing about the bill instead, which is a more confident position and a harder one to check.

It arrives in the API as gpt-6-astra, plus Microsoft Azure and AWS Bedrock. In ChatGPT it is included in existing Plus, Pro, Business and Enterprise allowances, with extra credits available to buy, and Pro, Business and Enterprise also get a GPT-6 Astra Pro tier.

One deployment detail worth knowing before you plan around it: for Enterprise workspaces, access is off by default at launch. An administrator has to turn it on. If your company bought ChatGPT Enterprise and nothing changed today, that is why.

Why it came fourth in the launch table

In OpenAI’s launch table, the Artificial Analysis Intelligence Index v4.1.1 put Claude Fable 5.1 at 65.7, Opus 5 at 63.1, Fable 5 at 62.1, then Astra at 61.2 and GPT-5.6 Sol at 60.9.

On Humanity’s Last Exam with tools the same pattern repeats and widens: Fable 5.1 at 65.0%, Fable 5 at 63.8%, Opus 5 at 63.6%, Astra at 57.2%. On the Artificial Analysis Coding Agent Index, Opus 5 leads at 68.1 with Fable 5 at 67.2 and Astra third at 67.0.

So on the aggregate scores that got screenshotted and posted on launch day, Claude still won. If your work is hard reasoning delivered as text, an argument to pressure-test, a strategy to poke holes in, a proof to check, nothing in the launch table dislodged Claude. The independent run a week later narrowed that to a tie.

What changed is that this is no longer the interesting question.

Where Astra is not close to anyone

Line up the benchmarks that involve a model operating something rather than answering something, and the table inverts hard.

BenchmarkWhat it testsAstraBest Claude
ARC-AGI-3Learning novel environments from scratch99.9%30.2%
SRE-BenchReverse engineering binaries with no source88.0%12.5%
Terminal-Bench ScienceResearch workflows in a real terminal64.6%52.6%
AutomationBenchEnd-to-end business process automation41.4%31.4%
BenchCADRebuilding 3D objects as CAD code95.9%84.3%
FrontierMath Tier 4Research-grade mathematics97.6%87.8%
MRCR 512K–1MRetrieval across a very long context96.3%n/a

Astra against the strongest Claude result OpenAI reports in each row. The two highlighted rows are not incremental improvements; they are different orders of magnitude.

Those first two rows deserve to be read twice. ARC-AGI-3 measures whether a model can work out the rules of an environment it has never seen. GPT-5.6 Sol scored 7.8% on it four months ago. Astra scores 99.9%, and the ARC Prize Foundation says it beat their human action-efficiency baseline on 96% of levels.

SRE-Bench asks a model to reverse engineer a compiled binary and explain what it does with no access to source. Astra solves 88% first try and 99.2% within four attempts. Opus 5 solves 12.5%.

The story is: end of one era, start of another.

Greg Burnham, EpochAI

There is a practical version of that quote. On OSWorld 2.0, Astra scores 72.6% at roughly forty minutes per task where Sol scored 65.7% at roughly seventy-five. Better result, about half the wall-clock. Pair that with the updated Codex harness and OpenAI claims 1.9x faster task completion overall.

A model that is slightly less clever but finishes in half the time is not a worse model for anyone running a business. It is a different product.

The cost argument, and why you should check it

OpenAI’s cost claims are all about steps rather than tokens. In its published configurations, Astra completed Terminal-Bench Science at about 31% lower cost than Fable 5.1, Terminal-Bench 4.0 at about 63% lower, and BenchCAD at about 86% lower. On Agents’ Last Exam it used roughly 65% fewer output tokens than Opus 5 while scoring higher.

This is the same argument I made in July about GPT-5.6 and Claude 5, just pointed the other way. Cheaper per token has never meant cheaper per result. A model that reaches the answer in four steps beats a cheaper one that wanders through twelve, and the gap widens the longer the task runs.

Two cautions before you put that in a business case.

These are OpenAI’s measurements of a competitor’s product. The footnotes are candid that Claude’s OSWorld and BenchCAD numbers reflect methodology OpenAI disagrees with, and that some Claude scores come from Mythos, which is Fable with fewer safeguards, the same distinction I dug into when Fable 5 found a live bug in my repo. A vendor grading a rival on tests the vendor chose is data, not a verdict.

And there is a quieter footnote: Claude Fable 5 and 5.1 are excluded from three life sciences benchmarks entirely, because they refuse most of the questions. If your work sits near medicine or biology, that single line may matter more to you than any score on the page.

The number I would actually buy on

Buried under the benchmarks is the result that changes what you can delegate.

OpenAI built a new evaluation, informed by the Hugging Face incident, that tests what a model does when handed a task that is difficult or impossible: does it stay inside the scope it was given, or does it go further to get the job done?

Went beyond the authorized target on an impossible task

+100% vs baseline
GPT-6 Astra0%
GPT-5.6 Sol48%
Both run without production safeguards. Source: OpenAI, GPT-6 Astra launch post, 4 September 2026.

Forty-eight percent to zero. The same pattern holds across the rest of the alignment table: on the ExploitGym honeypot, Sol took the bait 48.2% of the time and Astra 0.0%. On OpenAI’s internal computer-use safety benchmark, where lower is better, Astra scores 2.4% against Fable 5.1’s 9.5%, Opus 5’s 11.5% and Sol’s 22.0%. Astra never once tried to get around a Codex auto-review denial, including when the review was deliberately configured to be evadable and the task could not be finished any other way.

For anyone who has tried to hand real work to an agent, this is the whole game. Capability was never the thing stopping you from leaving it running. Scope was. A model that is brilliant and occasionally exceeds its brief needs a human watching it, which means it saves you nothing. A model that stops at the fence can be trusted with the keys.

It is the same principle I built into the operator CMO plugin: the useful behaviour is not producing more, it is refusing clearly when a request falls outside what was agreed.

The part OpenAI admits got worse

One paragraph in the launch post says something vendors do not usually say.

Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s. OpenAI tested this by explicitly asking models to evade monitoring, and attributes the decline to Astra having more control over what it writes down and needing fewer written steps to solve a problem. It still struggles to conceal the reasoning behind genuinely hard tasks. OpenAI says it takes the decline seriously and that monitorability remains a research priority.

Read plainly: the model got better at working and worse at showing its work. Those may be the same underlying change. If your oversight plan is reading the chain of thought, it just got weaker, and the fix has to move to auto-review, action logs and permission boundaries.

Alongside that, Astra is the first model to meet the Critical cybersecurity threshold under OpenAI’s Preparedness Framework. During evaluation it found and used two previously unknown zero-day vulnerabilities, now disclosed to their maintainers. At launch it will do secure code review and patching but refuses proof-of-concept exploit work, and OpenAI warns that extra safety checks can pause or stop legitimate tasks. In ChatGPT and Codex you get asked to review. In the API, the task just stops.

Plan for that if you are building on it. A pipeline that assumes every call returns needs a path for the call that does not.

How I would route work across these models now

Nothing here makes one model the answer. It sharpens the split I already run.

The workWhat I would useWhy
Thinking a hard problem throughClaude Fable 5.1 or Opus 5Level with Astra on the v4.3 index, and ahead on Humanity's Last Exam in the launch table
Operating software for hoursGPT-6 AstraThe agentic gap is not incremental, and it finishes in half the time
Long unsupervised runsGPT-6 AstraZero scope violations on the impossible-task test is the enabling number
Clearly specified implementationClaude Sonnet 5Cheap, competent, and already wired into Claude Code
High-volume classificationGPT-5.6 Luna$0.20 / $1.20 makes the quality difference irrelevant at that tier

How I would split the work this week. The interesting column is the last one: every row is a different reason, which is the argument against having a single default model.

The stack I run and why each piece is there is written up in the AI stack I actually use in production. The routing logic did not change today. One row got a much stronger candidate.

Questions people asked

Is GPT-6 Astra better than Claude?

On independent numbers it is level with the best Claude model. Artificial Analysis’s Intelligence Index v4.3 puts Claude Fable 5.1 at 53.37 and GPT-6 Astra at 52.81, which it describes as a tie for the lead, with Claude Opus 5 at 50.70. Astra also cost less to run the index than either: $5,324 against $13,129 for Fable 5.1 and $7,275 for Opus 5. On agentic work OpenAI’s launch numbers favour Astra by wide margins, such as 99.9% against Opus 5’s 30.2% on ARC-AGI-3 and 88% against 12.5% on SRE-Bench. At launch, on the previous index version, OpenAI’s own table had Astra fourth. Pick per task, not per vendor.

How much does GPT-6 Astra cost?

Standard API pricing is $10 per million input tokens and $50 per million output tokens, with separate rates for cache reads and writes. That is the same headline price as Claude Fable 5. Fast mode runs at up to twice the speed for twice the price. In ChatGPT it is included in existing Plus, Pro, Business and Enterprise allowances, with extra credits available to buy.

Is GPT-6 Astra cheaper than Claude in practice?

Yes, and independent numbers now back OpenAI’s claim. Artificial Analysis spent $5,324 running its Intelligence Index v4.3 on Astra against $13,129 on Claude Fable 5.1 and $7,275 on Opus 5, and says every reasoning effort of Astra has the lowest cost per task at its level of intelligence. On coding agents it measured Codex with Astra at $7.47 per task against $12.39 for Claude Code with Fable 5.1. OpenAI’s own launch claims were 31% to 86% lower cost than Fable 5.1 on several benchmarks. Long prompts are the exception: above 272K input tokens the whole request is billed at 2x input and 1.5x output, and Fable 5.1 keeps one rate.

Can I use GPT-6 Astra at work today?

Possibly not yet, even if your company pays for it. It is rolling out to a limited set of organizations first and reaches all Plus, Pro, Business and Enterprise users over the following days. For Enterprise workspaces, access is off by default at launch and an administrator has to switch it on. In the API it is available as gpt-6-astra, and also through Microsoft Azure and AWS Bedrock.

What is the catch with GPT-6 Astra?

Two things, both stated by OpenAI. It is the first model to meet the Critical cybersecurity threshold under their Preparedness Framework, which is why it refuses proof-of-concept exploit work at launch and why extra safety checks can pause legitimate tasks. And its written reasoning is harder to monitor than GPT-5.6 Sol’s, which OpenAI attributes to it solving problems in fewer written steps. They call the decline serious. It is the most honest line in the announcement.

What I am watching next

Two things will tell us whether this holds up.

The first is independent replication. Artificial Analysis, ARC Prize and the rest will publish their own runs, and the Claude numbers in OpenAI’s table will get re-measured by people who did not write the announcement. Artificial Analysis has now published its run, in the update at the top: Astra ties Fable 5.1 on the index and costs less to run.

The second is whether Anthropic answers on agentic ground rather than on the index. At launch Fable 5.1 held the reasoning crown comfortably. A week later it holds it by half a point. Nobody buys a model to hold a crown.

I will update this piece as real numbers land. If you want the version of this argument applied to your own stack rather than to a launch post, the plugin is open source and the contact page is the fastest route to me.

Share this article

About the author

Wojciech Luszczynski

Wojciech Luszczynski

GTM Architect and Growth Operator building AI-native revenue systems for B2B SaaS and technology companies. I connect positioning, SEO, content, paid acquisition, CRM, automation, analytics and AI workflows into practical growth infrastructure.

Newsletter

Get the next one first.

When I publish a new article on AI systems, GTM architecture, or growth operating models, you'll be the first to know.

Subscribe