GPT-6 Astra costs twice as much as Claude Opus 5 per token. It cost less to run the index.
GPT-6 Astra scores 52.81 to Opus 5's 50.70 at $10 / $50 against $5 / $25, yet cost 27% less to run the full index. Where each wins, and the 272K catch.
GTM Architect & Growth Operator · Now · 16 September 2026
TL;DR · Key insights
- On Artificial Analysis's Intelligence Index v4.3, GPT-6 Astra scores 52.81 and Claude Opus 5 scores 50.70. Just over two points to Astra, which sits level with Claude Fable 5.1 at the top.
- Astra lists at exactly twice Opus 5's price: $10 / $50 per million tokens against $5 / $25. Yet running the full index cost $5,324 on Astra and $7,275 on Opus 5, because Astra wrote 60M tokens to Opus 5's 140M.
- Above 272K input tokens Astra's whole request moves to $20 / $75. Anthropic bills Opus 5's 1M window at one rate, so on long prompts Opus 5 costs a quarter of Astra on input.
- On coding agents, Codex with Astra scores 62 at $7.47 per task and Claude Code with Opus 5 scores 60 at $10.79. Opus 5 is Anthropic's recommended default; Astra is the stronger model and cheaper per task on short prompts.
On a price list, Claude Opus 5 wins every line against GPT-6 Astra. It is exactly half the price.
On a finished job the order flips, because Astra writes far fewer tokens to get there. There is one exception, and it sits at 272K input tokens.
| Info | GPT-6 Astra | Claude Opus 5 |
|---|---|---|
| Input / output per 1M | $10 / $50 | $5 / $25 |
| Cached input per 1M | $1.00 | $0.50 |
| Above 272K input tokens, in / out | $20 / $75 | $5 / $25 |
| Intelligence Index v4.3 | 52.81 | 50.70 |
| Coding Agent Index, own harness | 62 (Codex) | 60 (Claude Code) |
| Cost per coding task | $7.47 | $10.79 |
| Cost to run the full index | $5,324 | $7,275 |
| Tokens generated on the index | 60M | 140M |
| Output speed, tokens/s | about 53 | about 50 |
| Context / max output | 1,050,000 / 128,000 | 1M / 128K |
Prices and long-context rules from OpenAI's GPT-6 Astra model page and Anthropic's pricing page. Index to two decimals from Artificial Analysis's v4.3 chart data; run cost, tokens and speed from its model pages; coding scores and cost per task from its Claude Code vs Codex comparison. Read on 16 September 2026.
Two points apart on the independent index
Artificial Analysis scores Astra at 52.81 and Opus 5 at 50.70 on Intelligence Index v4.3. Just over two points separate them, and that puts Astra level with Claude Fable 5.1 at the top while Opus 5 sits a step below both.
The coding agent numbers point the same way. Codex running Astra scores 62 on Artificial Analysis’s Coding Agent Index, and Claude Code running Opus 5 scores 60.
OpenAI’s launch post showed far bigger gaps on agentic work: 99.9% against Opus 5’s 30.2% on ARC-AGI-3, and 88% against 12.5% on SRE-Bench. Those are OpenAI’s measurements of a competitor. The same table had Opus 5 ahead on Humanity’s Last Exam with tools, 63.6% to 57.2%.
Twice the price per token, cheaper per task
Every rate on Astra’s price list is exactly double Opus 5’s. Input is $10 against $5, output $50 against $25, cached input $1 against $0.50.
The bill does not follow the price list. Running the full index cost Artificial Analysis $5,324 on Astra and $7,275 on Opus 5, 27% less for the dearer model, because Astra wrote 60M tokens to Opus 5’s 140M. Coding agents show a similar gap: $7.47 per task for Codex with Astra and $10.79 for Claude Code with Opus 5.
Artificial Analysis says every reasoning effort of Astra has the lowest cost per task at its level of intelligence. Against Opus 5 that shows up plainly. Astra costs more per token and less per job.
Per token, Opus 5 is half the price. Per finished task on short prompts, Astra came out about 30% cheaper in both of Artificial Analysis’s measurements. Check which of those two numbers your invoice follows.
The 272K line flips it back
OpenAI prices any Astra prompt above 272K input tokens at 2x input and cache rates and 1.5x output, for the whole request. That is $20 input, $2 cached and $75 output. Anthropic bills Opus 5’s full 1M window at one rate.
Past that line Opus 5 costs a quarter of Astra on input and a third on output, and Astra’s token savings would have to be very large to close that gap. For long documents, large codebases held in context and any workload that routinely crosses 272K, Opus 5 is the cheaper model. The arithmetic is in the 272K pricing cliff.
Speed is close
Artificial Analysis measures Astra at about 53 output tokens per second and Opus 5 at about 50, both slower than average among the models it tracks. Astra still tends to finish first, because it has fewer tokens to write.
What Opus 5 still has
- Anthropic’s default. Anthropic’s models page says to start with Opus 5 when you are unsure, and to move up to Fable 5.1 only when that falls short.
- Long prompts, where the 272K rule makes it several times cheaper.
- The Claude stack. Claude Code, Anthropic’s platform and any agreements that already cover Anthropic.
- Readable reasoning. OpenAI itself says Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s. If your oversight depends on reading a model’s reasoning, that counts against Astra.
Which one I would use
| The work | Model | Why |
|---|---|---|
| Long agentic runs, prompts under 272K | GPT-6 Astra | Higher scores and about 30% cheaper per finished task in both independent measurements. |
| Prompts that regularly cross 272K tokens | Claude Opus 5 | One rate across 1M tokens, while Astra's whole request doubles on input. |
| Team already working in Claude Code | Claude Opus 5 | Anthropic's recommended default, in the harness you already run. |
| Oversight that relies on reading reasoning | Claude Opus 5 | OpenAI says Astra's written reasoning got harder to monitor. |
| Choosing a default for a new build | Test both | Two points apart. Prompt length and harness decide it. |
Astra wins short-prompt agentic work on score and cost per task. Opus 5 wins long prompts and Claude-first teams.
For the model Astra ties at the top, see GPT-6 Astra against Claude Fable 5.1. For the step above Opus 5 inside Anthropic’s lineup, Opus 5 against Fable 5.1, and for the harnesses, Claude Code against Codex.
What would change my mind
A cost-per-task comparison in the same harness. Astra’s coding score was measured in Codex and Opus 5’s in Claude Code, and the harness matters as much as the model.
An Opus 5 point release. A two-point index gap is the size one release can close.
A change to OpenAI’s 272K rule. It is the single line that makes Opus 5 the cheaper model on long prompts.
Questions people asked
Is GPT-6 Astra better than Claude Opus 5?
On independent benchmarks, modestly. Artificial Analysis’s Intelligence Index v4.3 scores GPT-6 Astra at 52.81 and Claude Opus 5 at 50.70, and on its Coding Agent Index Codex with Astra scores 62 against Claude Code with Opus 5 at 60. OpenAI’s own launch numbers show much wider gaps on agentic tasks, such as 99.9% against 30.2% on ARC-AGI-3, while Opus 5 led Humanity’s Last Exam with tools, 63.6% to 57.2%. Treat the vendor figures as claims; the independent gap is about two points.
Is GPT-6 Astra more expensive than Claude Opus 5?
Per token, yes, exactly twice: $10 per million input tokens and $50 output against Opus 5’s $5 and $25, and $1 against $0.50 for cached input. Per task it is usually cheaper. Artificial Analysis spent $5,324 running its index on Astra and $7,275 on Opus 5, and measured Codex with Astra at $7.47 per coding task against $10.79 for Claude Code with Opus 5. Above 272K input tokens that reverses: Astra’s whole request moves to $20 input and $75 output while Opus 5 keeps one rate.
Which is faster, GPT-6 Astra or Claude Opus 5?
They generate at similar speeds: Artificial Analysis measures about 53 output tokens per second for Astra and about 50 for Opus 5, both slower than average. Astra tends to finish sooner because it writes far fewer tokens, 60M across the index against Opus 5’s 140M.
Should I use GPT-6 Astra or Claude Opus 5?
GPT-6 Astra for long agentic runs with prompts under 272K tokens, where it scores higher and costs less per finished task. Claude Opus 5 for prompts that cross 272K tokens, for teams whose work already runs in Claude Code or on Anthropic’s platform, and as the default Anthropic itself recommends.
If you are choosing between OpenAI and Anthropic for an agentic product, the contact page is the fastest route to me.