Skip to content
wojciech.io
All insights
AI SystemsAIOpenAIClaudeAI Systems

GPT-6 Astra costs twice as much as Claude Opus 5 per token. It cost less to run the index.

GPT-6 Astra scores 52.81 to Opus 5's 50.70 at $10 / $50 against $5 / $25, yet cost 27% less to run the full index. Where each wins, and the 272K catch.

Wojciech Łuszczyński

Wojciech Łuszczyński

GTM Architect & Growth Operator · Now · 16 September 2026

TL;DR · Key insights

  • On Artificial Analysis's Intelligence Index v4.3, GPT-6 Astra scores 52.81 and Claude Opus 5 scores 50.70. Just over two points to Astra, which sits level with Claude Fable 5.1 at the top.
  • Astra lists at exactly twice Opus 5's price: $10 / $50 per million tokens against $5 / $25. Yet running the full index cost $5,324 on Astra and $7,275 on Opus 5, because Astra wrote 60M tokens to Opus 5's 140M.
  • Above 272K input tokens Astra's whole request moves to $20 / $75. Anthropic bills Opus 5's 1M window at one rate, so on long prompts Opus 5 costs a quarter of Astra on input.
  • On coding agents, Codex with Astra scores 62 at $7.47 per task and Claude Code with Opus 5 scores 60 at $10.79. Opus 5 is Anthropic's recommended default; Astra is the stronger model and cheaper per task on short prompts.

On a price list, Claude Opus 5 wins every line against GPT-6 Astra. It is exactly half the price.

On a finished job the order flips, because Astra writes far fewer tokens to get there. There is one exception, and it sits at 272K input tokens.

InfoGPT-6 AstraClaude Opus 5
Input / output per 1M$10 / $50$5 / $25
Cached input per 1M$1.00$0.50
Above 272K input tokens, in / out$20 / $75$5 / $25
Intelligence Index v4.352.8150.70
Coding Agent Index, own harness62 (Codex)60 (Claude Code)
Cost per coding task$7.47$10.79
Cost to run the full index$5,324$7,275
Tokens generated on the index60M140M
Output speed, tokens/sabout 53about 50
Context / max output1,050,000 / 128,0001M / 128K

Prices and long-context rules from OpenAI's GPT-6 Astra model page and Anthropic's pricing page. Index to two decimals from Artificial Analysis's v4.3 chart data; run cost, tokens and speed from its model pages; coding scores and cost per task from its Claude Code vs Codex comparison. Read on 16 September 2026.

Two points apart on the independent index

Artificial Analysis scores Astra at 52.81 and Opus 5 at 50.70 on Intelligence Index v4.3. Just over two points separate them, and that puts Astra level with Claude Fable 5.1 at the top while Opus 5 sits a step below both.

The coding agent numbers point the same way. Codex running Astra scores 62 on Artificial Analysis’s Coding Agent Index, and Claude Code running Opus 5 scores 60.

OpenAI’s launch post showed far bigger gaps on agentic work: 99.9% against Opus 5’s 30.2% on ARC-AGI-3, and 88% against 12.5% on SRE-Bench. Those are OpenAI’s measurements of a competitor. The same table had Opus 5 ahead on Humanity’s Last Exam with tools, 63.6% to 57.2%.

Twice the price per token, cheaper per task

Every rate on Astra’s price list is exactly double Opus 5’s. Input is $10 against $5, output $50 against $25, cached input $1 against $0.50.

The bill does not follow the price list. Running the full index cost Artificial Analysis $5,324 on Astra and $7,275 on Opus 5, 27% less for the dearer model, because Astra wrote 60M tokens to Opus 5’s 140M. Coding agents show a similar gap: $7.47 per task for Codex with Astra and $10.79 for Claude Code with Opus 5.

Artificial Analysis says every reasoning effort of Astra has the lowest cost per task at its level of intelligence. Against Opus 5 that shows up plainly. Astra costs more per token and less per job.

Key takeaway

Per token, Opus 5 is half the price. Per finished task on short prompts, Astra came out about 30% cheaper in both of Artificial Analysis’s measurements. Check which of those two numbers your invoice follows.

The 272K line flips it back

OpenAI prices any Astra prompt above 272K input tokens at 2x input and cache rates and 1.5x output, for the whole request. That is $20 input, $2 cached and $75 output. Anthropic bills Opus 5’s full 1M window at one rate.

Past that line Opus 5 costs a quarter of Astra on input and a third on output, and Astra’s token savings would have to be very large to close that gap. For long documents, large codebases held in context and any workload that routinely crosses 272K, Opus 5 is the cheaper model. The arithmetic is in the 272K pricing cliff.

Speed is close

Artificial Analysis measures Astra at about 53 output tokens per second and Opus 5 at about 50, both slower than average among the models it tracks. Astra still tends to finish first, because it has fewer tokens to write.

What Opus 5 still has

  • Anthropic’s default. Anthropic’s models page says to start with Opus 5 when you are unsure, and to move up to Fable 5.1 only when that falls short.
  • Long prompts, where the 272K rule makes it several times cheaper.
  • The Claude stack. Claude Code, Anthropic’s platform and any agreements that already cover Anthropic.
  • Readable reasoning. OpenAI itself says Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s. If your oversight depends on reading a model’s reasoning, that counts against Astra.

Which one I would use

The workModelWhy
Long agentic runs, prompts under 272KGPT-6 AstraHigher scores and about 30% cheaper per finished task in both independent measurements.
Prompts that regularly cross 272K tokensClaude Opus 5One rate across 1M tokens, while Astra's whole request doubles on input.
Team already working in Claude CodeClaude Opus 5Anthropic's recommended default, in the harness you already run.
Oversight that relies on reading reasoningClaude Opus 5OpenAI says Astra's written reasoning got harder to monitor.
Choosing a default for a new buildTest bothTwo points apart. Prompt length and harness decide it.

Astra wins short-prompt agentic work on score and cost per task. Opus 5 wins long prompts and Claude-first teams.

For the model Astra ties at the top, see GPT-6 Astra against Claude Fable 5.1. For the step above Opus 5 inside Anthropic’s lineup, Opus 5 against Fable 5.1, and for the harnesses, Claude Code against Codex.

What would change my mind

A cost-per-task comparison in the same harness. Astra’s coding score was measured in Codex and Opus 5’s in Claude Code, and the harness matters as much as the model.

An Opus 5 point release. A two-point index gap is the size one release can close.

A change to OpenAI’s 272K rule. It is the single line that makes Opus 5 the cheaper model on long prompts.

Questions people asked

Is GPT-6 Astra better than Claude Opus 5?

On independent benchmarks, modestly. Artificial Analysis’s Intelligence Index v4.3 scores GPT-6 Astra at 52.81 and Claude Opus 5 at 50.70, and on its Coding Agent Index Codex with Astra scores 62 against Claude Code with Opus 5 at 60. OpenAI’s own launch numbers show much wider gaps on agentic tasks, such as 99.9% against 30.2% on ARC-AGI-3, while Opus 5 led Humanity’s Last Exam with tools, 63.6% to 57.2%. Treat the vendor figures as claims; the independent gap is about two points.

Is GPT-6 Astra more expensive than Claude Opus 5?

Per token, yes, exactly twice: $10 per million input tokens and $50 output against Opus 5’s $5 and $25, and $1 against $0.50 for cached input. Per task it is usually cheaper. Artificial Analysis spent $5,324 running its index on Astra and $7,275 on Opus 5, and measured Codex with Astra at $7.47 per coding task against $10.79 for Claude Code with Opus 5. Above 272K input tokens that reverses: Astra’s whole request moves to $20 input and $75 output while Opus 5 keeps one rate.

Which is faster, GPT-6 Astra or Claude Opus 5?

They generate at similar speeds: Artificial Analysis measures about 53 output tokens per second for Astra and about 50 for Opus 5, both slower than average. Astra tends to finish sooner because it writes far fewer tokens, 60M across the index against Opus 5’s 140M.

Should I use GPT-6 Astra or Claude Opus 5?

GPT-6 Astra for long agentic runs with prompts under 272K tokens, where it scores higher and costs less per finished task. Claude Opus 5 for prompts that cross 272K tokens, for teams whose work already runs in Claude Code or on Anthropic’s platform, and as the default Anthropic itself recommends.

If you are choosing between OpenAI and Anthropic for an agentic product, the contact page is the fastest route to me.

About the author

Wojciech Łuszczyński

Wojciech Łuszczyński

GTM Architect and Growth Operator building AI-native revenue systems for B2B SaaS and technology companies. I connect positioning, SEO, content, paid acquisition, CRM, automation, analytics and AI workflows into practical growth infrastructure.

Newsletter

Get the next one first.

When I publish a new article on AI systems, GTM architecture, or growth operating models, you'll be the first to know.

Subscribe