Skip to content
wojciech.io
All insights
AI SystemsAIOpenAIAI Systems

GPT-6 Sol beats Astra on the SWE benchmark and loses terminal work by thirteen points

Sol lists at a fifth of Astra's price and scores 48 to 53. In the same Codex harness it wins DeepSWE and loses Terminal-Bench. Route by task, not tier.

Wojciech Luszczynski

Wojciech Luszczynski

GTM Architect & Growth Operator · Now · 23 September 2026 · 7 min read

Share

TL;DR · Key insights

  • GPT-6 Sol lists at $2 / $10 per million tokens and GPT-6 Astra at $10 / $50. Five times the price for five points on Artificial Analysis's index, 48 against 53.
  • Per task the gap narrows to about three times: $1.06 against $3.26, because Astra is the more concise of the two at 60M tokens across the run against Sol's 77M.
  • In the same Codex harness Sol scores 69% on DeepSWE v1.1 against Astra's 68%, then loses Terminal-Bench 4.0 by thirteen points, 43% to 56%. One model, two opposite answers.
  • Sol generates at about 125 output tokens per second against Astra's 53, and finished the coding benchmark in 22.3 minutes per task against 29.4.

OpenAI now sells two models for the same job at a five-fold price difference. The tier chart says pick Astra when the work is hard. The benchmark table says that advice is too coarse to use.

In the same harness, on the same day, Sol wins one coding benchmark and loses another by thirteen points. That is a routing rule, not a tier.

InfoGPT-6 SolGPT-6 Astra
Input / output per 1M$2 / $10$10 / $50
Cached input per 1M$0.20$1.00
Above 272K input tokens, in / out$4 / $15$20 / $75
Intelligence Index v4.3.24853
Cost per index task$1.06$3.26
Tokens generated on the index77M60M
Output speed, tokens/sabout 125about 53
Coding Agent Index, Codex harness5762
Cost per coding task$2.99$7.47
Time per coding task22.3m29.4m

Prices and the 272K rule from OpenAI's model pages. Index score, cost per task, tokens and speed from Artificial Analysis's model pages; coding figures from its Coding Agent Index v1.5. Read on 23 September 2026.

Five points, five times the price

Artificial Analysis scores Astra at 53 and Sol at 48 on Intelligence Index v4.3.2. Astra sits sixth of 212 models, Sol eighteenth.

On the price list that difference costs five times as much per token. On a finished task it costs about three times, $3.26 against $1.06, and the reason is that Astra is the more concise of the two: 60M tokens across the index run against Sol’s 77M, which claws back some of what the rate card takes away. Astra thinks less. It charges more for every word of it.

The benchmark split is the useful part

Both models were measured in the Codex harness on the same three coding benchmarks. The result does not line up in a single order.

BenchmarkGPT-6 SolGPT-6 AstraWhat it means
DeepSWE v1.169%68%Repository-scale software engineering. Sol is level with a model costing five times as much.
Terminal-Bench 4.043%56%Agentic terminal and environment work. Thirteen points, the widest gap in the set.
SWE-Atlas-QnA58%62%Codebase question answering. Four points to Astra.
Coding Agent Index5762The composite. Astra's terminal lead carries most of it.
Tokens per task9.8M3.3MSol writes three times as much to get there, and still bills less.

Both models measured inside Codex, so the harness is held constant. The composite hides a split that matters more than the composite does.

Key takeaway

Sol matches Astra on writing and fixing code, and falls apart on driving a terminal. If your agent’s hard part is the environment rather than the diff, the cheaper model is not the one to route to.

Speed is the other five-fold difference

Sol generates at about 125 output tokens per second. Astra manages about 53, which Artificial Analysis flags as notably slow among the models it tracks.

On coding tasks that came out as 22.3 minutes against 29.4. The gap is smaller than the raw rate suggests, because Sol writes three times as many tokens to get to the same place. Still, if a person is watching a progress indicator, Sol is the model that keeps the loop feeling alive.

The 272K rule catches both

Any prompt over 272K input tokens is repriced for the whole request on both models: 2x input and cache rates, 1.5x output. Sol goes to $4 / $15, Astra to $20 / $75.

The ratio between them does not move. So this is not a reason to pick one over the other: it is a reason to keep prompts under the line on either, and the arithmetic is in the 272K pricing cliff.

Which one I would use

The workModelWhy
Writing and fixing code in a repoGPT-6 SolLevel on DeepSWE v1.1, 69% to 68%, at 40% of the cost per task.
Terminal, shell and environment agentsGPT-6 AstraThirteen points ahead on Terminal-Bench 4.0. The one place the price gap is earned.
Anything a person waits onGPT-6 SolAbout 125 output tokens per second against 53.
High-volume runs with a verifierGPT-6 SolRun everything at the cheap tier, re-run the failures on Astra. The split pays for itself below a 30% failure rate.
The hardest single reasoning stepGPT-6 AstraFive index points, and it gets there on a third fewer tokens.

Sol is the default and Astra is the exception, with terminal work as the exception that is easy to name in advance.

For the previous generation of this decision, Astra against GPT-5.6 Sol still holds on the Astra side. Across vendors, Sol against Claude Opus 5.5 is the same question with a different price ladder.

What would change my mind

A Terminal-Bench result that closes. Thirteen points is the only gap here that changes an architecture, and one point release could halve it.

Sol at a lower effort level. Every figure here is max effort. Sol’s advantage is cost, and most of its cost is thinking it may not need.

A harness other than Codex. Both models were measured inside the same harness, which is the right control, but it is one harness.

Questions people asked

Is GPT-6 Astra better than GPT-6 Sol?

By five points on the general index and not uniformly on coding. Artificial Analysis scores Astra at 53 and Sol at 48 on Intelligence Index v4.3.2. In its Codex-harness coding measurements Astra leads on the index, 62 to 57, and on Terminal-Bench 4.0 by thirteen points, but Sol edges it on DeepSWE v1.1, 69% to 68%. The tier ordering holds on average and breaks on specific task shapes.

How much cheaper is GPT-6 Sol than GPT-6 Astra?

Five times per token and about three times per task. Sol lists at $2 per million input tokens and $10 output against Astra’s $10 and $50, and cached input is $0.20 against $1. Artificial Analysis measured $1.06 per index task on Sol and $3.26 on Astra; on coding tasks in the Codex harness the figures were $2.99 and $7.47.

Which is faster, GPT-6 Sol or GPT-6 Astra?

Sol, by more than double. Artificial Analysis measures about 125 output tokens per second for Sol and about 53 for Astra, which it calls notably slow. On coding tasks Sol averaged 22.3 minutes against Astra’s 29.4.

Should I use GPT-6 Sol or GPT-6 Astra?

Sol as the default for volume: agentic coding, retrieval, anything you run thousands of times or can verify and re-run. Astra for terminal and environment work, where it leads by thirteen points, and for the hard steps where five index points change the outcome. Both reprice above 272K input tokens, so long-context work is expensive on either.

If you want the routing rule between these two written into a working agent stack, the contact page is the fastest route to me.

Share this article

About the author

Wojciech Luszczynski

Wojciech Luszczynski

GTM Architect and Growth Operator building AI-native revenue systems for B2B SaaS and technology companies. I connect positioning, SEO, content, paid acquisition, CRM, automation, analytics and AI workflows into practical growth infrastructure.

Newsletter

Get the next one first.

When I publish a new article on AI systems, GTM architecture, or growth operating models, you'll be the first to know.

Subscribe