GPT-5.6 vs Claude 5: the first tests say the biggest shift isn't a new benchmark leader
On Intelligence Index v4.3 both score 38, but Sonnet 5 cost 22 times more to run the benchmark. Updated September 2026: when Luna wins and when Sonnet 5 does.
GTM Architect & Growth Operator · Now · 12 July 2026 · 19 min read
TL;DR · Key insights
- Updated 16 September 2026 for Artificial Analysis Intelligence Index v4.3. GPT-5.6 Luna scores 37.50 and Claude Sonnet 5 scores 38.36, both shown as 38. Running the index cost $320 on Luna and $6,998 on Sonnet 5.
- Luna lists at $0.20 / $1.20 per million tokens against Sonnet 5's $2 / $10. The run-cost gap is wider than the price gap because Sonnet 5 generated 370M tokens on the index against Luna's 150M.
- Of the five models compared here, the July order holds on v4.3: Claude Fable 5 at 49.70, GPT-5.6 Sol 47.06, Terra 42.25, Sonnet 5 38.36, Luna 37.50. Three newer models now sit above all five: Fable 5.1, GPT-6 Astra and Opus 5.
- Stop asking which model is best. Route the work: Luna to Terra to Sol, or Sonnet 5 to execute and a stronger model to plan. The winner is the architecture, not the row with the highest number.
GPT-5.6 Luna vs Claude Sonnet 5 in September 2026
Most people land on this page with one question, so the current answer comes first.
| Info | GPT-5.6 Luna | Claude Sonnet 5 |
|---|---|---|
| Intelligence Index v4.3 | 37.50 | 38.36 |
| Input / output per 1M | $0.20 / $1.20 | $2 / $10 |
| Cost to run the full index | $320 | $6,998 |
| Tokens generated on the index | 150M | 370M |
Index to two decimals from Artificial Analysis's v4.3 chart data; run cost and tokens from its model pages; list prices from OpenAI and Anthropic. Read on 16 September 2026.
On capability they are level. Under a point separates them, and both display as 38.
On cost they are not close. Luna’s list price is a tenth of Sonnet 5’s on input, and the gap widens on a real workload: Artificial Analysis spent 22 times as much running its index on Sonnet 5, which wrote about two and a half times as many tokens to get through it.
That makes Luna the default for volume work. Sonnet 5 earns its price somewhere else, when the work already runs in Claude Code or on Anthropic’s platform and moving a working pipeline would cost more than the difference.
For the tier below, see Haiku 4.5 against Luna. For the step up, Sonnet 5 against Opus 5. If you are still on Sonnet 4.6, Luna against Sonnet 4.6.
The launch comparison, July 2026
OpenAI and Anthropic shipped their newest model generations within weeks of each other. On the surface it looks like another round of benchmark war. It is not. The interesting move is structural, and it changes how I would build on either stack.
OpenAI stopped selling one “best GPT.” GPT-5.6 is a family of three:
- GPT-5.6 Sol, the flagship,
- GPT-5.6 Terra, the balanced capability-and-cost tier,
- GPT-5.6 Luna, the fast, low-cost tier for volume.
Anthropic positions two:
- Claude Fable 5 as its most capable widely released model,
- Claude Sonnet 5 as the practical agentic model for everyday work, automation and coding.
All five carry roughly a million tokens of context and up to 128,000 output tokens. Context window has stopped being a differentiator. What separates them now is reliability, cost per finished task, tool use, speed, and how much supervision they need to run without me watching.
| Model | Index, July round | Index v4.3 | In / out per 1M | Most logical role |
|---|---|---|---|---|
| Claude Fable 5 | 60 | 49.70 | $10 / $50 | Hardest problems, analysis, planning, high-stakes coding |
| GPT-5.6 Sol | 59 | 47.06 | $4 / $20 | Frontier capability at better economics, strong Codex integration |
| GPT-5.6 Terra | 55 | 42.25 | $2 / $12 | Default model for professional daily work |
| Claude Sonnet 5 | 53 | 38.36 | $2 / $10 | Clearly specified agentic execution and Claude Code |
| GPT-5.6 Luna | 51 | 37.50 | $0.20 / $1.20 | Fast, high-volume, low-cost workloads |
Five models at a glance. Artificial Analysis Intelligence Index in the July round and on v4.3, which use different scales, current API list prices, and where each one earns its keep. The rest of this article is why the highlighted row is not the whole story.
The July finding: Fable led, Sol was one point behind for a third of the price
In the July round, Claude Fable 5 scored 60 on the Artificial Analysis Intelligence Index and GPT-5.6 Sol at max reasoning scored 59. Artificial Analysis put Sol’s cost at about $1.04 per task in that evaluation, roughly one-third of Fable’s. Sol also took the lead on the Coding Agent Index with a score of 80.
Artificial Analysis Intelligence Index
+2% vs baselineOn v4.3 the gap is 2.6 points, 49.70 against 47.06, and running the index cost Artificial Analysis $11,161 on Fable 5 against $3,465 on Sol. Still not a knockout. Still a change in market economics.
Fable can still be the right call for the single hardest problem, where model cost matters less than the probability of a correct answer. Sol looks like the better default for repeated frontier-level work, especially anything you run many times a day.
What GPT-5.6 actually changes
Sol: close to Fable, materially cheaper
GPT-5.6 Sol lists at $4 per million input tokens and $20 output. Fable 5 lists at $10 and $50.
Token price alone never tells the truth, because models spend different amounts of reasoning, take different numbers of steps and need different amounts of human correction. But the independent cost-per-task number says Sol’s advantage is more than a pricing-table trick.
Sol cost per task
$1.04
Artificial Analysis eval
Fable's cost vs Sol
~3x
cheaper
same benchmark
Sol, Coding Agent Index
80
current leader
Claire Vo ran Sol, Terra, Luna, Fable and Sonnet across PRDs, prototypes, wireframes, debugging and agentic voice. Sol won the overall test, strongest on PRDs, prototyping and browser use. Fable’s problem was not raw capability. It was excessive precision and a harder collaboration style during the work itself.
Early developer reports point the same way. People describe Sol as close to Fable, but faster, cheaper and more dependable on changes that touch several layers of a system at once. Anecdotes, not controlled experiments, but the direction matches the independent numbers.
Fable did not lose on capability. It lost on how much work it was to work with.
Terra may be the most important model in the release
Sol gets the attention. Terra may have the bigger operational impact.
Terra lists at $2 per million input tokens and $12 output. OpenAI describes it as competitive with GPT-5.5 at half the cost. Artificial Analysis scored it 55 in the July round, four points behind Sol. On v4.3 it is 42.25, 4.8 points behind.
That is enough for a large share of real production work:
Business analysis
TerraReads, summarises and reasons over messy inputs well enough to brief a decision.
Document and research work
TerraThe volume layer of professional output: drafts, extraction, synthesis across sources.
Code generation and refactoring
TerraModerate-difficulty changes with clear scope, not the ambiguous architecture calls.
Tool use and agent execution
TerraMulti-step runs where the plan is defined and the model executes it.
Not everyone is sold. Some testers note Terra scores below GPT-5.5 on several individual benchmarks and behaves more like a distilled mid-sized model than a cheaper Sol. Fair. A strong aggregate score does not mean even quality across every workload. But for a company the question is not “does Terra win every eval.” It is:
Does Terra finish enough of our tasks that paying for Sol by default becomes unnecessary?
For a lot of workloads the answer is yes.
Luna shows how fast capability is getting cheap
GPT-5.6 Luna lists at $0.20 per million input tokens and $1.20 output. It scored 51.2 in the July round, only just below Sonnet 5 at maximum reasoning, and on v4.3 the two are under a point apart.
That does not mean Luna beats Sonnet on every task. Sonnet can be much better at specific agentic, coding or long-running work. What Luna shows is the pricing pressure GPT-5.6 creates: capability that recently needed an expensive frontier model now sits in a twenty-cent input tier.
Where Luna is the rational pick
Judge Luna by the work it unlocks economically, not by whether it beats Fable on the hardest possible problem.
High volume, low ambiguity, clear pass or fail. The classic cheap-model job.
First drafts, variants, rewrites at scale where a human edits the output anyway.
Triage a task, decide difficulty, then escalate the hard ones to a stronger tier.
So the best architecture is not Sol for everything. It routes work by difficulty and risk.
Fable 5 still has the strongest pure-capability case
Of these five, Claude Fable 5 stays first on the general index and performs particularly well on long analytical tasks, agentic coding, prototyping, spreadsheet work and problems that need self-verification. Anthropic’s early partners described it as able to finish work that used to take many prompts, and to reflect on and check its own output at the highest effort setting.
Fable also scored 80% on SWE-Bench Pro against Sol’s 64.6%. But Sol led Terminal-Bench 2.1 at 88.8% against Fable’s 83.1%.
SWE-Bench Pro
+24% vs baselineTerminal-Bench 2.1
+7% vs baselineThe two benchmarks measure different jobs, and the harness matters as much as the model:
- SWE-Bench rewards fixing defined issues in a repo.
- Terminal-Bench tests work in a terminal environment.
- Codex and Claude Code add their own harnesses, system prompts, context management and tool strategies on top.
The same underlying model performs differently through a raw API, inside Claude Code, inside Codex, or inside a third-party agent framework. Judge the environment, not just the weights.
Fable’s real disadvantages are cost and access. After its temporary suspension the model came back globally, but plenty of users criticised Anthropic for moving Fable use outside normal subscription allowances and onto extra paid credits. Fable may be the best specialist without being the best default employee.
Sonnet 5: a capable model with a positioning problem
Sonnet 5 was meant to answer a simple question: which Claude model should people use every day?
The product story is attractive. Sonnet 5 is available across Claude plans, the default for Free and Pro, runs in Claude Code, and costs $2 per million input tokens and $10 output. That started as introductory pricing due to expire on 31 August 2026, but Anthropic has confirmed the scheduled rise to $3 and $15 will not happen and $2/$10 is now the standard price. Anthropic pitches it as materially more agentic than Sonnet 4.6: better at finishing multi-step tasks, checking its own output and working on messy existing codebases. Early-access partners reported strong debugging, tool use, testing and brownfield results.
The independent picture is more mixed.
That does not make Sonnet 5 weak. It suggests planning highly ambiguous work is not its natural role. A more rational split:
| The job | My pick | Why |
|---|---|---|
| Architecture, planning, high-uncertainty work | Fable or Sol | The expensive part is being wrong. Pay for the model least likely to send you down a bad path. |
| Implementing a clearly specified task | Sonnet 5 | Scope is set, execution is the work. This is where its agentic strength shows and the price is right. |
| Cost-effective daily analysis and execution | Terra | Frontier-adjacent quality at half the flagship cost, for the bulk of professional output. |
| Scale and low-risk workflows | Luna | Volume that never justified a frontier model before now runs at twenty cents per million in. |
Match the model to the risk of the task, not to a single leaderboard row.

The real contest: ChatGPT Work and Codex vs Claude Cowork and Claude Code
Model-only comparisons describe real productivity worse every month.
OpenAI put GPT-5.6 across ChatGPT, ChatGPT Work, Codex and the API. Paid users pick Sol, Terra or Luna inside Work and Codex and control reasoning effort. Max and ultra modes exist on eligible plans, and the API supports programmatic tool calling and parallel subagents. Anthropic offers a comparable system through Claude Chat, Cowork and Claude Code.
OpenAI: ChatGPT Work + Codex
tighter economicsSol, Terra and Luna selectable inside Work and Codex, effort control, max and ultra modes, programmatic tool calling and parallel subagents. The pull is tighter integration with the rest of ChatGPT and more aggressive model economics.
Anthropic: Cowork + Claude Code
mature dev UXThe same span of jobs through Claude Chat, Cowork and Claude Code. The edge stays Claude Code: still the most mature developer experience, especially for working naturally with an existing, messy codebase.
Both stacks move between the same jobs: conversation and analysis, documents and files, application-level work, editing a repository, running code, delegating to subagents. Anthropic’s edge stays Claude Code, still the most mature and widely preferred developer experience, especially for working naturally with an existing codebase. OpenAI’s answer is tighter integration between Codex and the rest of ChatGPT, plus more aggressive model economics.
The first feedback does not crown a universal winner. Some developers still prefer Claude for frontend work, sticking to project conventions and generating code closer to intent. Others report Codex with Sol is more reliable when a problem crosses several layers of a system, and less likely to get stuck in a local solution.
What the numbers actually say
Two conclusions survive all the caveats.
First, the capability gap between the flagship and the cheaper tiers is shrinking. In July one point separated Fable 5 and Sol on the general index, and four separated Sol and Terra. On v4.3 those gaps are 2.6 and 4.8 points. The premium models are no longer a different league, just a different price.
Second, token price is not the number that matters. A cheaper model can be more expensive per outcome when it generates more reasoning tokens, takes more steps or needs more correction. Cost per finished result is the only figure that touches a P&L, and it does not appear on any pricing page.
What the price page shows
Cost per token
What actually hits your budget
Cost per finished result
How I would route the work
For most professional work: Terra
Terra is the best starting point for analysis, research, document work, content and moderately hard coding. Sol should be an escalation path, not the universal default.
For the hardest problems: Sol or Fable
Fable keeps a marginal overall lead and performs strongly on long, complex tasks. Sol delivers close capability at much better economics and is the more pragmatic organisational default.
Inside Claude Code: Sonnet executes, Fable plans
Sonnet 5 fits clearly defined implementation. On ambiguous architectural work, using a stronger model to prepare the plan usually costs less than correcting the cheaper one three times.
For high-volume automation: Luna
Do not judge Luna by whether it beats Fable on the hardest problem. Its value is enabling whole categories of work that never economically justified frontier AI before.
Final verdict
OpenAI did not need to defeat Fable decisively. It only had to get close at a much lower cost, then distribute the same generation of capability across three economic tiers. As a result, “which model is best?” is becoming a weaker question.
Better ones:
- Which model actually finishes my specific workflow?
- How much supervision does it need?
- What does a correct outcome cost, not a million tokens?
- Can it operate across my documents, applications and repositories?
- When should the work escalate automatically to a stronger model?
Frontier AI is no longer one model. It is becoming a tiered production system. In that system the winner is not the model with the highest number on one chart. It is the architecture that routes the right task to the right level of intelligence.
FAQ: GPT-5.6 vs Claude 5
Is GPT-5.6 Luna better than Claude Sonnet 5?
They are level on capability. On Artificial Analysis’s Intelligence Index v4.3 Luna scores 37.50 and Sonnet 5 at maximum effort 38.36, both shown as 38. Luna lists at $0.20 per million input tokens and $1.20 output, against Sonnet 5’s $2 and $10, and running the whole index cost $320 on Luna against $6,998 on Sonnet 5, because Sonnet 5 generated 370M tokens to Luna’s 150M. Use Luna for high-volume, well-defined work where a wrong answer is cheap to catch: classification, extraction, routing, first drafts. Use Sonnet 5 where the work already runs in Claude Code or on Anthropic’s platform and moving it would cost more than the price gap. Judge both by cost per finished task, not price per token.
Which AI model is the best right now, GPT-5.6 or Claude 5?
Of the five models in this comparison, Claude Fable 5 still scores highest on Artificial Analysis’s Intelligence Index v4.3 at 49.70, with GPT-5.6 Sol at 47.06: the same order as in July, on a lower scale. Three newer models now sit above all five: Claude Fable 5.1 at 53.37, GPT-6 Astra at 52.81 and Claude Opus 5 at 50.70. There is no single winner: the right model depends on the task, the cost per finished result and how much supervision you can spare.
Is GPT-5.6 cheaper than Claude Fable 5?
Yes. GPT-5.6 Sol lists at $4 per million input tokens and $20 output against Fable 5’s $10 and $50, and running Artificial Analysis’s Intelligence Index v4.3 cost $3,465 on Sol against $11,161 on Fable 5. Sol’s price is promotional through at least 21 November 2026. Terra ($2 / $12) and Luna ($0.20 / $1.20) undercut every Claude tier further.
Should I use Claude Sonnet 5 in Claude Code?
For clearly specified implementation, yes: it is quick, it is built into Claude Code, and Anthropic pitches it as strong on debugging, tool use and messy existing codebases. For ambiguous architecture, plan with a stronger model first. Sonnet 5 is verbose: on the v4.3 index run it generated 370M tokens, about four times the 90M median, so planning the hardest work with it is often a false economy.
What is the best GPT-5.6 model for everyday professional work?
Terra for most of it: analysis, research, document work and moderately hard coding at half the flagship’s input price. Keep Sol as an escalation path for the hardest problems, and use Luna for high-volume, low-risk workloads like classification, extraction and routing.
Sources and further reading
Primary sources
Top-level sources for the numbers cited above. Deep-linked permalinks change; these are the pages to check the live figures against.
If you want the operator context around this, I have written up my own benchmark of Claude Fable 5 against Opus and Sonnet and the AI production stack I actually ship on. Same principle as here: the model is one layer, the system that routes work around it is the job.

