Skip to content
wojciech.io
All model comparisons
AI SystemsAIOpenAIClaudeAI Systems

GPT-5.6 vs Claude 5: the first tests say the biggest shift isn't a new benchmark leader

On Intelligence Index v4.3 both score 38, but Sonnet 5 cost 22 times more to run the benchmark. Updated September 2026: when Luna wins and when Sonnet 5 does.

Wojciech Luszczynski

Wojciech Luszczynski

GTM Architect & Growth Operator · Now · 12 July 2026 · 19 min read

Share

TL;DR · Key insights

  • Updated 16 September 2026 for Artificial Analysis Intelligence Index v4.3. GPT-5.6 Luna scores 37.50 and Claude Sonnet 5 scores 38.36, both shown as 38. Running the index cost $320 on Luna and $6,998 on Sonnet 5.
  • Luna lists at $0.20 / $1.20 per million tokens against Sonnet 5's $2 / $10. The run-cost gap is wider than the price gap because Sonnet 5 generated 370M tokens on the index against Luna's 150M.
  • Of the five models compared here, the July order holds on v4.3: Claude Fable 5 at 49.70, GPT-5.6 Sol 47.06, Terra 42.25, Sonnet 5 38.36, Luna 37.50. Three newer models now sit above all five: Fable 5.1, GPT-6 Astra and Opus 5.
  • Stop asking which model is best. Route the work: Luna to Terra to Sol, or Sonnet 5 to execute and a stronger model to plan. The winner is the architecture, not the row with the highest number.

GPT-5.6 Luna vs Claude Sonnet 5 in September 2026

Most people land on this page with one question, so the current answer comes first.

InfoGPT-5.6 LunaClaude Sonnet 5
Intelligence Index v4.337.5038.36
Input / output per 1M$0.20 / $1.20$2 / $10
Cost to run the full index$320$6,998
Tokens generated on the index150M370M

Index to two decimals from Artificial Analysis's v4.3 chart data; run cost and tokens from its model pages; list prices from OpenAI and Anthropic. Read on 16 September 2026.

On capability they are level. Under a point separates them, and both display as 38.

On cost they are not close. Luna’s list price is a tenth of Sonnet 5’s on input, and the gap widens on a real workload: Artificial Analysis spent 22 times as much running its index on Sonnet 5, which wrote about two and a half times as many tokens to get through it.

That makes Luna the default for volume work. Sonnet 5 earns its price somewhere else, when the work already runs in Claude Code or on Anthropic’s platform and moving a working pipeline would cost more than the difference.

For the tier below, see Haiku 4.5 against Luna. For the step up, Sonnet 5 against Opus 5. If you are still on Sonnet 4.6, Luna against Sonnet 4.6.

The launch comparison, July 2026

OpenAI and Anthropic shipped their newest model generations within weeks of each other. On the surface it looks like another round of benchmark war. It is not. The interesting move is structural, and it changes how I would build on either stack.

OpenAI stopped selling one “best GPT.” GPT-5.6 is a family of three:

  • GPT-5.6 Sol, the flagship,
  • GPT-5.6 Terra, the balanced capability-and-cost tier,
  • GPT-5.6 Luna, the fast, low-cost tier for volume.

Anthropic positions two:

  • Claude Fable 5 as its most capable widely released model,
  • Claude Sonnet 5 as the practical agentic model for everyday work, automation and coding.

All five carry roughly a million tokens of context and up to 128,000 output tokens. Context window has stopped being a differentiator. What separates them now is reliability, cost per finished task, tool use, speed, and how much supervision they need to run without me watching.

ModelIndex, July roundIndex v4.3In / out per 1MMost logical role
Claude Fable 56049.70$10 / $50Hardest problems, analysis, planning, high-stakes coding
GPT-5.6 Sol5947.06$4 / $20Frontier capability at better economics, strong Codex integration
GPT-5.6 Terra5542.25$2 / $12Default model for professional daily work
Claude Sonnet 55338.36$2 / $10Clearly specified agentic execution and Claude Code
GPT-5.6 Luna5137.50$0.20 / $1.20Fast, high-volume, low-cost workloads

Five models at a glance. Artificial Analysis Intelligence Index in the July round and on v4.3, which use different scales, current API list prices, and where each one earns its keep. The rest of this article is why the highlighted row is not the whole story.

The July finding: Fable led, Sol was one point behind for a third of the price

In the July round, Claude Fable 5 scored 60 on the Artificial Analysis Intelligence Index and GPT-5.6 Sol at max reasoning scored 59. Artificial Analysis put Sol’s cost at about $1.04 per task in that evaluation, roughly one-third of Fable’s. Sol also took the lead on the Coding Agent Index with a score of 80.

Artificial Analysis Intelligence Index

+2% vs baseline
Claude Fable 560
GPT-5.6 Sol59
July round. One point apart at the top: on the general index Fable kept a hair's-width lead, the kind of gap that disappears the moment you look at cost.

On v4.3 the gap is 2.6 points, 49.70 against 47.06, and running the index cost Artificial Analysis $11,161 on Fable 5 against $3,465 on Sol. Still not a knockout. Still a change in market economics.

Fable can still be the right call for the single hardest problem, where model cost matters less than the probability of a correct answer. Sol looks like the better default for repeated frontier-level work, especially anything you run many times a day.

What GPT-5.6 actually changes

Sol: close to Fable, materially cheaper

GPT-5.6 Sol lists at $4 per million input tokens and $20 output. Fable 5 lists at $10 and $50.

Token price alone never tells the truth, because models spend different amounts of reasoning, take different numbers of steps and need different amounts of human correction. But the independent cost-per-task number says Sol’s advantage is more than a pricing-table trick.

Sol cost per task

$1.04

Artificial Analysis eval

Fable's cost vs Sol

~3x

cheaper

same benchmark

Sol, Coding Agent Index

80

current leader

Sol's case is not that it wins every chart. It is that it lands next to Fable on quality while costing a fraction per completed job.

Claire Vo ran Sol, Terra, Luna, Fable and Sonnet across PRDs, prototypes, wireframes, debugging and agentic voice. Sol won the overall test, strongest on PRDs, prototyping and browser use. Fable’s problem was not raw capability. It was excessive precision and a harder collaboration style during the work itself.

Early developer reports point the same way. People describe Sol as close to Fable, but faster, cheaper and more dependable on changes that touch several layers of a system at once. Anecdotes, not controlled experiments, but the direction matches the independent numbers.

Fable did not lose on capability. It lost on how much work it was to work with.

The recurring note in the first cross-model tests

Terra may be the most important model in the release

Sol gets the attention. Terra may have the bigger operational impact.

Terra lists at $2 per million input tokens and $12 output. OpenAI describes it as competitive with GPT-5.5 at half the cost. Artificial Analysis scored it 55 in the July round, four points behind Sol. On v4.3 it is 42.25, 4.8 points behind.

That is enough for a large share of real production work:

Business analysis

Terra

Reads, summarises and reasons over messy inputs well enough to brief a decision.

Document and research work

Terra

The volume layer of professional output: drafts, extraction, synthesis across sources.

Code generation and refactoring

Terra

Moderate-difficulty changes with clear scope, not the ambiguous architecture calls.

Tool use and agent execution

Terra

Multi-step runs where the plan is defined and the model executes it.

Not every benchmark, but enough production tasks that paying for Sol by default stops making sense.

Not everyone is sold. Some testers note Terra scores below GPT-5.5 on several individual benchmarks and behaves more like a distilled mid-sized model than a cheaper Sol. Fair. A strong aggregate score does not mean even quality across every workload. But for a company the question is not “does Terra win every eval.” It is:

Does Terra finish enough of our tasks that paying for Sol by default becomes unnecessary?

For a lot of workloads the answer is yes.

Luna shows how fast capability is getting cheap

GPT-5.6 Luna lists at $0.20 per million input tokens and $1.20 output. It scored 51.2 in the July round, only just below Sonnet 5 at maximum reasoning, and on v4.3 the two are under a point apart.

That does not mean Luna beats Sonnet on every task. Sonnet can be much better at specific agentic, coding or long-running work. What Luna shows is the pricing pressure GPT-5.6 creates: capability that recently needed an expensive frontier model now sits in a twenty-cent input tier.

Where Luna is the rational pick

Judge Luna by the work it unlocks economically, not by whether it beats Fable on the hardest possible problem.

Classification and extraction

High volume, low ambiguity, clear pass or fail. The classic cheap-model job.

Bulk content processing

First drafts, variants, rewrites at scale where a human edits the output anyway.

Simple agents and routing

Triage a task, decide difficulty, then escalate the hard ones to a stronger tier.

So the best architecture is not Sol for everything. It routes work by difficulty and risk.

Incoming task
triage
Lunavolume · low risk
too hard
Terradaily professional work
high stakes
Solhardest problems only
The tiered routing that beats a single-model default: cheapest model that can finish the job, escalate only when it can't

Fable 5 still has the strongest pure-capability case

Of these five, Claude Fable 5 stays first on the general index and performs particularly well on long analytical tasks, agentic coding, prototyping, spreadsheet work and problems that need self-verification. Anthropic’s early partners described it as able to finish work that used to take many prompts, and to reflect on and check its own output at the highest effort setting.

Fable also scored 80% on SWE-Bench Pro against Sol’s 64.6%. But Sol led Terminal-Bench 2.1 at 88.8% against Fable’s 83.1%.

SWE-Bench Pro

+24% vs baseline
Claude Fable 580%
GPT-5.6 Sol64.6%
Fixing defined repository issues: Fable's home ground, a clear lead.

Terminal-Bench 2.1

+7% vs baseline
GPT-5.6 Sol88.8%
Claude Fable 583.1%
Work inside a terminal environment: Sol takes it. Same two models, opposite result. This is why 'the best coding model' is a useless label.

The two benchmarks measure different jobs, and the harness matters as much as the model:

  • SWE-Bench rewards fixing defined issues in a repo.
  • Terminal-Bench tests work in a terminal environment.
  • Codex and Claude Code add their own harnesses, system prompts, context management and tool strategies on top.

The same underlying model performs differently through a raw API, inside Claude Code, inside Codex, or inside a third-party agent framework. Judge the environment, not just the weights.

Fable’s real disadvantages are cost and access. After its temporary suspension the model came back globally, but plenty of users criticised Anthropic for moving Fable use outside normal subscription allowances and onto extra paid credits. Fable may be the best specialist without being the best default employee.

Sonnet 5: a capable model with a positioning problem

Sonnet 5 was meant to answer a simple question: which Claude model should people use every day?

The product story is attractive. Sonnet 5 is available across Claude plans, the default for Free and Pro, runs in Claude Code, and costs $2 per million input tokens and $10 output. That started as introductory pricing due to expire on 31 August 2026, but Anthropic has confirmed the scheduled rise to $3 and $15 will not happen and $2/$10 is now the standard price. Anthropic pitches it as materially more agentic than Sonnet 4.6: better at finishing multi-step tasks, checking its own output and working on messy existing codebases. Early-access partners reported strong debugging, tool use, testing and brownfield results.

The independent picture is more mixed.

That does not make Sonnet 5 weak. It suggests planning highly ambiguous work is not its natural role. A more rational split:

The jobMy pickWhy
Architecture, planning, high-uncertainty workFable or SolThe expensive part is being wrong. Pay for the model least likely to send you down a bad path.
Implementing a clearly specified taskSonnet 5Scope is set, execution is the work. This is where its agentic strength shows and the price is right.
Cost-effective daily analysis and executionTerraFrontier-adjacent quality at half the flagship cost, for the bulk of professional output.
Scale and low-risk workflowsLunaVolume that never justified a frontier model before now runs at twenty cents per million in.

Match the model to the risk of the task, not to a single leaderboard row.

Operator working late at a desk with a laptop, notebook and a small shelf of books, city lights behind the window
The choice is no longer which model is smartest. It is which environment finishes your work.

The real contest: ChatGPT Work and Codex vs Claude Cowork and Claude Code

Model-only comparisons describe real productivity worse every month.

OpenAI put GPT-5.6 across ChatGPT, ChatGPT Work, Codex and the API. Paid users pick Sol, Terra or Luna inside Work and Codex and control reasoning effort. Max and ultra modes exist on eligible plans, and the API supports programmatic tool calling and parallel subagents. Anthropic offers a comparable system through Claude Chat, Cowork and Claude Code.

OpenAI: ChatGPT Work + Codex

tighter economics

Sol, Terra and Luna selectable inside Work and Codex, effort control, max and ultra modes, programmatic tool calling and parallel subagents. The pull is tighter integration with the rest of ChatGPT and more aggressive model economics.

Anthropic: Cowork + Claude Code

mature dev UX

The same span of jobs through Claude Chat, Cowork and Claude Code. The edge stays Claude Code: still the most mature developer experience, especially for working naturally with an existing, messy codebase.

This is no longer a choice between chatbots. It is a choice between execution environments for digital work.

Both stacks move between the same jobs: conversation and analysis, documents and files, application-level work, editing a repository, running code, delegating to subagents. Anthropic’s edge stays Claude Code, still the most mature and widely preferred developer experience, especially for working naturally with an existing codebase. OpenAI’s answer is tighter integration between Codex and the rest of ChatGPT, plus more aggressive model economics.

The first feedback does not crown a universal winner. Some developers still prefer Claude for frontend work, sticking to project conventions and generating code closer to intent. Others report Codex with Sol is more reliable when a problem crosses several layers of a system, and less likely to get stuck in a local solution.

What the numbers actually say

Two conclusions survive all the caveats.

First, the capability gap between the flagship and the cheaper tiers is shrinking. In July one point separated Fable 5 and Sol on the general index, and four separated Sol and Terra. On v4.3 those gaps are 2.6 and 4.8 points. The premium models are no longer a different league, just a different price.

Second, token price is not the number that matters. A cheaper model can be more expensive per outcome when it generates more reasoning tokens, takes more steps or needs more correction. Cost per finished result is the only figure that touches a P&L, and it does not appear on any pricing page.

What the price page shows

Cost per token

What actually hits your budget

Cost per finished result

The metric that decides everything, and the one no launch post puts on the slide

How I would route the work

For most professional work: Terra

Terra is the best starting point for analysis, research, document work, content and moderately hard coding. Sol should be an escalation path, not the universal default.

For the hardest problems: Sol or Fable

Fable keeps a marginal overall lead and performs strongly on long, complex tasks. Sol delivers close capability at much better economics and is the more pragmatic organisational default.

Inside Claude Code: Sonnet executes, Fable plans

Sonnet 5 fits clearly defined implementation. On ambiguous architectural work, using a stronger model to prepare the plan usually costs less than correcting the cheaper one three times.

For high-volume automation: Luna

Do not judge Luna by whether it beats Fable on the hardest problem. Its value is enabling whole categories of work that never economically justified frontier AI before.

Final verdict

OpenAI did not need to defeat Fable decisively. It only had to get close at a much lower cost, then distribute the same generation of capability across three economic tiers. As a result, “which model is best?” is becoming a weaker question.

Better ones:

  • Which model actually finishes my specific workflow?
  • How much supervision does it need?
  • What does a correct outcome cost, not a million tokens?
  • Can it operate across my documents, applications and repositories?
  • When should the work escalate automatically to a stronger model?

Frontier AI is no longer one model. It is becoming a tiered production system. In that system the winner is not the model with the highest number on one chart. It is the architecture that routes the right task to the right level of intelligence.

FAQ: GPT-5.6 vs Claude 5

Is GPT-5.6 Luna better than Claude Sonnet 5?

They are level on capability. On Artificial Analysis’s Intelligence Index v4.3 Luna scores 37.50 and Sonnet 5 at maximum effort 38.36, both shown as 38. Luna lists at $0.20 per million input tokens and $1.20 output, against Sonnet 5’s $2 and $10, and running the whole index cost $320 on Luna against $6,998 on Sonnet 5, because Sonnet 5 generated 370M tokens to Luna’s 150M. Use Luna for high-volume, well-defined work where a wrong answer is cheap to catch: classification, extraction, routing, first drafts. Use Sonnet 5 where the work already runs in Claude Code or on Anthropic’s platform and moving it would cost more than the price gap. Judge both by cost per finished task, not price per token.

Which AI model is the best right now, GPT-5.6 or Claude 5?

Of the five models in this comparison, Claude Fable 5 still scores highest on Artificial Analysis’s Intelligence Index v4.3 at 49.70, with GPT-5.6 Sol at 47.06: the same order as in July, on a lower scale. Three newer models now sit above all five: Claude Fable 5.1 at 53.37, GPT-6 Astra at 52.81 and Claude Opus 5 at 50.70. There is no single winner: the right model depends on the task, the cost per finished result and how much supervision you can spare.

Is GPT-5.6 cheaper than Claude Fable 5?

Yes. GPT-5.6 Sol lists at $4 per million input tokens and $20 output against Fable 5’s $10 and $50, and running Artificial Analysis’s Intelligence Index v4.3 cost $3,465 on Sol against $11,161 on Fable 5. Sol’s price is promotional through at least 21 November 2026. Terra ($2 / $12) and Luna ($0.20 / $1.20) undercut every Claude tier further.

Should I use Claude Sonnet 5 in Claude Code?

For clearly specified implementation, yes: it is quick, it is built into Claude Code, and Anthropic pitches it as strong on debugging, tool use and messy existing codebases. For ambiguous architecture, plan with a stronger model first. Sonnet 5 is verbose: on the v4.3 index run it generated 370M tokens, about four times the 90M median, so planning the hardest work with it is often a false economy.

What is the best GPT-5.6 model for everyday professional work?

Terra for most of it: analysis, research, document work and moderately hard coding at half the flagship’s input price. Keep Sol as an escalation path for the hardest problems, and use Luna for high-volume, low-risk workloads like classification, extraction and routing.

Sources and further reading

Primary sources

Top-level sources for the numbers cited above. Deep-linked permalinks change; these are the pages to check the live figures against.

If you want the operator context around this, I have written up my own benchmark of Claude Fable 5 against Opus and Sonnet and the AI production stack I actually ship on. Same principle as here: the model is one layer, the system that routes work around it is the job.

Share this article

About the author

Wojciech Luszczynski

Wojciech Luszczynski

GTM Architect and Growth Operator building AI-native revenue systems for B2B SaaS and technology companies. I connect positioning, SEO, content, paid acquisition, CRM, automation, analytics and AI workflows into practical growth infrastructure.

Newsletter

Get the next one first.

When I publish a new article on AI systems, GTM architecture, or growth operating models, you'll be the first to know.

Subscribe