Which model, and what it costs
to finish the job.
Every pair I have run, indexed by model. Published price, the independent score, and the number that decides it in practice: what one finished task cost, not what a million tokens list for.
30 comparisons across 22 models. Each one is rewritten when the numbers move, not left to rot.

Claude Opus 5.5 beats GPT-6 Astra on every score and still costs more to finish the job

Anthropic cut every price on Claude Opus 5.5. The bill against Opus 5 still went up.

Claude Haiku 4.5 is the cheap tier. GPT-6 Luna just made it the expensive one.
Pick the model you run
Models you can buy today. Each one lists what I have measured it against.
Claude Fable 5.1
Anthropic · 7 comparisonsClaude Sonnet 5
Anthropic · 7 comparisonsClaude Haiku 4.5
Anthropic · 3 comparisonsClaude Opus 5.5
Anthropic · 3 comparisonsClaude Code
Anthropic · 1 comparisonClaude Mythos 5.1
Anthropic · 1 comparisonGPT-6 Astra
OpenAI · 6 comparisonsGPT-6 Sol
OpenAI · 2 comparisonsCodex
OpenAI · 1 comparisonGPT-6 Luna
OpenAI · 1 comparisonGemini 3.8 Flash
Google · 2 comparisonsMuse Spark 1.3
Meta · 1 comparisonGrok 4.6
xAI · 1 comparisonDeepSeek V4.1 Flash
DeepSeek · 1 comparisonModels that have been replaced
Still here, because people run what they already bought. Each article says what took its place.
GPT-5.6 Luna
Claude Sonnet 4.6
GPT-5.6 Terra
Claude Mythos 5
Claude Opus 4.8
In the order I ran them
Newest first, with the verdict on one line.
The systems behind these numbers are the actual work.
Model choice is one decision inside a revenue system. The rest of what I build and run is in the field notes.
Read the field notes →