Skip to content
wojciech.io
Model comparisons

Which model, and what it costs to finish the job.

Every pair I have run, indexed by model. Published price, the independent score, and the number that decides it in practice: what one finished task cost, not what a million tokens list for.

30 comparisons across 22 models. Each one is rewritten when the numbers move, not left to rot.

Pick the model you run

Models you can buy today. Each one lists what I have measured it against.

Claude Haiku 4.5

Anthropic · 3 comparisons

Claude Opus 5.5

Anthropic · 3 comparisons

Claude Code

Anthropic · 1 comparison

Claude Mythos 5.1

Anthropic · 1 comparison

GPT-6 Sol

OpenAI · 2 comparisons

Codex

OpenAI · 1 comparison

GPT-6 Luna

OpenAI · 1 comparison

Gemini 3.8 Flash

Google · 2 comparisons

Muse Spark 1.3

Meta · 1 comparison

Grok 4.6

xAI · 1 comparison

DeepSeek V4.1 Flash

DeepSeek · 1 comparison

Models that have been replaced

Still here, because people run what they already bought. Each article says what took its place.

In the order I ran them

Newest first, with the verdict on one line.

30 comparisons
GPT-6 AstravsClaude Opus 5.5better and dearer per task28 Sept 20268 min readClaude Opus 5.5vsOpus 5cheaper per token, dearer per task23 Sept 20269 min readGPT-6 LunavsClaude Haiku 4.5a tenth of the price, twice the score23 Sept 20267 min readGPT-6 SolvsClaude Opus 5.5the index and the invoice disagree23 Sept 20268 min readGPT-6 SolvsGPT-6 Astraties on code, loses the terminal23 Sept 20267 min readClaude CodevsCodexcoding agent benchmark and cost16 Sept 20267 min readClaude Fable 5.1vsFable 5what changed, what breaks16 Sept 20267 min readGPT-5.6 LunavsClaude Fable 5.115 points, 50x price16 Sept 20266 min readGPT-5.6 TerravsClaude Fable 5.111 points at 5x price16 Sept 20266 min readClaude Haiku 4.5vsGPT-5.6 Lunafive times dearer, 20 points behind16 Sept 20266 min read

The systems behind these numbers are the actual work.

Model choice is one decision inside a revenue system. The rest of what I build and run is in the field notes.

Read the field notes →