Skip to content
wojciech.io
All insights
AI SystemsAIAI Systems

Fable 5 scored 60 in July and 49.7 today. The model did not change. The index did.

What v4.3 measures, what changed on 7 September, the current ranking to two decimals, and why AI model scores from July cannot be compared with today's.

Wojciech Łuszczyński

Wojciech Łuszczyński

GTM Architect & Growth Operator · Now · 16 September 2026

TL;DR · Key insights

  • Artificial Analysis published Intelligence Index v4.3 on 7 September 2026. It upgraded Terminal-Bench from 2.1 to 4.0 and replaced a banking task with AutomationBench-AA, built with Zapier on a private test set.
  • The index combines ten evaluations under fixed category weights: agents 30%, coding 20%, general 30%, scientific reasoning 20%. Evaluations with private tasks or answers now carry 45% of the weight, up from 40%.
  • The scale moves between versions. Claude Fable 5 read 60 in the July round and reads 49.70 on v4.3 without any change to the model. A score is only comparable with other scores from the same version.
  • At the top of v4.3: Claude Fable 5.1 at 53.37, GPT-6 Astra at 52.81, Claude Opus 5 at 50.70. Artificial Analysis calls v4.3 a continuation of its rollout of v5, so expect the scale to move again.

Two articles on this site quote Claude Fable 5 with two different Intelligence Index scores. The July comparison says 60. A piece from early September says 62.1. Artificial Analysis’s own page today says 49.70.

None of those is a typo, and the model never changed. The yardstick did, twice.

This is the explainer I needed before writing any of the comparisons: what Artificial Analysis’s Intelligence Index measures, what version 4.3 changed, and how to read a score without comparing it to a number from a different ruler.

What the index is made of

Version 4.3 combines ten independently run evaluations:

  • AA-Briefcase and GDPval-AA v2
  • AutomationBench-AA, new in v4.3
  • Terminal-Bench 4.0, upgraded in v4.3
  • SciCode, Humanity’s Last Exam, GDP.pdf and CritPt
  • AA-Omniscience, Artificial Analysis’s knowledge and hallucination benchmark
  • AA-LCR v1.1, for long-context reasoning

They roll up into four categories with fixed weights, and those did not change in v4.3:

CategoryWeight
Agents30%
General30%
Coding20%
Scientific reasoning20%

Category weights as stated in Artificial Analysis's v4.3 announcement, read on 16 September 2026.

Half the index is agents and coding. That matters when you read it: a model that is excellent at single-turn questions and weak at multi-step work will score lower here than its chat reputation suggests.

What changed on 7 September

Two evaluations moved.

Terminal-Bench went from 2.1 to 4.0. Artificial Analysis describes 4.0 as having harder tasks worked through a terminal, across software, machine learning, science, operations, security, hardware and media, with recalibrated time and compute allowances. It runs all 66 tasks three times and reports average pass@1.

τ³-Banking was replaced by AutomationBench-AA, Artificial Analysis’s implementation of Zapier’s business workflow automation benchmark, run on a held-out test set built with Zapier.

That held-out set had a knock-on effect. The weight on evaluations with private tasks or answers rose from 40% to 45%. Private test sets are harder to train toward, which is the stated point: Artificial Analysis says each change in v4.2 and v4.3 adds more private test sets to prevent gaming and reduces saturation.

Why the scores dropped

Harder evaluations lower everyone’s score at once. That is why Fable 5 went from 60 to 49.70: a different mix of evaluations, harder tasks and more of them hidden, not a worse model.

It is also why mixing versions produces nonsense. Put Fable 5’s July 60 next to Sonnet 5’s v4.3 38 and you get a 22-point gap. On a single version the gap is about 12. Both numbers are real; only one comparison is.

Key takeaway

A score belongs to a version. Compare models within one version, read the gaps rather than the absolute values, and check the version before you quote anyone’s number, including mine.

The current ranking

These are the values embedded in Artificial Analysis’s v4.3 announcement, to two decimals. The public pages round them, which is why two models separated by half a point both display as 53.

#Model and variantIndex v4.3
1Claude Fable 5.1 (max, with fallback)53.37
2GPT-6 Astra (max)52.81
3Claude Opus 5 (max)50.70
4Claude Fable 5 (with fallback)49.70
5Muse Spark 1.3 (max)48.17
6GPT-5.6 Sol (max)47.06
7GLM-5.3 (max)44.45
8Grok 4.6 (high)44.41
9Kimi K3 (max)43.78
10GPT-5.6 Terra (max)42.25
11GLM-5.3-Flash41.91
12Gemini 3.8 Flash (high)41.19
13Qwen3.8 2.4T A95B40.04
14GPT-5.6 Luna (max)37.50
15DeepSeek V4 Pro 0813 (max)36.28

From the dataset embedded in Artificial Analysis's Intelligence Index v4.3 announcement, read on 16 September 2026. The announcement lists a selection of models, not every one Artificial Analysis tracks; Claude Sonnet 5, for example, scores 38 on its own model page and is not in this list.

Two things are worth reading off this table.

The variant matters. “Max” means maximum reasoning effort. “With fallback” on the Fable models refers to Anthropic’s refusal fallback, under which a declined request is retried on another model, and Artificial Analysis’s variant names say which: Fable 5 was run with Opus 4.8 as its fallback. The same model at a lower effort setting scores lower, and a comparison that silently mixes effort levels is not a comparison.

The top is tight. Fable 5.1 and Astra are 0.56 apart. Artificial Analysis itself describes them as tied for first. Treat anything under a point or so as a tie, and do not let a ranking position decide something a single point cannot.

How I use it in the comparisons

Every model comparison I publish from now on quotes one version only, names it, and dates the read. When a new version lands I update the numbers rather than mixing them, the way I did across Sonnet 5 against Opus 5, GPT-5.6 Sol against Sonnet 5 and Opus 5 against Fable 5.1.

The index is also not the only number that matters. Artificial Analysis runs a separate Coding Agent Index that scores models inside real coding harnesses, which I went through in Claude Code against Codex. For decisions that end in an invoice, cost per completed task usually matters more than either.

Questions people asked

What is the Artificial Analysis Intelligence Index?

A composite score built from independently run evaluations. Version 4.3 combines ten, grouped into agents 30%, coding 20%, general 30% and scientific reasoning 20%.

What changed in Intelligence Index v4.3?

Published on 7 September 2026, it upgraded Terminal-Bench from 2.1 to 4.0 and replaced τ³-Banking with AutomationBench-AA on a private test set built with Zapier. The weight on private evaluations rose from 40% to 45%; category weights did not change.

Why do AI model scores from July differ from today’s?

Because the version changed and the evaluations got harder. Claude Fable 5 read 60 in July, 62.1 on v4.1.1 and 49.70 on v4.3 with no change to the model. Compare scores only within one version.

Which AI model scores highest on the Intelligence Index v4.3?

Claude Fable 5.1 at 53.37 and GPT-6 Astra at 52.81, treated as tied for first. Then Claude Opus 5 at 50.70, Claude Fable 5 at 49.70, Muse Spark 1.3 at 48.17 and GPT-5.6 Sol at 47.06.

If you need benchmark numbers translated into a model decision for your own workload, the contact page is the fastest route to me.

About the author

Wojciech Łuszczyński

Wojciech Łuszczyński

GTM Architect and Growth Operator building AI-native revenue systems for B2B SaaS and technology companies. I connect positioning, SEO, content, paid acquisition, CRM, automation, analytics and AI workflows into practical growth infrastructure.

Newsletter

Get the next one first.

When I publish a new article on AI systems, GTM architecture, or growth operating models, you'll be the first to know.

Subscribe