THE MODEL SCOREBOARD

Big models.
Benchmarks that matter.

Compare how leading AI models perform on coding, reasoning, and real work. Explore the numbers, understand the tests, and find a starting point for your own evaluation.

16 models · 4 benchmarksSnapshot checked Source: Artificial Analysis
Compare model pricing & context windows
THE BENCHMARK

Artificial Analysis Intelligence Index v4.3

A composite of ten evaluations spanning reasoning, coding, and professional tasks. Index points are not a percentage of tasks passed.

View results at Artificial Analysis
Leading this selectionHigher is better · index points
  1. Claude Fable 5.153
  2. GPT-6 Astra53
  3. Claude Opus 551
  4. Muse Spark 1.348
  5. GPT-5.6 Sol47
Bars are relative to the highest score shown.
16 of 16 modelsRanked by overall · within the filtered selection
AI model benchmarks. Ranks apply to the selected benchmark and filtered models. Equal displayed scores share a rank; the next distinct score receives the next rank without gaps.
RankModel & evaluation settings
1TiedClaude Fable 5.1 — model sourceAnthropicAdaptive reasoning · xhigh effort · default fallback5355.1%58.7%1745
1TiedGPT-6 Astra — model sourceOpenAIReasoning · max effort5359.1%54.7%1580
2Claude Opus 5 — model sourceAnthropicAdaptive reasoning · max effort5149.0%54.9%1735
3Muse Spark 1.3 — model sourceMetaReasoning · max effort4833.3%48.7%1703
4GPT-5.6 Sol — model sourceOpenAIReasoning · max effort4739.9%49.5%1624
5GLM-5.3 — model sourceZ AIReasoning · max effort4541.9%42.3%1676
6TiedGrok 4.6 — model sourceSpaceXAIReasoning · high effort4421.2%42.9%1643
6TiedKimi K3 — model sourceKimiReasoning · max effort4412.6%46.9%1584
7GPT-5.6 Terra — model sourceOpenAIReasoning · max effort4235.4%42.9%1477
8Gemini 3.8 Flash — model sourceGoogleReasoning · high effort4119.7%47.8%1464
9Qwen3.8 2.4T A95B — model sourceAlibabaReasoning · source configuration4011.1%42.4%1628
10DeepSeek V4 Pro 0813 — model sourceDeepSeek0813 release · reasoning · max effort3614.1%41.0%1493
11MiniMax-M3 — model sourceMiniMaxReasoning · source configuration302.0%39.0%1304
12Nemotron 3 Ultra — model sourceNVIDIA550B A55B · source configuration230.5%28.4%1091
13Mistral Medium 3.5 — model sourceMistralSource configuration150.0%13.8%875
14gpt-oss-120b — model sourceOpenAIReasoning · high effort120.0%19.6%745

Ranks recalculate for the selected benchmark and filters. Equal displayed scores are marked “Tied” and share a rank, with no skipped ranks (1, 1, 2). “Not available” means no verified score in this snapshot; 0.0% is a reported score. On smaller screens, scroll the table sideways.