Big models.
Benchmarks that matter.
Compare how leading AI models perform on coding, reasoning, and real work. Explore the numbers, understand the tests, and find a starting point for your own evaluation.
16 models · 4 benchmarksSnapshot checked Source: Artificial Analysis
Compare model pricing & context windows THE BENCHMARK
Artificial Analysis Intelligence Index v4.3
A composite of ten evaluations spanning reasoning, coding, and professional tasks. Index points are not a percentage of tasks passed.
View results at Artificial AnalysisLeading this selectionHigher is better · index points
- Claude Fable 5.153
- GPT-6 Astra53
- Claude Opus 551
- Muse Spark 1.348
- GPT-5.6 Sol47
16 of 16 modelsRanked by overall · within the filtered selection
| Rank | Model & evaluation settings | ||||
|---|---|---|---|---|---|
| 1Tied | Claude Fable 5.1 — model sourceAnthropicAdaptive reasoning · xhigh effort · default fallback | 53 | 55.1% | 58.7% | 1745 |
| 1Tied | GPT-6 Astra — model sourceOpenAIReasoning · max effort | 53 | 59.1% | 54.7% | 1580 |
| 2 | Claude Opus 5 — model sourceAnthropicAdaptive reasoning · max effort | 51 | 49.0% | 54.9% | 1735 |
| 3 | Muse Spark 1.3 — model sourceMetaReasoning · max effort | 48 | 33.3% | 48.7% | 1703 |
| 4 | GPT-5.6 Sol — model sourceOpenAIReasoning · max effort | 47 | 39.9% | 49.5% | 1624 |
| 5 | GLM-5.3 — model sourceZ AIReasoning · max effort | 45 | 41.9% | 42.3% | 1676 |
| 6Tied | Grok 4.6 — model sourceSpaceXAIReasoning · high effort | 44 | 21.2% | 42.9% | 1643 |
| 6Tied | Kimi K3 — model sourceKimiReasoning · max effort | 44 | 12.6% | 46.9% | 1584 |
| 7 | GPT-5.6 Terra — model sourceOpenAIReasoning · max effort | 42 | 35.4% | 42.9% | 1477 |
| 8 | Gemini 3.8 Flash — model sourceGoogleReasoning · high effort | 41 | 19.7% | 47.8% | 1464 |
| 9 | Qwen3.8 2.4T A95B — model sourceAlibabaReasoning · source configuration | 40 | 11.1% | 42.4% | 1628 |
| 10 | DeepSeek V4 Pro 0813 — model sourceDeepSeek0813 release · reasoning · max effort | 36 | 14.1% | 41.0% | 1493 |
| 11 | MiniMax-M3 — model sourceMiniMaxReasoning · source configuration | 30 | 2.0% | 39.0% | 1304 |
| 12 | Nemotron 3 Ultra — model sourceNVIDIA550B A55B · source configuration | 23 | 0.5% | 28.4% | 1091 |
| 13 | Mistral Medium 3.5 — model sourceMistralSource configuration | 15 | 0.0% | 13.8% | 875 |
| 14 | gpt-oss-120b — model sourceOpenAIReasoning · high effort | 12 | 0.0% | 19.6% | 745 |
Ranks recalculate for the selected benchmark and filters. Equal displayed scores are marked “Tied” and share a rank, with no skipped ranks (1, 1, 2). “Not available” means no verified score in this snapshot; 0.0% is a reported score. On smaller screens, scroll the table sideways.