What is Humanity's Last Exam?
A very hard, broad exam of expert-level questions across many fields, designed to remain difficult for frontier models.
Humanity's Last Exam scores by model
1
Claude Fable 5.1Anthropic60.9% ↗
2
Claude Opus 5Anthropic56.6% ↗
3
Claude Opus 4.8Anthropic49.8% ↗
4
Claude Opus 4.7Anthropic46.9% ↗
6
Claude Sonnet 5Anthropic43.2% ↗
7
GPT-5.5 ProOpenAI43.1% ↗
8
GPT-5.4 ProOpenAI42.7% ↗
11
GPT-5.2 ProOpenAI36.6% ↗
12
Gemini 3 Flash PreviewGoogle33.7% ↗
13
Dola Seed 2.0 ProByteDance Seed33.3% ↗
14
Claude Sonnet 4.6Anthropic33.2% ↗
15
Solar Open 2 250B-A15BUpstage28.8% ↗
16
NVIDIA Nemotron 3 Ultra 550B-A55B NVFP4NVIDIA26.1% ↗
17
DeepSeek V3.2DeepSeek25.1% ↗
18
Gemini 2.5 ProGoogle21.6% ↗
19
Grok 4 Fast ReasoningxAI20% ↗
20
Gemma 4 31BGoogle19.5% ↗
21
NVIDIA Nemotron 3 Super 120B-A12BNVIDIA18.26% ↗
22
Claude Sonnet 4.5Anthropic17.7% ↗
23
Gemini 3.1 Flash-LiteGoogle16% ↗
24
Gemini 2.5 FlashGoogle11% ↗
25
NVIDIA Nemotron 3 Nano 30B-A3B BF16NVIDIA10.6% ↗
26
NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4NVIDIA10.47% ↗
27
Gemma 4 26B A4BGoogle8.7% ↗
28
Gemini 2.5 Flash-LiteGoogle6.9% ↗
29
Gemma 4 12B UnifiedGoogle5.2% ↗