What is MMLU?
Massive Multitask Language Understanding: multiple-choice questions across 57 subjects (STEM, humanities, law, and more) testing broad knowledge and reasoning.
MMLU scores by model
1
DeepSeek R1DeepSeek90.8% ↗
3
Dola Seed 2.0 ProByteDance Seed90.1% ↗
4
LongCat Flash ChatMeituan LongCat89.71% ↗
5
DeepSeek V3DeepSeek88.5% ↗
6
Qwen3.5 397B-A17BQwen87.8% ↗
7
GPT-4.1 miniOpenAI87.5% ↗
8
Amazon Nova PremierAmazon87.4% ↗
9
Solar Open 2 250B-A15BUpstage86.2% ↗
10
Qwen3.6 27BQwen86.2% ↗
11
Qwen3.5 27BQwen86.1% ↗
12
Amazon Nova ProAmazon85.9% ↗
15
Gemma 4 31BGoogle85.2% ↗
16
Qwen3.6 35B-A3BQwen85.2% ↗
17
DeepSeek V3.2DeepSeek85% ↗
18
MAI-Thinking-1Microsoft85% ↗
21
NVIDIA Nemotron 3 Super 120B-A12BNVIDIA83.73% ↗
22
Gemma 4 26B A4BGoogle82.6% ↗
23
GPT-4o miniOpenAI82% ↗
24
NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4NVIDIA81.62% ↗
25
Amazon Nova 2 LiteAmazon80.9% ↗
26
Llama 4 MaverickMeta80.5% ↗
27
Amazon Nova LiteAmazon80.5% ↗
28
GPT-4.1 nanoOpenAI80.1% ↗
29
NVIDIA Nemotron 3 Nano 30B-A3B BF16NVIDIA78.3% ↗
30
Mistral Small 4Mistral AI78% ↗
31
Gemma 4 12B UnifiedGoogle77.2% ↗
32
Llama 4 ScoutMeta74.3% ↗
33
Phi-4 ReasoningMicrosoft74.3% ↗
34
Gemma 4 E4BGoogle69.4% ↗
35
Llama 3.3 70B InstructMeta68.9% ↗
36
Command R7BCohere65.2% ↗
37
Gemma 4 E2BGoogle60% ↗