ModelsCompareBest forBenchmarksStatusPricingAPI

What is MMLU?

Massive Multitask Language Understanding: multiple-choice questions across 57 subjects (STEM, humanities, law, and more) testing broad knowledge and reasoning.

MMLU scores by model

1
DeepSeek R1DeepSeek90.8%
2
GPT-4.1OpenAI90.2%
3
Dola Seed 2.0 ProByteDance Seed90.1%
4
LongCat Flash ChatMeituan LongCat89.71%
5
DeepSeek V3DeepSeek88.5%
7
GPT-4.1 miniOpenAI87.5%
10
Qwen3.6 27BQwen86.2%
11
Qwen3.5 27BQwen86.1%
12
Amazon Nova ProAmazon85.9%
13
GPT-4oOpenAI85.7%
14
Command ACohere85.5%
15
Gemma 4 31BGoogle85.2%
16
17
DeepSeek V3.2DeepSeek85%
18
MAI-Thinking-1Microsoft85%
19
Phi-4Microsoft84.8%
20
ERNIE 5.0Baidu83.8%
22
Gemma 4 26B A4BGoogle82.6%
23
GPT-4o miniOpenAI82%
25
27
28
GPT-4.1 nanoOpenAI80.1%
30
Mistral Small 4Mistral AI78%
32
33
Phi-4 ReasoningMicrosoft74.3%
34
Gemma 4 E4BGoogle69.4%
36
Command R7BCohere65.2%
37
Gemma 4 E2BGoogle60%