ModelsCompareBest forBenchmarksStatusPricingAPI

What is AIME?

Problems from the American Invitational Mathematics Examination, a competition-level math benchmark used to test multi-step quantitative reasoning.

AIME scores by model

1
GPT-5.2OpenAI100%
2
GPT-5.2 ProOpenAI100%
3
Dola Seed 2.0 ProByteDance Seed99%
4
Step 3.5 FlashStepFun98.3%
5
MAI-Thinking-1Microsoft97%
6
GLM-5Z.ai96.7%
8
GPT-5OpenAI95%
9
GPT-5.1OpenAI94.2%
10
DeepSeek V3.2DeepSeek94.2%
11
GLM-4.5Z.ai93.3%
13
GLM-4.6Z.ai91.7%
16
gpt-oss-120bOpenAI90%
17
Command A+Cohere90%
18
gpt-oss-20bOpenAI89.2%
19
o3OpenAI89.2%
21
ERNIE 5.0Baidu89.06%
22
Gemini 2.5 ProGoogle88.3%
23
GPT-5 miniOpenAI87.5%
24
GPT-5 nanoOpenAI85%
25
Ministral 3 14BMistral AI85%
26
Claude Sonnet 4.5Anthropic84.2%
27
GLM-4.5-AirZ.ai83.3%
28
29
Ministral 3 8BMistral AI78.7%
30
Ministral 3 3BMistral AI72.1%
32
DeepSeek R1DeepSeek70%
34
Phi-4 ReasoningMicrosoft62.9%
35
LongCat Flash ChatMeituan LongCat61.25%
37
Claude 3.7 SonnetAnthropic49.2%
38
DeepSeek V3DeepSeek25%
39
GPT-4oOpenAI11.7%
40
Claude 3.5 SonnetAnthropic3.3%