ModelsCompareBest forBenchmarksStatusPricingAPI

Best LLM for reasoning

Ranked by GPQA Diamond, the most widely reported graduate-level reasoning benchmark. Every score links to its source.

1
GPT-6 AstraOpenAI96%
2
3
4
GPT-5.6 SolOpenAI94.6%
5
GPT-5.4 ProOpenAI94.4%
7
Claude Opus 4.7Anthropic94.2%
8
9
10
Claude Opus 5Anthropic93.9%
11
Claude Opus 4.8Anthropic93.6%
12
Kimi K3Moonshot AI93.5%
13
Grok 4.5xAI93.4%
14
GPT-5.2 ProOpenAI93.2%
15
GPT-5.6 TerraOpenAI92.9%
16
GPT-5.4OpenAI92.8%
17
18
Qwen3.8-MaxQwen92.6%
20
GPT-5.2OpenAI92.4%
21
Dola Seed 2.0 ProByteDance Seed92.4%
22
GPT-5.6 LunaOpenAI92.3%
23
GLM-5.2Z.ai91.9%
25
Claude Opus 4.6Anthropic91.3%
26
27
DeepSeek V4 ProDeepSeek90.9%
28
MiniMax M3MiniMax90.9%
29
GLM-5.3Z.ai90.9%
30
Kimi K2.6Moonshot AI90.8%

Try any of these through one API with automatic failover: Respan gateway.