Compare / head-to-head

Grok 4vsKimi K2 Thinking

Grok 4 leads 3 of 3 shared benchmarks. Both offer a 256K-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
3 – 0
Grok 4 leads
Cheaper per token
—
list price, input + output
Larger context
Tie
both 256K tokens
Providers
2 providers
xAI · Moonshot AI
Grok 4
xAI · released 2025-07-09
textvision
Context
256K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
17 · 8 core
Kimi K2 Thinking
Moonshot AI · released 2025-11-06
text
Context
256K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
17 · 9 core
Quality

Benchmark matrix

BenchmarkGrok 4Kimi K2 ThinkingΔEdge
Reported by both · 3
GPQA (no tools)87.5% ↗84.5% ↗+3 ptGrok 4
SimpleBench (SimpleBench)60.5% ↗39.6% ↗+20.9 ptGrok 4
WeirdML (Håvard Tveit Ihle)45.7% ↗42.8% ↗+2.9 ptGrok 4
Only Grok 4 reports · 14
GPQA Diamond87% ↗epoch runnot reported——
AIME 202591.7% ↗not reported——
Humanity's Last Exam (no tools)25.4% ↗not reported——
LMArena Elo1410.8 ↗not reported——
ARC-AGI-215.9% ↗not reported——
BALROG (BALROG)43.6% ↗not reported——
Chess Puzzles (Epoch AI run)28% ↗epoch runnot reported——
HMMT 2025 (no tools)90% ↗not reported——
Humanity's Last Exam (with tools)38.6% ↗not reported——
LiveCodeBench (Jan - May)79% ↗not reported——
LMArena Vision (LMArena)1184.2 ↗not reported——
OTIS Mock AIME 2024-2025 (Epoch AI run)84% ↗epoch runnot reported——
SAGE (Vals AI)25.1% ↗not reported——
USAMO 2025 (no tools)37.5% ↗not reported——
Only Kimi K2 Thinking reports · 14
AIME25 (no tools)not reported94.5% ↗——
BrowseComp (w/ tools)not reported60.2% ↗——
BrowseComp-ZH (w/ tools)not reported62.3% ↗——
HLE (Text-only) (no tools)not reported23.9% ↗——
HLE (Text-only) (w/ tools)not reported44.9% ↗——
HMMT25 (no tools)not reported89.4% ↗——
IMO-AnswerBench (no tools)not reported78.6% ↗——
LiveCodeBenchV6 (no tools)not reported83.1% ↗——
MMLU-Pro (no tools)not reported84.6% ↗——
Multi-SWE-bench (w/ tools)not reported41.9% ↗——
SciCode (no tools)not reported44.8% ↗——
SWE-bench Multilingual (w/ tools)not reported61.1% ↗——
SWE-bench Verified (w/ tools)not reported71.3% ↗——
Terminal-Bench (w/ simulated tools (JSON))not reported47.1% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Grok 4 minus Kimi K2 Thinking in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGrok 4Kimi K2 ThinkingEdge
Context window256K tokens256K tokensTie
Max output———
Input price / 1M———
Output price / 1M———
Cached input / 1M———
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
———
Modalitiestext · visiontextGrok 4
Released2025-07-092025-11-06—
Cited benchmark scores1717—
Reliability

Provider status

All providers →
More matchups

Kimi K2 Thinking vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Grok 4 and Kimi K2 Thinking on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.