Compare / head-to-head
Gemini 3.1 Pro Previewvs
Kimi K2.5
Gemini 3.1 Pro Preview leads 10 of 11 shared benchmarks. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
10 – 1
Gemini 3.1 Pro Preview leads
Cheaper per token
—
list price, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Google · Moonshot AI
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Kimi K2.5
Moonshot AI · released 2026-01-27
textvisionvideo
- Context
- 256K
- Max out
- —
- Input /1M
- —
- Output /1M
- —
- Cached /1M
- —
- Scores
- 26 · 14 core
Quality
Benchmark matrix
11 shared · 25 only Gemini 3.1 Pro Preview · 15 only Kimi K2.5| Benchmark | Gemini 3.1 Pro Preview | Kimi K2.5 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 11 | ||||
| LMArena Elo | 1480.1 ↗ | 1450.1 ↗ | +30 | Gemini 3.1 Pro Preview |
| LMArena Vision (LMArena) | 1279.3 ↗ | 1252.2 ↗ | +27.1 | Gemini 3.1 Pro Preview |
| LMArena WebDev (LMArena) | 1446.2 ↗ | 1435.8 ↗ | +10.4 | Gemini 3.1 Pro Preview |
| SAGE (Vals AI) | 48.7% ↗ | 49.9% ↗ | -1.2 pt | Kimi K2.5 |
| SimpleBench (SimpleBench) | 79.6% ↗ | 46.8% ↗ | +32.8 pt | Gemini 3.1 Pro Preview |
| SimpleQA Verified | 73.5% ↗epoch run | 34.3% ↗epoch run | +39.2 pt | Gemini 3.1 Pro Preview |
| SWE-Bench Pro | 54.2% ↗ | 50.7% ↗ | +3.5 pt | Gemini 3.1 Pro Preview |
| SWE-bench Verified (Epoch AI run) | 75.6% ↗epoch run | 73.8% ↗epoch run | +1.8 pt | Gemini 3.1 Pro Preview |
| Toolathlon-Verified (HKUST) | 61.1% ↗ | 33% ↗ | +28.1 pt | Gemini 3.1 Pro Preview |
| Vending-Bench 2 (Andon Labs) | 3774.25 ↗ | 1198.46 ↗ | +2575.8 | Gemini 3.1 Pro Preview |
| WeirdML (Håvard Tveit Ihle) | 72.1% ↗ | 45.6% ↗ | +26.5 pt | Gemini 3.1 Pro Preview |
| Only Gemini 3.1 Pro Preview reports · 25 | ||||
| GPQA Diamond | 94.3% ↗ | not reported | — | — |
| AIME 2026 | 98.3% ↗matharena ⚠ | not reported | — | — |
| APEX-Agents (Mercor) | 35.3% ↗ | not reported | — | — |
| ARC-AGI-2 | 77.1% ↗ | not reported | — | — |
| BALROG (BALROG) | 57% ↗ | not reported | — | — |
| Chess Puzzles (Epoch AI run) | 55% ↗epoch run | not reported | — | — |
| DeepSWE v1.1 (Datacurve) | 11.7% ↗ | not reported | — | — |
| EBR-bench (Epoch AI run) | 14.3% ↗epoch run | not reported | — | — |
| FrontierMath Tier 4 v2 (Epoch AI run) | 26.8% ↗epoch run | not reported | — | — |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 59.6% ↗epoch run | not reported | — | — |
| Furniture Assembly (Epoch AI run) | 26.7% ↗epoch run | not reported | — | — |
| GSO Opt@1 (GSO) | 21.6% ↗ | not reported | — | — |
| LiveBench Agentic Coding (LiveBench) | 44.1% ↗ | not reported | — | — |
| LiveBench Coding (LiveBench) | 76.5% ↗ | not reported | — | — |
| LiveBench Data Analysis (LiveBench) | 78.5% ↗ | not reported | — | — |
| LiveBench Instruction Following (LiveBench) | 79.1% ↗ | not reported | — | — |
| LiveBench Language (LiveBench) | 85.4% ↗ | not reported | — | — |
| LiveBench Mathematics (LiveBench) | 91% ↗ | not reported | — | — |
| LiveBench Reasoning (LiveBench) | 84% ↗ | not reported | — | — |
| LMArena Agent (LMArena) | -0.0771 ↗ | not reported | — | — |
| MirrorCode (Epoch AI run) | 8.9% ↗epoch run | not reported | — | — |
| Mystery Game Puzzles (Epoch AI run) | 34% ↗epoch run | not reported | — | — |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 95.6% ↗epoch run | not reported | — | — |
| tau2-bench Banking Knowledge (Sierra) | 26% ↗ | not reported | — | — |
| Terminal-Bench 4.0 (Vals AI) | 2.5% ↗ | not reported | — | — |
| Only Kimi K2.5 reports · 15 | ||||
| SWE-bench Verified | not reported | 76.8% ↗ | — | — |
| MMLU-Pro | not reported | 87.1% ↗ | — | — |
| AIME 2025 | not reported | 96.1% ↗ | — | — |
| BrowseComp | not reported | 60.6% ↗ | — | — |
| GPQA-Diamond | not reported | 87.6% ↗ | — | — |
| HLE-Full | not reported | 30.1% ↗ | — | — |
| HLE-Full (w/ tools) | not reported | 50.2% ↗ | — | — |
| HMMT 2025 (Feb) | not reported | 95.4% ↗ | — | — |
| IMO-AnswerBench | not reported | 81.8% ↗ | — | — |
| LiveCodeBench (v6) | not reported | 85% ↗ | — | — |
| MMMU-Pro | not reported | 78.5% ↗ | — | — |
| OSWorld-Verified (XLANG) | not reported | 63.3% ↗ | — | — |
| SWE-Bench Multilingual | not reported | 73% ↗ | — | — |
| Terminal Bench 2.0 | not reported | 50.8% ↗ | — | — |
| VideoMMMU | not reported | 86.6% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.1 Pro Preview minus Kimi K2.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Gemini 3.1 Pro Preview | Kimi K2.5 | Edge |
|---|---|---|---|
| Context window | 1.0M tokens | 256K tokens | Gemini 3.1 Pro Preview |
| Max output | 66K tokens | — | — |
| Input price / 1M | $2 ↗ | — | — |
| Output price / 1M | $12 ↗ | — | — |
| Cached input / 1M | $0.2 ↗ | — | — |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $14 | — | — |
| Modalities | text · vision · audio | text · vision · video | Tie |
| Released | 2026-02-19 | 2026-01-27 | — |
| Cited benchmark scores | 43 | 26 | — |
Reliability
Provider status
Live from /statusMore matchups
Gemini 3.1 Pro Preview vs …
Models sharing the most benchmarksMore matchups
Kimi K2.5 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Gemini 3.1 Pro Preview and Kimi K2.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.