Compare / head-to-head
Claude Opus 4.5vs
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview leads 18 of 26 shared benchmarks. Gemini 3.1 Pro Preview is 2.1x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
7 – 18
Gemini 3.1 Pro Preview leads · 1 tied
Cheaper per token
Gemini 3.1 Pro Preview
2.1x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Quality
Benchmark matrix
26 shared · 12 only Claude Opus 4.5 · 10 only Gemini 3.1 Pro Preview| Benchmark | Claude Opus 4.5 | Gemini 3.1 Pro Preview | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 26 | ||||
| GPQA Diamond | 87% ↗ | 94.3% ↗ | -7.3 pt | Gemini 3.1 Pro Preview |
| LMArena Elo | 1473.6 ↗ | 1480.1 ↗ | -6.5 | Gemini 3.1 Pro Preview |
| BALROG (BALROG) | 43.5% ↗ | 57% ↗ | -13.5 pt | Gemini 3.1 Pro Preview |
| Chess Puzzles (Epoch AI run) | 12% ↗epoch run | 55% ↗epoch run | -43 pt | Gemini 3.1 Pro Preview |
| EBR-bench (Epoch AI run) | 14.3% ↗epoch run | 14.3% ↗epoch run | 0 | Tie |
| FrontierMath Tier 4 v2 (Epoch AI run) | 4.9% ↗epoch run | 26.8% ↗epoch run | -21.9 pt | Gemini 3.1 Pro Preview |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 34.4% ↗epoch run | 59.6% ↗epoch run | -25.2 pt | Gemini 3.1 Pro Preview |
| Furniture Assembly (Epoch AI run) | 28.3% ↗epoch run | 26.7% ↗epoch run | +1.6 pt | Claude Opus 4.5 |
| GSO Opt@1 (GSO) | 24.5% ↗ | 21.6% ↗ | +2.9 pt | Claude Opus 4.5 |
| LiveBench Agentic Coding (LiveBench) | 39.7% ↗ | 44.1% ↗ | -4.4 pt | Gemini 3.1 Pro Preview |
| LiveBench Coding (LiveBench) | 79.7% ↗ | 76.5% ↗ | +3.2 pt | Claude Opus 4.5 |
| LiveBench Data Analysis (LiveBench) | 74.4% ↗ | 78.5% ↗ | -4.1 pt | Gemini 3.1 Pro Preview |
| LiveBench Instruction Following (LiveBench) | 62.5% ↗ | 79.1% ↗ | -16.6 pt | Gemini 3.1 Pro Preview |
| LiveBench Language (LiveBench) | 81.3% ↗ | 85.4% ↗ | -4.1 pt | Gemini 3.1 Pro Preview |
| LiveBench Mathematics (LiveBench) | 90.4% ↗ | 91% ↗ | -0.6 pt | Gemini 3.1 Pro Preview |
| LiveBench Reasoning (LiveBench) | 80.1% ↗ | 84% ↗ | -3.9 pt | Gemini 3.1 Pro Preview |
| LMArena WebDev (LMArena) | 1493.3 ↗ | 1446.2 ↗ | +47.1 | Claude Opus 4.5 |
| Mystery Game Puzzles (Epoch AI run) | 22% ↗epoch run | 34% ↗epoch run | -12 pt | Gemini 3.1 Pro Preview |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 86.1% ↗epoch run | 95.6% ↗epoch run | -9.5 pt | Gemini 3.1 Pro Preview |
| SAGE (Vals AI) | 52.1% ↗ | 48.7% ↗ | +3.4 pt | Claude Opus 4.5 |
| SimpleBench (SimpleBench) | 62% ↗ | 79.6% ↗ | -17.6 pt | Gemini 3.1 Pro Preview |
| SimpleQA Verified | 45.7% ↗epoch run | 73.5% ↗epoch run | -27.8 pt | Gemini 3.1 Pro Preview |
| SWE-bench Verified (Epoch AI run) | 76.7% ↗epoch run | 75.6% ↗epoch run | +1.1 pt | Claude Opus 4.5 |
| tau2-bench Banking Knowledge (Sierra) | 24.7% ↗ | 26% ↗ | -1.3 pt | Gemini 3.1 Pro Preview |
| Vending-Bench 2 (Andon Labs) | 4967.06 ↗ | 3774.25 ↗ | +1192.8 | Claude Opus 4.5 |
| WeirdML (Håvard Tveit Ihle) | 63.7% ↗ | 72.1% ↗ | -8.4 pt | Gemini 3.1 Pro Preview |
| Only Claude Opus 4.5 reports · 12 | ||||
| SWE-bench Verified | 80.9% ↗ | not reported | — | — |
| ARC-AGI-2 (Verified) | 37.6% ↗ | not reported | — | — |
| MCP Atlas | 62.3% ↗ | not reported | — | — |
| MMMLU | 90.8% ↗ | not reported | — | — |
| MMMU (validation) | 80.7% ↗ | not reported | — | — |
| OSWorld | 66.3% ↗ | not reported | — | — |
| tau2-bench Airline (Sierra) | 84% ↗ | not reported | — | — |
| tau2-bench Retail (Sierra) | 79.6% ↗ | not reported | — | — |
| tau2-bench Telecom (Sierra) | 92.3% ↗ | not reported | — | — |
| Terminal-Bench 2.0 | 59.3% ↗ | not reported | — | — |
| τ2-bench (Retail) | 88.9% ↗ | not reported | — | — |
| τ2-bench (Telecom) | 98.2% ↗ | not reported | — | — |
| Only Gemini 3.1 Pro Preview reports · 10 | ||||
| AIME 2026 | not reported | 98.3% ↗matharena ⚠ | — | — |
| APEX-Agents (Mercor) | not reported | 35.3% ↗ | — | — |
| ARC-AGI-2 | not reported | 77.1% ↗ | — | — |
| DeepSWE v1.1 (Datacurve) | not reported | 11.7% ↗ | — | — |
| LMArena Agent (LMArena) | not reported | -0.0771 ↗ | — | — |
| LMArena Vision (LMArena) | not reported | 1279.3 ↗ | — | — |
| MirrorCode (Epoch AI run) | not reported | 8.9% ↗epoch run | — | — |
| SWE-Bench Pro | not reported | 54.2% ↗ | — | — |
| Terminal-Bench 4.0 (Vals AI) | not reported | 2.5% ↗ | — | — |
| Toolathlon-Verified (HKUST) | not reported | 61.1% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Opus 4.5 | Gemini 3.1 Pro Preview | Edge |
|---|---|---|---|
| Context window | 200K tokens | 1.0M tokens | Gemini 3.1 Pro Preview |
| Max output | 64K tokens | 66K tokens | Gemini 3.1 Pro Preview |
| Input price / 1M | $5 ↗ | $2 ↗ | Gemini 3.1 Pro Preview |
| Output price / 1M | $25 ↗ | $12 ↗ | Gemini 3.1 Pro Preview |
| Cached input / 1M | $0.5 ↗ | $0.2 ↗ | Gemini 3.1 Pro Preview |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $30 | $14 | Gemini 3.1 Pro Preview |
| Modalities | text · vision | text · vision · audio | Gemini 3.1 Pro Preview |
| Released | 2025-11-24 | 2026-02-19 | — |
| Cited benchmark scores | 39 | 43 | — |
Reliability
Provider status
Live from /statusMore matchups
Claude Opus 4.5 vs …
Models sharing the most benchmarksMore matchups
Gemini 3.1 Pro Preview vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Opus 4.5 and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.