Compare / head-to-head

Claude Opus 4.5vsGemini 3.1 Pro Preview

Gemini 3.1 Pro Preview leads 18 of 26 shared benchmarks. Gemini 3.1 Pro Preview is 2.1x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
7 – 18
Gemini 3.1 Pro Preview leads · 1 tied
Cheaper per token
Gemini 3.1 Pro Preview
2.1x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Context
200K
Max out
64K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
39 · 9 core
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.5Gemini 3.1 Pro PreviewΔEdge
Reported by both · 26
GPQA Diamond87% ↗94.3% ↗-7.3 ptGemini 3.1 Pro Preview
LMArena Elo1473.6 ↗1480.1 ↗-6.5Gemini 3.1 Pro Preview
BALROG (BALROG)43.5% ↗57% ↗-13.5 ptGemini 3.1 Pro Preview
Chess Puzzles (Epoch AI run)12% ↗epoch run55% ↗epoch run-43 ptGemini 3.1 Pro Preview
EBR-bench (Epoch AI run)14.3% ↗epoch run14.3% ↗epoch run0Tie
FrontierMath Tier 4 v2 (Epoch AI run)4.9% ↗epoch run26.8% ↗epoch run-21.9 ptGemini 3.1 Pro Preview
FrontierMath Tiers 1-3 v2 (Epoch AI run)34.4% ↗epoch run59.6% ↗epoch run-25.2 ptGemini 3.1 Pro Preview
Furniture Assembly (Epoch AI run)28.3% ↗epoch run26.7% ↗epoch run+1.6 ptClaude Opus 4.5
GSO Opt@1 (GSO)24.5% ↗21.6% ↗+2.9 ptClaude Opus 4.5
LiveBench Agentic Coding (LiveBench)39.7% ↗44.1% ↗-4.4 ptGemini 3.1 Pro Preview
LiveBench Coding (LiveBench)79.7% ↗76.5% ↗+3.2 ptClaude Opus 4.5
LiveBench Data Analysis (LiveBench)74.4% ↗78.5% ↗-4.1 ptGemini 3.1 Pro Preview
LiveBench Instruction Following (LiveBench)62.5% ↗79.1% ↗-16.6 ptGemini 3.1 Pro Preview
LiveBench Language (LiveBench)81.3% ↗85.4% ↗-4.1 ptGemini 3.1 Pro Preview
LiveBench Mathematics (LiveBench)90.4% ↗91% ↗-0.6 ptGemini 3.1 Pro Preview
LiveBench Reasoning (LiveBench)80.1% ↗84% ↗-3.9 ptGemini 3.1 Pro Preview
LMArena WebDev (LMArena)1493.3 ↗1446.2 ↗+47.1Claude Opus 4.5
Mystery Game Puzzles (Epoch AI run)22% ↗epoch run34% ↗epoch run-12 ptGemini 3.1 Pro Preview
OTIS Mock AIME 2024-2025 (Epoch AI run)86.1% ↗epoch run95.6% ↗epoch run-9.5 ptGemini 3.1 Pro Preview
SAGE (Vals AI)52.1% ↗48.7% ↗+3.4 ptClaude Opus 4.5
SimpleBench (SimpleBench)62% ↗79.6% ↗-17.6 ptGemini 3.1 Pro Preview
SimpleQA Verified45.7% ↗epoch run73.5% ↗epoch run-27.8 ptGemini 3.1 Pro Preview
SWE-bench Verified (Epoch AI run)76.7% ↗epoch run75.6% ↗epoch run+1.1 ptClaude Opus 4.5
tau2-bench Banking Knowledge (Sierra)24.7% ↗26% ↗-1.3 ptGemini 3.1 Pro Preview
Vending-Bench 2 (Andon Labs)4967.06 ↗3774.25 ↗+1192.8Claude Opus 4.5
WeirdML (Håvard Tveit Ihle)63.7% ↗72.1% ↗-8.4 ptGemini 3.1 Pro Preview
Only Claude Opus 4.5 reports · 12
SWE-bench Verified80.9% ↗not reported——
ARC-AGI-2 (Verified)37.6% ↗not reported——
MCP Atlas62.3% ↗not reported——
MMMLU90.8% ↗not reported——
MMMU (validation)80.7% ↗not reported——
OSWorld66.3% ↗not reported——
tau2-bench Airline (Sierra)84% ↗not reported——
tau2-bench Retail (Sierra)79.6% ↗not reported——
tau2-bench Telecom (Sierra)92.3% ↗not reported——
Terminal-Bench 2.059.3% ↗not reported——
τ2-bench (Retail)88.9% ↗not reported——
τ2-bench (Telecom)98.2% ↗not reported——
Only Gemini 3.1 Pro Preview reports · 10
AIME 2026not reported98.3% ↗matharena ⚠——
APEX-Agents (Mercor)not reported35.3% ↗——
ARC-AGI-2not reported77.1% ↗——
DeepSWE v1.1 (Datacurve)not reported11.7% ↗——
LMArena Agent (LMArena)not reported-0.0771 ↗——
LMArena Vision (LMArena)not reported1279.3 ↗——
MirrorCode (Epoch AI run)not reported8.9% ↗epoch run——
SWE-Bench Pronot reported54.2% ↗——
Terminal-Bench 4.0 (Vals AI)not reported2.5% ↗——
Toolathlon-Verified (HKUST)not reported61.1% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.5Gemini 3.1 Pro PreviewEdge
Context window200K tokens1.0M tokensGemini 3.1 Pro Preview
Max output64K tokens66K tokensGemini 3.1 Pro Preview
Input price / 1M$5 ↗$2 ↗Gemini 3.1 Pro Preview
Output price / 1M$25 ↗$12 ↗Gemini 3.1 Pro Preview
Cached input / 1M$0.5 ↗$0.2 ↗Gemini 3.1 Pro Preview
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$14Gemini 3.1 Pro Preview
Modalitiestext · visiontext · vision · audioGemini 3.1 Pro Preview
Released2025-11-242026-02-19—
Cited benchmark scores3943—
Reliability

Provider status

All providers →
More matchups

Claude Opus 4.5 vs …

More matchups

Gemini 3.1 Pro Preview vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.5 and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.