Compare / head-to-head

Claude Opus 4.5vsClaude Opus 4.7

Claude Opus 4.7 leads 24 of 26 shared benchmarks. Both list the same combined token price. Claude Opus 4.7 has the larger context window (1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
2 – 24
Claude Opus 4.7 leads
Cheaper per token
Tie
list price, input + output
Larger context
Claude Opus 4.7
1M tokens
Providers
Anthropic
same provider
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Context
200K
Max out
64K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
39 · 9 core
Claude Opus 4.7
Anthropic · released 2026-04-16
textvision
Context
1M
Max out
128K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
46 · 16 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.5Claude Opus 4.7ΔEdge
Reported by both · 26
SWE-bench Verified80.9% ↗87.6% ↗-6.7 ptClaude Opus 4.7
GPQA Diamond87% ↗94.2% ↗-7.2 ptClaude Opus 4.7
LMArena Elo1473.6 ↗1483.4 ↗-9.8Claude Opus 4.7
Chess Puzzles (Epoch AI run)12% ↗epoch run30% ↗epoch run-18 ptClaude Opus 4.7
EBR-bench (Epoch AI run)14.3% ↗epoch run19% ↗epoch run-4.7 ptClaude Opus 4.7
FrontierMath Tier 4 v2 (Epoch AI run)4.9% ↗epoch run31.7% ↗epoch run-26.8 ptClaude Opus 4.7
FrontierMath Tiers 1-3 v2 (Epoch AI run)34.4% ↗epoch run70.2% ↗epoch run-35.8 ptClaude Opus 4.7
Furniture Assembly (Epoch AI run)28.3% ↗epoch run33.3% ↗epoch run-5 ptClaude Opus 4.7
GSO Opt@1 (GSO)24.5% ↗42.2% ↗-17.7 ptClaude Opus 4.7
LiveBench Agentic Coding (LiveBench)39.7% ↗50.7% ↗-11 ptClaude Opus 4.7
LiveBench Coding (LiveBench)79.7% ↗82.1% ↗-2.4 ptClaude Opus 4.7
LiveBench Data Analysis (LiveBench)74.4% ↗78.3% ↗-3.9 ptClaude Opus 4.7
LiveBench Instruction Following (LiveBench)62.5% ↗66.7% ↗-4.2 ptClaude Opus 4.7
LiveBench Language (LiveBench)81.3% ↗77.9% ↗+3.4 ptClaude Opus 4.5
LiveBench Mathematics (LiveBench)90.4% ↗92.9% ↗-2.5 ptClaude Opus 4.7
LiveBench Reasoning (LiveBench)80.1% ↗87.2% ↗-7.1 ptClaude Opus 4.7
LMArena WebDev (LMArena)1493.3 ↗1557.3 ↗-64Claude Opus 4.7
Mystery Game Puzzles (Epoch AI run)22% ↗epoch run28% ↗epoch run-6 ptClaude Opus 4.7
OTIS Mock AIME 2024-2025 (Epoch AI run)86.1% ↗epoch run97.8% ↗epoch run-11.7 ptClaude Opus 4.7
SAGE (Vals AI)52.1% ↗56.1% ↗-4 ptClaude Opus 4.7
SimpleBench (SimpleBench)62% ↗61.7% ↗+0.3 ptClaude Opus 4.5
SimpleQA Verified45.7% ↗epoch run51.7% ↗epoch run-6 ptClaude Opus 4.7
SWE-bench Verified (Epoch AI run)76.7% ↗epoch run83.5% ↗epoch run-6.8 ptClaude Opus 4.7
tau2-bench Banking Knowledge (Sierra)24.7% ↗40.2% ↗-15.5 ptClaude Opus 4.7
Vending-Bench 2 (Andon Labs)4967.06 ↗10936.76 ↗-5969.7Claude Opus 4.7
WeirdML (Håvard Tveit Ihle)63.7% ↗76.4% ↗-12.7 ptClaude Opus 4.7
Only Claude Opus 4.5 reports · 12
ARC-AGI-2 (Verified)37.6% ↗not reported——
BALROG (BALROG)43.5% ↗not reported——
MCP Atlas62.3% ↗not reported——
MMMLU90.8% ↗not reported——
MMMU (validation)80.7% ↗not reported——
OSWorld66.3% ↗not reported——
tau2-bench Airline (Sierra)84% ↗not reported——
tau2-bench Retail (Sierra)79.6% ↗not reported——
tau2-bench Telecom (Sierra)92.3% ↗not reported——
Terminal-Bench 2.059.3% ↗not reported——
τ2-bench (Retail)88.9% ↗not reported——
τ2-bench (Telecom)98.2% ↗not reported——
Only Claude Opus 4.7 reports · 12
AIME 2026not reported95.8% ↗matharena ⚠——
Humanity's Last Exam (no tools)not reported46.9% ↗——
APEX-Agents (Mercor)not reported49.2% ↗——
BrowseCompnot reported79.8% ↗——
Humanity's Last Exam (with tools)not reported54.7% ↗——
LMArena Vision (LMArena)not reported1298.1 ↗——
MirrorCode (Epoch AI run)not reported31.1% ↗epoch run——
OSWorld-Verifiednot reported82.8% ↗——
SWE-bench Multilingualnot reported80.5% ↗——
SWE-bench Multimodalnot reported34.5% ↗——
SWE-Bench Pronot reported64.3% ↗——
Terminal-Bench 2.1not reported66.1% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Claude Opus 4.7 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.5Claude Opus 4.7Edge
Context window200K tokens1M tokensClaude Opus 4.7
Max output64K tokens128K tokensClaude Opus 4.7
Input price / 1M$5 ↗$5 ↗Tie
Output price / 1M$25 ↗$25 ↗Tie
Cached input / 1M$0.5 ↗$0.5 ↗Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$30Tie
Modalitiestext · visiontext · visionTie
Released2025-11-242026-04-16—
Cited benchmark scores3946—
Reliability

Anthropic status

All providers →
More matchups

Claude Opus 4.5 vs …

More matchups

Claude Opus 4.7 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.5 and Claude Opus 4.7 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.