ModelsCompareBest forBenchmarksStatusPricingAPI
Compare / head-to-head

Claude Opus 4.6vsClaude Sonnet 4.5

Claude Opus 4.6 leads 16 of 16 shared benchmarks. Claude Sonnet 4.5 is 1.7x cheaper per token. Claude Opus 4.6 has the larger context window (1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
16 – 0
Claude Opus 4.6 leads
Cheaper per token
Claude Sonnet 4.5
1.7x cheaper, input + output
Larger context
Claude Opus 4.6
1M tokens
Provider uptime (30d)
100%
Anthropic
Claude Opus 4.6
Anthropic · released 2026-02-05
textvision
Context
1M
Max out
128K
Input /1M
$5
Output /1M
$25
Cached /1M
$0.5
Scores
19 · 9 core
Claude Sonnet 4.5
Anthropic · released 2025-09-29
textvision
Context
200K
Max out
64K
Input /1M
$3
Output /1M
$15
Cached /1M
$0.3
Scores
21 · 10 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.6Claude Sonnet 4.5ΔEdge
Reported by both · 16
SWE-bench Verified80.8% 77.2% +3.6 ptClaude Opus 4.6
GPQA Diamond91.3% 83.4% +7.9 ptClaude Opus 4.6
ARC-AGI-2 (Verified)68.8% 13.6% +55.2 ptClaude Opus 4.6
FrontierMath Tier 4 v2 (Epoch AI run)26.8% epoch2.4% epoch+24.4 ptClaude Opus 4.6
FrontierMath Tiers 1-3 v2 (Epoch AI run)66% epoch23.9% epoch+42.1 ptClaude Opus 4.6
MCP Atlas59.5% 43.8% +15.7 ptClaude Opus 4.6
MMMLU91.1% 89.5% +1.6 ptClaude Opus 4.6
MMMU-Pro (no tools)73.9% 63.4% +10.5 ptClaude Opus 4.6
MMMU-Pro (with tools)77.3% 68.9% +8.4 ptClaude Opus 4.6
OSWorld-Verified72.7% 61.4% +11.3 ptClaude Opus 4.6
OTIS Mock AIME 2024-2025 (Epoch AI run)94.4% epoch77.8% epoch+16.6 ptClaude Opus 4.6
SimpleQA Verified47% epoch30.7% epoch+16.3 ptClaude Opus 4.6
SWE-bench Verified (Epoch AI run)78.7% epoch71.3% epoch+7.4 ptClaude Opus 4.6
Tau2-bench Retail91.9% 86.2% +5.7 ptClaude Opus 4.6
Tau2-bench Telecom99.3% 98% +1.3 ptClaude Opus 4.6
Terminal-Bench 2.065.4% 51% +14.4 ptClaude Opus 4.6
Only Claude Opus 4.6 reports · 2
AIME 202696.7% matharenanot reported
LMArena Elo1497.5 not reported
Only Claude Sonnet 4.5 reports · 4
AIME 2025not reported84.2% matharena
Humanity's Last Exam (no tools)not reported17.7%
GDPval-AAnot reported1276
Humanity's Last Exam (with tools)not reported33.6%
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.6 minus Claude Sonnet 4.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.6Claude Sonnet 4.5Edge
Context window1M tokens200K tokensClaude Opus 4.6
Max output128K tokens64K tokensClaude Opus 4.6
Input price / 1M$5 $3 Claude Sonnet 4.5
Output price / 1M$25 $15 Claude Sonnet 4.5
Cached input / 1M$0.5 $0.3 Claude Sonnet 4.5
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$18Claude Sonnet 4.5
Modalitiestext · visiontext · visionTie
Released2026-02-052025-09-29
Cited benchmark scores1921
Reliability

Anthropic status

All providers
More matchups

Claude Opus 4.6 vs …

More matchups

Claude Sonnet 4.5 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.6 and Claude Sonnet 4.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.