Compare / head-to-head
Claude Opus 4.5vs
Claude Sonnet 5.5
Claude Sonnet 5.5 leads 16 of 18 shared benchmarks. Claude Sonnet 5.5 is 2.5x cheaper per token. Claude Sonnet 5.5 has the larger context window (1M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
2 – 16
Claude Sonnet 5.5 leads
Cheaper per token
Claude Sonnet 5.5
2.5x cheaper, input + output
Larger context
Claude Sonnet 5.5
1M tokens
Providers
Anthropic
same provider
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Claude Sonnet 5.5
Anthropic · released 2026-09-28
textvision
Quality
Benchmark matrix
18 shared · 20 only Claude Opus 4.5 · 11 only Claude Sonnet 5.5| Benchmark | Claude Opus 4.5 | Claude Sonnet 5.5 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 18 | ||||
| GPQA Diamond | 87% ↗ | 95.6% ↗epoch run | -8.6 pt | Claude Sonnet 5.5 |
| LMArena Elo | 1473.6 ↗ | 1471 ↗ | +2.6 | Claude Opus 4.5 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 4.9% ↗epoch run | 80.5% ↗epoch run | -75.6 pt | Claude Sonnet 5.5 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 34.4% ↗epoch run | 88.8% ↗epoch run | -54.4 pt | Claude Sonnet 5.5 |
| Furniture Assembly (Epoch AI run) | 28.3% ↗epoch run | 75% ↗epoch run | -46.7 pt | Claude Sonnet 5.5 |
| LiveBench Agentic Coding (LiveBench) | 39.7% ↗ | 56.3% ↗ | -16.6 pt | Claude Sonnet 5.5 |
| LiveBench Coding (LiveBench) | 79.7% ↗ | 91.4% ↗ | -11.7 pt | Claude Sonnet 5.5 |
| LiveBench Data Analysis (LiveBench) | 74.4% ↗ | 78.6% ↗ | -4.2 pt | Claude Sonnet 5.5 |
| LiveBench Instruction Following (LiveBench) | 62.5% ↗ | 70.5% ↗ | -8 pt | Claude Sonnet 5.5 |
| LiveBench Language (LiveBench) | 81.3% ↗ | 83.4% ↗ | -2.1 pt | Claude Sonnet 5.5 |
| LiveBench Mathematics (LiveBench) | 90.4% ↗ | 96.7% ↗ | -6.3 pt | Claude Sonnet 5.5 |
| LiveBench Reasoning (LiveBench) | 80.1% ↗ | 91.6% ↗ | -11.5 pt | Claude Sonnet 5.5 |
| LMArena WebDev (LMArena) | 1493.3 ↗ | 1786.3 ↗ | -293 | Claude Sonnet 5.5 |
| Mystery Game Puzzles (Epoch AI run) | 22% ↗epoch run | 65% ↗epoch run | -43 pt | Claude Sonnet 5.5 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 86.1% ↗epoch run | 100% ↗epoch run | -13.9 pt | Claude Sonnet 5.5 |
| SAGE (Vals AI) | 52.1% ↗ | 51.8% ↗ | +0.3 pt | Claude Opus 4.5 |
| SimpleBench (SimpleBench) | 62% ↗ | 75.9% ↗ | -13.9 pt | Claude Sonnet 5.5 |
| SimpleQA Verified | 45.7% ↗epoch run | 46.5% ↗epoch run | -0.8 pt | Claude Sonnet 5.5 |
| Only Claude Opus 4.5 reports · 20 | ||||
| SWE-bench Verified | 80.9% ↗ | not reported | — | — |
| ARC-AGI-2 (Verified) | 37.6% ↗ | not reported | — | — |
| BALROG (BALROG) | 43.5% ↗ | not reported | — | — |
| Chess Puzzles (Epoch AI run) | 12% ↗epoch run | not reported | — | — |
| EBR-bench (Epoch AI run) | 14.3% ↗epoch run | not reported | — | — |
| GSO Opt@1 (GSO) | 24.5% ↗ | not reported | — | — |
| MCP Atlas | 62.3% ↗ | not reported | — | — |
| MMMLU | 90.8% ↗ | not reported | — | — |
| MMMU (validation) | 80.7% ↗ | not reported | — | — |
| OSWorld | 66.3% ↗ | not reported | — | — |
| SWE-bench Verified (Epoch AI run) | 76.7% ↗epoch run | not reported | — | — |
| tau2-bench Airline (Sierra) | 84% ↗ | not reported | — | — |
| tau2-bench Banking Knowledge (Sierra) | 24.7% ↗ | not reported | — | — |
| tau2-bench Retail (Sierra) | 79.6% ↗ | not reported | — | — |
| tau2-bench Telecom (Sierra) | 92.3% ↗ | not reported | — | — |
| Terminal-Bench 2.0 | 59.3% ↗ | not reported | — | — |
| Vending-Bench 2 (Andon Labs) | 4967.06 ↗ | not reported | — | — |
| WeirdML (Håvard Tveit Ihle) | 63.7% ↗ | not reported | — | — |
| τ2-bench (Retail) | 88.9% ↗ | not reported | — | — |
| τ2-bench (Telecom) | 98.2% ↗ | not reported | — | — |
| Only Claude Sonnet 5.5 reports · 11 | ||||
| APEX-Agents (Mercor) | not reported | 75.5% ↗ | — | — |
| Chartography (no tools) | not reported | 61.6% ↗ | — | — |
| CursorBench 4.0 | not reported | 55.5% ↗ | — | — |
| FrontierCode v1.1 (Main) | not reported | 46.2% ↗ | — | — |
| FrontierSWE V2 (Proximal Labs) | not reported | 61.9% ↗ | — | — |
| Humanity's Last Exam (with tools) | not reported | 64.5% ↗ | — | — |
| LMArena Agent (LMArena) | not reported | 0.1252 ↗ | — | — |
| LMArena Vision (LMArena) | not reported | 1268.3 ↗ | — | — |
| OSWorld 2.1 (partial) | not reported | 80.1% ↗ | — | — |
| Terminal-Bench 4.0 | not reported | 70.6% ↗ | — | — |
| Terminal-Bench 4.0 (Vals AI) | not reported | 64.1% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Claude Sonnet 5.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Opus 4.5 | Claude Sonnet 5.5 | Edge |
|---|---|---|---|
| Context window | 200K tokens | 1M tokens | Claude Sonnet 5.5 |
| Max output | 64K tokens | 128K tokens | Claude Sonnet 5.5 |
| Input price / 1M | $5 ↗ | $2 ↗ | Claude Sonnet 5.5 |
| Output price / 1M | $25 ↗ | $10 ↗ | Claude Sonnet 5.5 |
| Cached input / 1M | $0.5 ↗ | $0.2 ↗ | Claude Sonnet 5.5 |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $30 | $12 | Claude Sonnet 5.5 |
| Modalities | text · vision | text · vision | Tie |
| Released | 2025-11-24 | 2026-09-28 | — |
| Cited benchmark scores | 39 | 29 | — |
Reliability
Anthropic status
Live from /statusMore matchups
Claude Opus 4.5 vs …
Models sharing the most benchmarksMore matchups
Claude Sonnet 5.5 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Opus 4.5 and Claude Sonnet 5.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.