Compare / head-to-head
Claude Sonnet 4.5vsClaude Sonnet 4.6
Claude Sonnet 4.6 leads 15 of 17 shared benchmarks. Both list the same combined token price. Claude Sonnet 4.6 has the larger context window (1M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
2 – 15
Claude Sonnet 4.6 leads
Cheaper per token
Tie
list price, input + output
Larger context
Claude Sonnet 4.6
1M tokens
Provider uptime (30d)
100%
Anthropic
Claude Sonnet 4.5
Anthropic · released 2025-09-29
textvision
Claude Sonnet 4.6
Anthropic · released 2026-02-17
textvision
Quality
Benchmark matrix
17 shared · 3 only Claude Sonnet 4.5 · 1 only Claude Sonnet 4.6| Benchmark | Claude Sonnet 4.5 | Claude Sonnet 4.6 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 17 | ||||
| SWE-bench Verified | 77.2% ↗ | 79.6% ↗ | -2.4 pt | Claude Sonnet 4.6 |
| GPQA Diamond | 83.4% ↗ | 89.9% ↗ | -6.5 pt | Claude Sonnet 4.6 |
| Humanity's Last Exam (no tools) | 17.7% ↗ | 33.2% ↗ | -15.5 pt | Claude Sonnet 4.6 |
| ARC-AGI-2 (Verified) | 13.6% ↗ | 58.3% ↗ | -44.7 pt | Claude Sonnet 4.6 |
| GDPval-AA | 1276 ↗ | 1633 ↗ | -357 | Claude Sonnet 4.6 |
| Humanity's Last Exam (with tools) | 33.6% ↗ | 49% ↗ | -15.4 pt | Claude Sonnet 4.6 |
| MCP Atlas | 43.8% ↗ | 61.3% ↗ | -17.5 pt | Claude Sonnet 4.6 |
| MMMLU | 89.5% ↗ | 89.3% ↗ | +0.2 pt | Claude Sonnet 4.5 |
| MMMU-Pro (no tools) | 63.4% ↗ | 74.5% ↗ | -11.1 pt | Claude Sonnet 4.6 |
| MMMU-Pro (with tools) | 68.9% ↗ | 75.6% ↗ | -6.7 pt | Claude Sonnet 4.6 |
| OSWorld-Verified | 61.4% ↗ | 72.5% ↗ | -11.1 pt | Claude Sonnet 4.6 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 77.8% ↗epoch | 85.8% ↗epoch | -8 pt | Claude Sonnet 4.6 |
| SimpleQA Verified | 30.7% ↗epoch | 35.5% ↗epoch | -4.8 pt | Claude Sonnet 4.6 |
| SWE-bench Verified (Epoch AI run) | 71.3% ↗epoch | 75.2% ↗epoch | -3.9 pt | Claude Sonnet 4.6 |
| Tau2-bench Retail | 86.2% ↗ | 91.7% ↗ | -5.5 pt | Claude Sonnet 4.6 |
| Tau2-bench Telecom | 98% ↗ | 97.9% ↗ | +0.1 pt | Claude Sonnet 4.5 |
| Terminal-Bench 2.0 | 51% ↗ | 59.1% ↗ | -8.1 pt | Claude Sonnet 4.6 |
| Only Claude Sonnet 4.5 reports · 3 | ||||
| AIME 2025 | 84.2% ↗matharena ⚠ | not reported | — | — |
| FrontierMath Tier 4 v2 (Epoch AI run) | 2.4% ↗epoch | not reported | — | — |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 23.9% ↗epoch | not reported | — | — |
| Only Claude Sonnet 4.6 reports · 1 | ||||
| LMArena Elo | not reported | 1458.3 ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 4.5 minus Claude Sonnet 4.6 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Sonnet 4.5 | Claude Sonnet 4.6 | Edge |
|---|---|---|---|
| Context window | 200K tokens | 1M tokens | Claude Sonnet 4.6 |
| Max output | 64K tokens | 128K tokens | Claude Sonnet 4.6 |
| Input price / 1M | $3 ↗ | $3 ↗ | Tie |
| Output price / 1M | $15 ↗ | $15 ↗ | Tie |
| Cached input / 1M | $0.3 ↗ | $0.3 ↗ | Tie |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $18 | $18 | Tie |
| Modalities | text · vision | text · vision | Tie |
| Released | 2025-09-29 | 2026-02-17 | — |
| Cited benchmark scores | 21 | 19 | — |
Reliability
Anthropic status
Last 60 days · refreshed daily on this page · live on /statusMore matchups
Claude Sonnet 4.5 vs …
Models sharing the most benchmarksMore matchups
Claude Sonnet 4.6 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Sonnet 4.5 and Claude Sonnet 4.6 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.