Compare / head-to-head
Claude Opus 4.8vs
Muse Spark 1.1
Claude Opus 4.8 leads 14 of 22 shared benchmarks. Muse Spark 1.1 is 5.5x cheaper per token. Muse Spark 1.1 has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
14 – 8
Claude Opus 4.8 leads
Cheaper per token
Muse Spark 1.1
5.5x cheaper, input + output
Larger context
Muse Spark 1.1
1.0M tokens
Providers
2 providers
Anthropic · Meta
Claude Opus 4.8
Anthropic · released 2026-05-28
textvision
Muse Spark 1.1
Meta · released 2026-07-09
textvisionvideoaudio
Quality
Benchmark matrix
22 shared · 18 only Claude Opus 4.8 · 8 only Muse Spark 1.1| Benchmark | Claude Opus 4.8 | Muse Spark 1.1 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 22 | ||||
| LMArena Elo | 1452.7 ↗ | 1491.3 ↗ | -38.6 | Muse Spark 1.1 |
| APEX-Agents (Mercor) | 48.9% ↗ | 31.8% ↗ | +17.1 pt | Claude Opus 4.8 |
| DeepSWE v1.1 (Datacurve) | 59% ↗ | 53.3% ↗ | +5.7 pt | Claude Opus 4.8 |
| Humanity's Last Exam (with tools) | 57.9% ↗ | 62.1% ↗ | -4.2 pt | Muse Spark 1.1 |
| LiveBench Agentic Coding (LiveBench) | 50.5% ↗ | 58.5% ↗ | -8 pt | Muse Spark 1.1 |
| LiveBench Coding (LiveBench) | 81.8% ↗ | 77.2% ↗ | +4.6 pt | Claude Opus 4.8 |
| LiveBench Data Analysis (LiveBench) | 66% ↗ | 72.5% ↗ | -6.5 pt | Muse Spark 1.1 |
| LiveBench Instruction Following (LiveBench) | 72% ↗ | 69.6% ↗ | +2.4 pt | Claude Opus 4.8 |
| LiveBench Language (LiveBench) | 79.7% ↗ | 74.3% ↗ | +5.4 pt | Claude Opus 4.8 |
| LiveBench Mathematics (LiveBench) | 94.3% ↗ | 87.1% ↗ | +7.2 pt | Claude Opus 4.8 |
| LiveBench Reasoning (LiveBench) | 89.2% ↗ | 87.7% ↗ | +1.5 pt | Claude Opus 4.8 |
| LMArena Agent (LMArena) | 0.0664 ↗ | -0.0488 ↗ | +0.1 | Claude Opus 4.8 |
| LMArena Vision (LMArena) | 1286.4 ↗ | 1281.1 ↗ | +5.3 | Claude Opus 4.8 |
| LMArena WebDev (LMArena) | 1555.5 ↗ | 1541.7 ↗ | +13.8 | Claude Opus 4.8 |
| OSWorld-Verified | 83.4% ↗ | 80.8% ↗ | +2.6 pt | Claude Opus 4.8 |
| SAGE (Vals AI) | 54.8% ↗ | 45.6% ↗ | +9.2 pt | Claude Opus 4.8 |
| SimpleQA Verified | 53% ↗epoch run | 57.8% ↗epoch run | -4.8 pt | Muse Spark 1.1 |
| SWE-Bench Pro | 69.2% ↗ | 61.5% ↗ | +7.7 pt | Claude Opus 4.8 |
| tau2-bench Banking Knowledge (Sierra) | 39.7% ↗ | 40.5% ↗ | -0.8 pt | Muse Spark 1.1 |
| Terminal-Bench 2.1 | 74.6% ↗ | 80% ↗ | -5.4 pt | Muse Spark 1.1 |
| Toolathlon-Verified (HKUST) | 76.2% ↗ | 75.6% ↗ | +0.6 pt | Claude Opus 4.8 |
| Vending-Bench 2 (Andon Labs) | 5787.43 ↗ | 6520.48 ↗ | -733 | Muse Spark 1.1 |
| Only Claude Opus 4.8 reports · 18 | ||||
| SWE-bench Verified | 88.6% ↗ | not reported | — | — |
| GPQA Diamond | 93.6% ↗ | not reported | — | — |
| AIME 2026 | 100% ↗matharena ⚠ | not reported | — | — |
| Humanity's Last Exam (no tools) | 49.8% ↗ | not reported | — | — |
| BrowseComp | 84.3% ↗ | not reported | — | — |
| Chess Puzzles (Epoch AI run) | 34% ↗epoch run | not reported | — | — |
| EBR-bench (Epoch AI run) | 28.6% ↗epoch run | not reported | — | — |
| FrontierMath Tier 4 v2 (Epoch AI run) | 56.1% ↗epoch run | not reported | — | — |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 80% ↗epoch run | not reported | — | — |
| Furniture Assembly (Epoch AI run) | 42.5% ↗epoch run | not reported | — | — |
| GSO Opt@1 (GSO) | 47.1% ↗ | not reported | — | — |
| Mystery Game Puzzles (Epoch AI run) | 36% ↗epoch run | not reported | — | — |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 98.3% ↗epoch run | not reported | — | — |
| SimpleBench (SimpleBench) | 64.8% ↗ | not reported | — | — |
| SWE-bench Multilingual | 84.4% ↗ | not reported | — | — |
| SWE-bench Multimodal | 38.4% ↗ | not reported | — | — |
| Terminal-Bench 4.0 (Vals AI) | 23.2% ↗ | not reported | — | — |
| WeirdML (Håvard Tveit Ihle) | 82.9% ↗ | not reported | — | — |
| Only Muse Spark 1.1 reports · 8 | ||||
| BabyVision | not reported | 76.3% ↗ | — | — |
| CharXiv Reasoning | not reported | 88.4% ↗ | — | — |
| DeepSWE v1.1 | not reported | 53.3% ↗ | — | — |
| Finance Agent v2 | not reported | 57.2% ↗ | — | — |
| JobBench | not reported | 54.7% ↗ | — | — |
| MCP Atlas | not reported | 88.1% ↗ | — | — |
| OSWorld-Verified (XLANG) | not reported | 80.7% ↗ | — | — |
| Toolathlon-Verified | not reported | 75.6% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.8 minus Muse Spark 1.1 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Opus 4.8 | Muse Spark 1.1 | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Muse Spark 1.1 |
| Max output | 128K tokens | — | — |
| Input price / 1M | $5 ↗ | $1.25 ↗ | Muse Spark 1.1 |
| Output price / 1M | $25 ↗ | $4.25 ↗ | Muse Spark 1.1 |
| Cached input / 1M | $0.5 ↗ | $0.15 ↗ | Muse Spark 1.1 |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $30 | $5.5 | Muse Spark 1.1 |
| Modalities | text · vision | text · vision · video · audio | Muse Spark 1.1 |
| Released | 2026-05-28 | 2026-07-09 | — |
| Cited benchmark scores | 47 | 30 | — |
Reliability
Provider status
Live from /statusMore matchups
Claude Opus 4.8 vs …
Models sharing the most benchmarksMore matchups
Muse Spark 1.1 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Opus 4.8 and Muse Spark 1.1 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.