Compare / head-to-head
Inklingvs
Muse Spark 1.3
Muse Spark 1.3 leads 17 of 17 shared benchmarks. Inkling is 1.1x cheaper per token. Muse Spark 1.3 has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
0 – 17
Muse Spark 1.3 leads
Cheaper per token
Inkling
1.1x cheaper, input + output
Larger context
Muse Spark 1.3
1.0M tokens
Providers
2 providers
Thinking Machines Lab · Meta
Inkling
Thinking Machines Lab · released 2026-07-15
textvisionaudio
Muse Spark 1.3
Meta · released 2026-09-02
textvisionvideo
Quality
Benchmark matrix
17 shared · 21 only Inkling · 12 only Muse Spark 1.3| Benchmark | Inkling | Muse Spark 1.3 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 17 | ||||
| LMArena Elo | 1441.4 ↗ | 1494.3 ↗ | -52.9 | Muse Spark 1.3 |
| APEX-Agents (Mercor) | 33.8% ↗ | 57.8% ↗ | -24 pt | Muse Spark 1.3 |
| Chess Puzzles (Epoch AI run) | 21% ↗epoch run | 38% ↗epoch run | -17 pt | Muse Spark 1.3 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 4.9% ↗epoch run | 46.3% ↗epoch run | -41.4 pt | Muse Spark 1.3 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 33.3% ↗epoch run | 74.4% ↗epoch run | -41.1 pt | Muse Spark 1.3 |
| LiveBench Agentic Coding (LiveBench) | 49.4% ↗ | 64.1% ↗ | -14.7 pt | Muse Spark 1.3 |
| LiveBench Coding (LiveBench) | 71% ↗ | 81.1% ↗ | -10.1 pt | Muse Spark 1.3 |
| LiveBench Data Analysis (LiveBench) | 72.8% ↗ | 79.6% ↗ | -6.8 pt | Muse Spark 1.3 |
| LiveBench Instruction Following (LiveBench) | 70.1% ↗ | 78% ↗ | -7.9 pt | Muse Spark 1.3 |
| LiveBench Language (LiveBench) | 73.5% ↗ | 82.8% ↗ | -9.3 pt | Muse Spark 1.3 |
| LiveBench Mathematics (LiveBench) | 88.4% ↗ | 95.9% ↗ | -7.5 pt | Muse Spark 1.3 |
| LiveBench Reasoning (LiveBench) | 78.3% ↗ | 89.7% ↗ | -11.4 pt | Muse Spark 1.3 |
| LMArena Agent (LMArena) | -0.1086 ↗ | 0.04 ↗ | -0.1 | Muse Spark 1.3 |
| LMArena WebDev (LMArena) | 1412.6 ↗ | 1656.6 ↗ | -244 | Muse Spark 1.3 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 88.9% ↗epoch run | 99.2% ↗epoch run | -10.3 pt | Muse Spark 1.3 |
| SimpleBench (SimpleBench) | 50% ↗ | 81.8% ↗ | -31.8 pt | Muse Spark 1.3 |
| Terminal-Bench 4.0 (Vals AI) | 0.5% ↗ | 24.7% ↗ | -24.2 pt | Muse Spark 1.3 |
| Only Inkling reports · 21 | ||||
| SWE-bench Verified | 77.6% ↗ | not reported | — | — |
| GPQA Diamond | 88.3% ↗epoch run | not reported | — | — |
| AIME 2026 | 97.1% ↗ | not reported | — | — |
| Audio MC | 56.6% ↗ | not reported | — | — |
| BrowseComp (w/ ctx management) | 77.1% ↗ | not reported | — | — |
| CharXiv RQ | 78.1% ↗ | not reported | — | — |
| CharXiv RQ (with python) | 82% ↗ | not reported | — | — |
| FrontierSWE V2 (Proximal Labs) | 4.1% ↗ | not reported | — | — |
| Global-MMLU-Lite | 88.7% ↗ | not reported | — | — |
| IFBench | 79.8% ↗ | not reported | — | — |
| MCP Atlas | 76% ↗ | not reported | — | — |
| MMAU | 77.2% ↗ | not reported | — | — |
| SAGE (Vals AI) | 36.6% ↗ | not reported | — | — |
| SimpleQA Verified | 43.9% ↗ | not reported | — | — |
| SWE-bench Pro (public) | 54.3% ↗ | not reported | — | — |
| tau2-bench Banking Knowledge (Sierra) | 25% ↗ | not reported | — | — |
| Terminal-Bench 2.1 (best harness) | 63.8% ↗ | not reported | — | — |
| Toolathlon Verified | 45.5% ↗ | not reported | — | — |
| Toolathlon-Verified (HKUST) | 45.5% ↗ | not reported | — | — |
| VoiceBench | 91.4% ↗ | not reported | — | — |
| WeirdML (Håvard Tveit Ihle) | 32.3% ↗ | not reported | — | — |
| Only Muse Spark 1.3 reports · 12 | ||||
| AutomationBench | not reported | 49.6% ↗ | — | — |
| DeepSearchQA | not reported | 90.3% ↗ | — | — |
| DeepSWE v1.1 | not reported | 75.4% ↗ | — | — |
| GDPval-AA v2 | not reported | 1754 ↗ | — | — |
| JobBench | not reported | 64.9% ↗ | — | — |
| LMArena Vision (LMArena) | not reported | 1290.2 ↗ | — | — |
| MRCR 256K-512K | not reported | 98.5% ↗ | — | — |
| MRCR 512K-1M | not reported | 98.1% ↗ | — | — |
| Mystery Game Puzzles (Epoch AI run) | not reported | 25% ↗epoch run | — | — |
| OSWorld 2.0 (partial) | not reported | 66.9% ↗ | — | — |
| SWE-Atlas Codebase QnA | not reported | 59.4% ↗ | — | — |
| Terminal-Bench 2.1 | not reported | 88.8% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Inkling minus Muse Spark 1.3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Inkling | Muse Spark 1.3 | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Muse Spark 1.3 |
| Max output | — | — | — |
| Input price / 1M | $1 ↗ | $1.25 ↗ | Inkling |
| Output price / 1M | $4.05 ↗ | $4.25 ↗ | Inkling |
| Cached input / 1M | $0.17 ↗ | $0.15 ↗ | Muse Spark 1.3 |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $5.05 | $5.5 | Inkling |
| Modalities | text · vision · audio | text · vision · video | Tie |
| Released | 2026-07-15 | 2026-09-02 | — |
| Cited benchmark scores | 39 | 29 | — |
Reliability
Provider status
Live from /statusMore matchups
Inkling vs …
Models sharing the most benchmarksMore matchups
Muse Spark 1.3 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Inkling and Muse Spark 1.3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.