Compare / head-to-head
Claude Fable 5vs
Gemini 3.1 Pro Preview
Claude Fable 5 leads 27 of 31 shared benchmarks. Gemini 3.1 Pro Preview is 4.3x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
27 – 4
Claude Fable 5 leads
Cheaper per token
Gemini 3.1 Pro Preview
4.3x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Fable 5
Anthropic · released 2026-06-09
textvision
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Quality
Benchmark matrix
31 shared · 13 only Claude Fable 5 · 5 only Gemini 3.1 Pro Preview| Benchmark | Claude Fable 5 | Gemini 3.1 Pro Preview | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 31 | ||||
| GPQA Diamond | 85.9% ↗epoch run | 94.3% ↗ | -8.4 pt | Gemini 3.1 Pro Preview |
| LMArena Elo | 1504.3 ↗ | 1480.1 ↗ | +24.2 | Claude Fable 5 |
| APEX-Agents (Mercor) | 63.6% ↗ | 35.3% ↗ | +28.3 pt | Claude Fable 5 |
| Chess Puzzles (Epoch AI run) | 41% ↗epoch run | 55% ↗epoch run | -14 pt | Gemini 3.1 Pro Preview |
| DeepSWE v1.1 (Datacurve) | 69.9% ↗ | 11.7% ↗ | +58.2 pt | Claude Fable 5 |
| EBR-bench (Epoch AI run) | 39.5% ↗epoch run | 14.3% ↗epoch run | +25.2 pt | Claude Fable 5 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 90.2% ↗epoch run | 26.8% ↗epoch run | +63.4 pt | Claude Fable 5 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 87% ↗epoch run | 59.6% ↗epoch run | +27.4 pt | Claude Fable 5 |
| Furniture Assembly (Epoch AI run) | 35.8% ↗epoch run | 26.7% ↗epoch run | +9.1 pt | Claude Fable 5 |
| GSO Opt@1 (GSO) | 76.5% ↗ | 21.6% ↗ | +54.9 pt | Claude Fable 5 |
| LiveBench Agentic Coding (LiveBench) | 62.2% ↗ | 44.1% ↗ | +18.1 pt | Claude Fable 5 |
| LiveBench Coding (LiveBench) | 86% ↗ | 76.5% ↗ | +9.5 pt | Claude Fable 5 |
| LiveBench Data Analysis (LiveBench) | 80.5% ↗ | 78.5% ↗ | +2 pt | Claude Fable 5 |
| LiveBench Instruction Following (LiveBench) | 75.8% ↗ | 79.1% ↗ | -3.3 pt | Gemini 3.1 Pro Preview |
| LiveBench Language (LiveBench) | 90.7% ↗ | 85.4% ↗ | +5.3 pt | Claude Fable 5 |
| LiveBench Mathematics (LiveBench) | 96% ↗ | 91% ↗ | +5 pt | Claude Fable 5 |
| LiveBench Reasoning (LiveBench) | 89.7% ↗ | 84% ↗ | +5.7 pt | Claude Fable 5 |
| LMArena Agent (LMArena) | 0.0821 ↗ | -0.0771 ↗ | +0.2 | Claude Fable 5 |
| LMArena Vision (LMArena) | 1308.8 ↗ | 1279.3 ↗ | +29.5 | Claude Fable 5 |
| LMArena WebDev (LMArena) | 1625.3 ↗ | 1446.2 ↗ | +179.1 | Claude Fable 5 |
| MirrorCode (Epoch AI run) | 63.9% ↗epoch run | 8.9% ↗epoch run | +55 pt | Claude Fable 5 |
| Mystery Game Puzzles (Epoch AI run) | 52% ↗epoch run | 34% ↗epoch run | +18 pt | Claude Fable 5 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 100% ↗epoch run | 95.6% ↗epoch run | +4.4 pt | Claude Fable 5 |
| SAGE (Vals AI) | 51.9% ↗ | 48.7% ↗ | +3.2 pt | Claude Fable 5 |
| SimpleBench (SimpleBench) | 81.9% ↗ | 79.6% ↗ | +2.3 pt | Claude Fable 5 |
| SimpleQA Verified | 70.7% ↗epoch run | 73.5% ↗epoch run | -2.8 pt | Gemini 3.1 Pro Preview |
| SWE-Bench Pro | 80% ↗ | 54.2% ↗ | +25.8 pt | Claude Fable 5 |
| tau2-bench Banking Knowledge (Sierra) | 39.7% ↗ | 26% ↗ | +13.7 pt | Claude Fable 5 |
| Terminal-Bench 4.0 (Vals AI) | 41.4% ↗ | 2.5% ↗ | +38.9 pt | Claude Fable 5 |
| Vending-Bench 2 (Andon Labs) | 5680.26 ↗ | 3774.25 ↗ | +1906 | Claude Fable 5 |
| WeirdML (Håvard Tveit Ihle) | 91.9% ↗ | 72.1% ↗ | +19.8 pt | Claude Fable 5 |
| Only Claude Fable 5 reports · 13 | ||||
| SWE-bench Verified | 95% ↗ | not reported | — | — |
| AutomationBench | 17.4% ↗ | not reported | — | — |
| Blueprint-Bench 2 | 38.6% ↗ | not reported | — | — |
| CursorBench | 72.9% ↗ | not reported | — | — |
| FrontierCode (Diamond) | 29.3% ↗ | not reported | — | — |
| FrontierSWE V2 (Proximal Labs) | 47% ↗ | not reported | — | — |
| GDP.pdf | 29.8% ↗ | not reported | — | — |
| Legal Agent Benchmark (Harvey's Held-Out Set) | 13.3% ↗ | not reported | — | — |
| OfficeQA Pro | 57.9% ↗ | not reported | — | — |
| OSWorld-Verified | 85% ↗ | not reported | — | — |
| OSWorld-Verified (XLANG) | 86% ↗ | not reported | — | — |
| SWE-rebench 2026-05-15 to 2026-07-01 (Nebius) | 64.5% ↗ | not reported | — | — |
| Terminal-Bench 2.1 | 84.3% ↗ | not reported | — | — |
| Only Gemini 3.1 Pro Preview reports · 5 | ||||
| AIME 2026 | not reported | 98.3% ↗matharena ⚠ | — | — |
| ARC-AGI-2 | not reported | 77.1% ↗ | — | — |
| BALROG (BALROG) | not reported | 57% ↗ | — | — |
| SWE-bench Verified (Epoch AI run) | not reported | 75.6% ↗epoch run | — | — |
| Toolathlon-Verified (HKUST) | not reported | 61.1% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Fable 5 minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Fable 5 | Gemini 3.1 Pro Preview | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Gemini 3.1 Pro Preview |
| Max output | 128K tokens | 66K tokens | Claude Fable 5 |
| Input price / 1M | $10 ↗ | $2 ↗ | Gemini 3.1 Pro Preview |
| Output price / 1M | $50 ↗ | $12 ↗ | Gemini 3.1 Pro Preview |
| Cached input / 1M | $1 ↗ | $0.2 ↗ | Gemini 3.1 Pro Preview |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $60 | $14 | Gemini 3.1 Pro Preview |
| Modalities | text · vision | text · vision · audio | Gemini 3.1 Pro Preview |
| Released | 2026-06-09 | 2026-02-19 | — |
| Cited benchmark scores | 44 | 43 | — |
Reliability
Provider status
Live from /statusMore matchups
Claude Fable 5 vs …
Models sharing the most benchmarksMore matchups
Gemini 3.1 Pro Preview vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Fable 5 and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.