Compare / head-to-head

DeepSeek V3.2-SpecialevsGLM-5

GLM-5 leads 4 of 4 shared benchmarks.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 4
GLM-5 leads
Cheaper per token
—
list price, input + output
Larger context
—
unknown
Providers
2 providers
DeepSeek · Z.ai
DeepSeek V3.2-Speciale
DeepSeek · released 2025-12-01
text
Context
—
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
9 · 4 core
GLM-5
Z.ai
text
Context
200K
Max out
131K
Input /1M
$1 ↗
Output /1M
$3.2 ↗
Cached /1M
$0.2
Scores
21 · 8 core
Quality

Benchmark matrix

BenchmarkDeepSeek V3.2-SpecialeGLM-5ΔEdge
Reported by both · 4
GPQA Diamond85.7% ↗87.8% ↗epoch run-2.1 ptGLM-5
AIME 202596% ↗96.7% ↗matharena ⚠-0.7 ptGLM-5
SimpleBench (SimpleBench)52.6% ↗53.2% ↗-0.6 ptGLM-5
WeirdML (Håvard Tveit Ihle)46.7% ↗48.2% ↗-1.5 ptGLM-5
Only DeepSeek V3.2-Speciale reports · 5
HLE (text-only)30.6% ↗not reported——
HMMT Feb 202599.2% ↗not reported——
HMMT Nov 202594.4% ↗not reported——
IMOAnswerBench84.5% ↗not reported——
LiveCodeBench (Pass@1-COT)88.7% ↗not reported——
Only GLM-5 reports · 13
SWE-bench Verifiednot reported77.8% ↗——
AIME 2026not reported96.7% ↗matharena ⚠——
LMArena Elonot reported1446.3 ↗——
Chess Puzzles (Epoch AI run)not reported10% ↗epoch run——
LMArena WebDev (LMArena)not reported1434 ↗——
OTIS Mock AIME 2024-2025 (Epoch AI run)not reported80% ↗epoch run——
SWE-bench Verified (Epoch AI run)not reported72.1% ↗epoch run——
tau2-bench Airline (Sierra)not reported82.5% ↗——
tau2-bench Banking Knowledge (Sierra)not reported9.8% ↗——
tau2-bench Retail (Sierra)not reported73.7% ↗——
tau2-bench Telecom (Sierra)not reported86.8% ↗——
Terminal-Bench 2.0not reported56.2% ↗——
Vending-Bench 2 (Andon Labs)not reported4432.12 ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is DeepSeek V3.2-Speciale minus GLM-5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecDeepSeek V3.2-SpecialeGLM-5Edge
Context window—200K tokens—
Max output—131K tokens—
Input price / 1M—$1 ↗—
Output price / 1M—$3.2 ↗—
Cached input / 1M—$0.2 ↗—
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
—$4.2—
ModalitiestexttextTie
Released2025-12-01——
Cited benchmark scores921—
Reliability

Provider status

All providers →
More matchups

DeepSeek V3.2-Speciale vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run DeepSeek V3.2-Speciale and GLM-5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.