Compare / head-to-head

DeepSeek V4 FlashvsGLM-5.2

DeepSeek V4 Flash leads 11 of 21 shared benchmarks. Both offer a 1M-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
11 – 10
DeepSeek V4 Flash leads
Cheaper per token
—
list price, input + output
Larger context
Tie
both 1M tokens
Providers
2 providers
DeepSeek · Z.ai
DeepSeek V4 Flash
DeepSeek · released 2026-07-31
text
Context
1M
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
28 · 7 core
GLM-5.2
Z.ai · released 2026-06-24
text
Context
1M
Max out
131K
Input /1M
$1.4 ↗
Output /1M
$4.4 ↗
Cached /1M
$0.26
Scores
53 · 12 core
Quality

Benchmark matrix

BenchmarkDeepSeek V4 FlashGLM-5.2ΔEdge
Reported by both · 21
GPQA Diamond91% ↗epoch run91.9% ↗epoch run-0.9 ptGLM-5.2
LMArena Elo1436 ↗1475.8 ↗-39.8GLM-5.2
Chess Puzzles (Epoch AI run)33% ↗epoch run21% ↗epoch run+12 ptDeepSeek V4 Flash
DeepSWE v1.1 (Datacurve)53.3% ↗43.8% ↗+9.5 ptDeepSeek V4 Flash
FrontierMath Tier 4 v2 (Epoch AI run)24.4% ↗epoch run29.3% ↗epoch run-4.9 ptGLM-5.2
FrontierMath Tiers 1-3 v2 (Epoch AI run)57.5% ↗epoch run59.2% ↗epoch run-1.7 ptGLM-5.2
LiveBench Agentic Coding (LiveBench)46.8% ↗51.8% ↗-5 ptGLM-5.2
LiveBench Coding (LiveBench)75% ↗79.7% ↗-4.7 ptGLM-5.2
LiveBench Data Analysis (LiveBench)79.3% ↗73.7% ↗+5.6 ptDeepSeek V4 Flash
LiveBench Instruction Following (LiveBench)65.5% ↗62.3% ↗+3.2 ptDeepSeek V4 Flash
LiveBench Language (LiveBench)79.2% ↗76.2% ↗+3 ptDeepSeek V4 Flash
LiveBench Mathematics (LiveBench)86.8% ↗89.8% ↗-3 ptGLM-5.2
LiveBench Reasoning (LiveBench)86.6% ↗78.6% ↗+8 ptDeepSeek V4 Flash
LMArena WebDev (LMArena)1581 ↗1605 ↗-24GLM-5.2
Mystery Game Puzzles (Epoch AI run)34% ↗epoch run19% ↗epoch run+15 ptDeepSeek V4 Flash
NL2Repo54.2% ↗48.9% ↗+5.3 ptDeepSeek V4 Flash
OTIS Mock AIME 2024-2025 (Epoch AI run)94.4% ↗epoch run86.4% ↗epoch run+8 ptDeepSeek V4 Flash
SimpleBench (SimpleBench)61.1% ↗58.8% ↗+2.3 ptDeepSeek V4 Flash
SimpleQA Verified33.6% ↗epoch run34.2% ↗epoch run-0.6 ptGLM-5.2
Toolathlon-Verified (HKUST)70.7% ↗59.9% ↗+10.8 ptDeepSeek V4 Flash
WeirdML (Håvard Tveit Ihle)63% ↗70.1% ↗-7.1 ptGLM-5.2
Only DeepSeek V4 Flash reports · 7
Agents' Last Exam25.2% ↗not reported——
AutomationBench Public25.1% ↗not reported——
Cybergym76.7% ↗not reported——
DeepSWE54.4% ↗not reported——
Terminal Bench 2.182.7% ↗not reported——
Terminal-Bench 4.0 (Vals AI)18.7% ↗not reported——
Toolathlon-Verified70.3% ↗not reported——
Only GLM-5.2 reports · 25
AIME 2026not reported90% ↗matharena ⚠——
Agents' Last Exam (ALE-CLI)not reported23.8% ↗——
APEX-Agents (Mercor)not reported45.2% ↗——
AutomationBench v1.0.6not reported26.2% ↗——
CyberGymnot reported77.2% ↗——
DeepSWE v1.1not reported46.2% ↗——
EBR-bench (Epoch AI run)not reported9.5% ↗epoch run——
ExploitBenchnot reported24.4% ↗——
ExploitGym 2hnot reported29 ↗——
ExploitGym 6hnot reported39 ↗——
FrontierSWEnot reported67.5% ↗——
GDPval-AA v2not reported1508 ↗——
HLE (with tools)not reported54.7% ↗——
LMArena Agent (LMArena)not reported0.0348 ↗——
PostTrainBenchnot reported31.7% ↗——
ProgramBench (Almost Solved)not reported9.5% ↗——
SWE-Bench Pronot reported62.1% ↗——
SWE-bench Verified (Epoch AI run)not reported78.7% ↗epoch run——
SWE-Marathon v1.1not reported19.4% ↗——
SWE-rebench 2026-05-15 to 2026-07-01 (Nebius)not reported62.9% ↗——
tau2-bench Banking Knowledge (Sierra)not reported37.1% ↗——
Terminal-Bench 2.1not reported81% ↗——
Terminal-Bench 3.0not reported4.6% ↗——
Toolathlon Verifiednot reported59.9% ↗——
Vending-Bench 2 (Andon Labs)not reported8313.78 ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is DeepSeek V4 Flash minus GLM-5.2 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecDeepSeek V4 FlashGLM-5.2Edge
Context window1M tokens1M tokensTie
Max output—131K tokens—
Input price / 1M—$1.4 ↗—
Output price / 1M—$4.4 ↗—
Cached input / 1M—$0.26 ↗—
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
—$5.8—
ModalitiestexttextTie
Released2026-07-312026-06-24—
Cited benchmark scores2853—
Reliability

Provider status

All providers →
More matchups

DeepSeek V4 Flash vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run DeepSeek V4 Flash and GLM-5.2 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.