Compare / head-to-head

LongCat Flash ChatvsMiniMax M1

LongCat Flash Chat leads 5 of 7 shared benchmarks. MiniMax M1 has the larger context window (1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
5 – 2
LongCat Flash Chat leads
Cheaper per token
—
list price, input + output
Larger context
MiniMax M1
1M tokens
Providers
2 providers
Meituan LongCat · MiniMax
LongCat Flash Chat
Meituan LongCat · released 2025-09-01
text
Context
128K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
32 · 9 core
MiniMax M1
MiniMax · released 2025-06-16
text
Context
1M
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
16 · 7 core
Quality

Benchmark matrix

BenchmarkLongCat Flash ChatMiniMax M1ΔEdge
Reported by both · 7
SWE-bench Verified60.4% ↗56% ↗+4.4 ptLongCat Flash Chat
GPQA Diamond73.23% ↗70% ↗+3.2 ptLongCat Flash Chat
MMLU-Pro82.68% ↗81.1% ↗+1.6 ptLongCat Flash Chat
AIME 202561.25% ↗76.9% ↗-15.7 ptMiniMax M1
LMArena Elo1422.9 ↗1363.9 ↗+59LongCat Flash Chat
AIME 202470.42% ↗86% ↗-15.6 ptMiniMax M1
ZebraLogic89.3% ↗86.8% ↗+2.5 ptLongCat Flash Chat
Only LongCat Flash Chat reports · 24
MMLU89.71% ↗not reported——
AceBench76.1% ↗not reported——
ArenaHard-V286.5% ↗not reported——
BeyondAIME43% ↗not reported——
C-Eval90.44% ↗not reported——
CMMLU84.34% ↗not reported——
COLLIE57.1% ↗not reported——
DROP79.06% ↗not reported——
GraphWalks-128k51.05% ↗not reported——
HumanEval+88.41% ↗not reported——
IFEval89.65% ↗not reported——
LiveCodeBench48.02% ↗not reported——
MATH50096.4% ↗not reported——
MBPP+79.63% ↗not reported——
Meeseeks-zh43.03% ↗not reported——
Safety: Criminal91.24% ↗not reported——
Safety: Harmful83.98% ↗not reported——
Safety: Misinformation81.72% ↗not reported——
Safety: Privacy93.98% ↗not reported——
Tau2-Bench (airline)58% ↗not reported——
Tau2-Bench (retail)71.27% ↗not reported——
Tau2-Bench (telecom)73.68% ↗not reported——
TerminalBench39.51% ↗not reported——
VitaBench24.3% ↗not reported——
Only MiniMax M1 reports · 9
FullStackBenchnot reported68.3% ↗——
HLE (no tools)not reported8.4% ↗——
LiveCodeBench (24/8~25/5)not reported65% ↗——
LongBench-v2not reported61.5% ↗——
MATH-500not reported96.8% ↗——
OpenAI-MRCR (128k)not reported73.4% ↗——
SimpleQAnot reported18.5% ↗——
TAU-bench (airline)not reported62% ↗——
TAU-bench (retail)not reported63.5% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is LongCat Flash Chat minus MiniMax M1 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecLongCat Flash ChatMiniMax M1Edge
Context window128K tokens1M tokensMiniMax M1
Max output———
Input price / 1M———
Output price / 1M———
Cached input / 1M———
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
———
ModalitiestexttextTie
Released2025-09-012025-06-16—
Cited benchmark scores3216—
Reliability

Provider status

All providers →
More matchups

LongCat Flash Chat vs …

More matchups

MiniMax M1 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run LongCat Flash Chat and MiniMax M1 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.