Compare / head-to-head

Phi-4-mini-flash-reasoningvsQwen3 4B Instruct 2507

Qwen3 4B Instruct 2507 leads 2 of 2 shared benchmarks. Qwen3 4B Instruct 2507 has the larger context window (262K tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 2
Qwen3 4B Instruct 2507 leads
Cheaper per token
—
list price, input + output
Larger context
Qwen3 4B Instruct 2507
262K tokens
Providers
2 providers
Microsoft · Qwen
Phi-4-mini-flash-reasoning
Microsoft · released 2025-07-09
text
Context
64K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
4 · 3 core
Qwen3 4B Instruct 2507
Qwen · released 2025-08-06
text
Context
262K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
18 · 7 core
Quality

Benchmark matrix

BenchmarkPhi-4-mini-flash-reasoningQwen3 4B Instruct 2507ΔEdge
Reported by both · 2
GPQA Diamond45.08% ↗45.8% ↗epoch run-0.7 ptQwen3 4B Instruct 2507
AIME2533.59% ↗47.4% ↗-13.8 ptQwen3 4B Instruct 2507
Only Phi-4-mini-flash-reasoning reports · 2
AIME2452.29% ↗not reported——
Math50092.45% ↗not reported——
Only Qwen3 4B Instruct 2507 reports · 16
MMLU-Pronot reported69.6% ↗——
Aider-Polyglotnot reported12.9% ↗——
Arena-Hard v2not reported43.4% ↗——
BFCL-v3not reported61.9% ↗——
Chess Puzzles (Epoch AI run)not reported4% ↗epoch run——
GPQAnot reported62% ↗——
HMMT25not reported31% ↗——
IFEvalnot reported83.4% ↗——
LiveBench 20241125not reported63% ↗——
LiveCodeBench v6 (25.02-25.05)not reported35.1% ↗——
MMLU-Reduxnot reported84.2% ↗——
MultiPL-Enot reported76.8% ↗——
OTIS Mock AIME 2024-2025 (Epoch AI run)not reported52.2% ↗epoch run——
SuperGPQAnot reported42.8% ↗——
TAU2-Retailnot reported40.4% ↗——
ZebraLogicnot reported80.2% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Phi-4-mini-flash-reasoning minus Qwen3 4B Instruct 2507 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecPhi-4-mini-flash-reasoningQwen3 4B Instruct 2507Edge
Context window64K tokens262K tokensQwen3 4B Instruct 2507
Max output———
Input price / 1M———
Output price / 1M———
Cached input / 1M———
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
———
ModalitiestexttextTie
Released2025-07-092025-08-06—
Cited benchmark scores418—
Reliability

Provider status

All providers →
Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Phi-4-mini-flash-reasoning and Qwen3 4B Instruct 2507 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.