Compare / head-to-head

Phi-4-mini-flash-reasoningvsQwen3 30B-A3B Thinking 2507

Qwen3 30B-A3B Thinking 2507 leads 2 of 2 shared benchmarks. Qwen3 30B-A3B Thinking 2507 has the larger context window (262K tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 2
Qwen3 30B-A3B Thinking 2507 leads
Cheaper per token
—
list price, input + output
Larger context
Qwen3 30B-A3B Thinking 2507
262K tokens
Providers
2 providers
Microsoft · Qwen
Phi-4-mini-flash-reasoning
Microsoft · released 2025-07-09
text
Context
64K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
4 · 3 core
Qwen3 30B-A3B Thinking 2507
Qwen · released 2025-07-31
text
Context
262K
Max out
—
Input /1M
$0.2 ↗
Output /1M
$2.4 ↗
Cached /1M
—
Scores
18 · 7 core
Quality

Benchmark matrix

BenchmarkPhi-4-mini-flash-reasoningQwen3 30B-A3B Thinking 2507ΔEdge
Reported by both · 2
GPQA Diamond45.08% ↗70.1% ↗epoch run-25 ptQwen3 30B-A3B Thinking 2507
AIME2533.59% ↗85% ↗-51.4 ptQwen3 30B-A3B Thinking 2507
Only Phi-4-mini-flash-reasoning reports · 2
AIME2452.29% ↗not reported——
Math50092.45% ↗not reported——
Only Qwen3 30B-A3B Thinking 2507 reports · 16
MMLU-Pronot reported80.9% ↗——
Arena-Hard v2not reported56% ↗——
BFCL-v3not reported72.4% ↗——
Chess Puzzles (Epoch AI run)not reported8% ↗epoch run——
GPQAnot reported73.4% ↗——
HMMT25not reported71.4% ↗——
IFEvalnot reported88.9% ↗——
LiveBench 20241125not reported76.8% ↗——
LiveCodeBench v6 (25.02-25.05)not reported66% ↗——
MMLU-Reduxnot reported91.4% ↗——
OJBenchnot reported25.1% ↗——
OTIS Mock AIME 2024-2025 (Epoch AI run)not reported70.3% ↗epoch run——
SuperGPQAnot reported56.8% ↗——
TAU2-Airlinenot reported58% ↗——
TAU2-Retailnot reported58.8% ↗——
TAU2-Telecomnot reported26.3% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Phi-4-mini-flash-reasoning minus Qwen3 30B-A3B Thinking 2507 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecPhi-4-mini-flash-reasoningQwen3 30B-A3B Thinking 2507Edge
Context window64K tokens262K tokensQwen3 30B-A3B Thinking 2507
Max output———
Input price / 1M—$0.2 ↗—
Output price / 1M—$2.4 ↗—
Cached input / 1M———
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
—$2.6—
ModalitiestexttextTie
Released2025-07-092025-07-31—
Cited benchmark scores418—
Reliability

Provider status

All providers →
Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Phi-4-mini-flash-reasoning and Qwen3 30B-A3B Thinking 2507 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.