Compare / head-to-head

InklingvsKimi K3

Kimi K3 leads 23 of 25 shared benchmarks. Inkling is 3.6x cheaper per token. Kimi K3 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
2 – 23
Kimi K3 leads
Cheaper per token
Inkling
3.6x cheaper, input + output
Larger context
Kimi K3
1.0M tokens
Providers
2 providers
Thinking Machines Lab · Moonshot AI
Inkling
Thinking Machines Lab · released 2026-07-15
textvisionaudio
Context
1M
Max out
—
Input /1M
$1 ↗
Output /1M
$4.05 ↗
Cached /1M
$0.17
Scores
39 · 12 core
Kimi K3
Moonshot AI · released 2026-07-16
textvision
Context
1.0M
Max out
1.0M
Input /1M
$3 ↗
Output /1M
$15 ↗
Cached /1M
$0.3
Scores
87 · 11 core
Quality

Benchmark matrix

BenchmarkInklingKimi K3ΔEdge
Reported by both · 25
GPQA Diamond88.3% ↗epoch run93.5% ↗-5.2 ptKimi K3
AIME 202697.1% ↗96.7% ↗matharena ⚠+0.4 ptInkling
LMArena Elo1441.4 ↗1488.4 ↗-47Kimi K3
APEX-Agents (Mercor)33.8% ↗50.6% ↗-16.8 ptKimi K3
Chess Puzzles (Epoch AI run)21% ↗epoch run39% ↗epoch run-18 ptKimi K3
FrontierMath Tier 4 v2 (Epoch AI run)4.9% ↗epoch run39% ↗epoch run-34.1 ptKimi K3
FrontierMath Tiers 1-3 v2 (Epoch AI run)33.3% ↗epoch run72.2% ↗epoch run-38.9 ptKimi K3
FrontierSWE V2 (Proximal Labs)4.1% ↗25.9% ↗-21.8 ptKimi K3
LiveBench Agentic Coding (LiveBench)49.4% ↗62.2% ↗-12.8 ptKimi K3
LiveBench Coding (LiveBench)71% ↗81.4% ↗-10.4 ptKimi K3
LiveBench Data Analysis (LiveBench)72.8% ↗78.7% ↗-5.9 ptKimi K3
LiveBench Instruction Following (LiveBench)70.1% ↗71.4% ↗-1.3 ptKimi K3
LiveBench Language (LiveBench)73.5% ↗85.5% ↗-12 ptKimi K3
LiveBench Mathematics (LiveBench)88.4% ↗84.4% ↗+4 ptInkling
LiveBench Reasoning (LiveBench)78.3% ↗90.7% ↗-12.4 ptKimi K3
LMArena Agent (LMArena)-0.1086 ↗0.0418 ↗-0.2Kimi K3
LMArena WebDev (LMArena)1412.6 ↗1657.8 ↗-245.2Kimi K3
OTIS Mock AIME 2024-2025 (Epoch AI run)88.9% ↗epoch run97.2% ↗epoch run-8.3 ptKimi K3
SAGE (Vals AI)36.6% ↗52.8% ↗-16.2 ptKimi K3
SimpleBench (SimpleBench)50% ↗60.7% ↗-10.7 ptKimi K3
SimpleQA Verified43.9% ↗50.6% ↗epoch run-6.7 ptKimi K3
tau2-bench Banking Knowledge (Sierra)25% ↗37.1% ↗-12.1 ptKimi K3
Terminal-Bench 4.0 (Vals AI)0.5% ↗17.2% ↗-16.7 ptKimi K3
Toolathlon-Verified (HKUST)45.5% ↗76.5% ↗-31 ptKimi K3
WeirdML (Håvard Tveit Ihle)32.3% ↗82.6% ↗-50.3 ptKimi K3
Only Inkling reports · 13
SWE-bench Verified77.6% ↗not reported——
Audio MC56.6% ↗not reported——
BrowseComp (w/ ctx management)77.1% ↗not reported——
CharXiv RQ78.1% ↗not reported——
CharXiv RQ (with python)82% ↗not reported——
Global-MMLU-Lite88.7% ↗not reported——
IFBench79.8% ↗not reported——
MCP Atlas76% ↗not reported——
MMAU77.2% ↗not reported——
SWE-bench Pro (public)54.3% ↗not reported——
Terminal-Bench 2.1 (best harness)63.8% ↗not reported——
Toolathlon Verified45.5% ↗not reported——
VoiceBench91.4% ↗not reported——
Only Kimi K3 reports · 55
Humanity's Last Exam (no tools)not reported43.5% ↗——
AA-Briefcasenot reported1548 ↗——
AA-LCRnot reported74.7% ↗——
Agents' Last Examnot reported28.3% ↗——
APEX-Agentsnot reported41% ↗——
AutomationBenchnot reported30.8% ↗——
BabyVision (with Python)not reported85.7% ↗——
BrowseComp (1M context, no compaction)not reported90.4% ↗——
BrowseComp (context compaction)not reported91.2% ↗——
CharXiv Reasoning (no tools)not reported84.8% ↗——
CharXiv Reasoning (with Python)not reported91.3% ↗——
CorpFin v2not reported71.6% ↗——
CritPtnot reported23.4% ↗——
DeepSearchQA (F1)not reported95% ↗——
DeepSWE v1.1not reported67.3% ↗——
DeepSWE v1.1 (Datacurve)not reported68.5% ↗——
DeepSWE v1.1 (Kimi Code)not reported67.5% ↗——
Finance Agent v2not reported54.4% ↗——
FrontierSWEnot reported81.2% ↗——
Furniture Assembly (Epoch AI run)not reported34.2% ↗epoch run——
GDPval-AA v2not reported1686 ↗——
Harvey Lab-AAnot reported94.6% ↗——
Humanity's Last Exam (with tools)not reported56% ↗——
JobBenchnot reported54.3% ↗——
Kimi Code Bench 2.0not reported72.9% ↗——
Legal Research Benchnot reported44.2% ↗——
MathVision (no tools)not reported94.3% ↗——
MathVision (with Python)not reported97.8% ↗——
MCP-Atlasnot reported84.2% ↗——
MCPMark-Verifiednot reported94.5% ↗——
MLS-Bench-Litenot reported48.3% ↗——
MMMU-Pro (no tools)not reported81.6% ↗——
MMMU-Pro (with Python)not reported83.4% ↗——
MMVUnot reported82.1% ↗——
Mystery Game Puzzles (Epoch AI run)not reported26% ↗epoch run——
OfficeQA Pronot reported63.3% ↗——
OmniDocBenchnot reported91.1% ↗——
OSWorld 2.0not reported58.3% ↗——
OSWorld-Verifiednot reported84.8% ↗——
PerceptionBenchnot reported58.5% ↗——
PostTrainBenchnot reported36.6% ↗——
ProgramBenchnot reported77.8% ↗——
ResearchRubricsnot reported76.2% ↗——
SaaS-Benchnot reported60.1% ↗——
SciCodenot reported58.7% ↗——
SpreadsheetBench 2not reported34.8% ↗——
SWE-Marathonnot reported42% ↗——
tau3-Bankingnot reported33.4% ↗——
Terminal-Bench 2.1not reported88.3% ↗——
Toolathlon-Verifiednot reported76.5% ↗——
Vending-Bench 2 (Andon Labs)not reported5165.04 ↗——
Video-MME (with subtitles)not reported90% ↗——
WorldVQA ForceAnswernot reported51% ↗——
ZeroBench pass@5 (no tools)not reported23% ↗——
ZeroBench pass@5 (with Python)not reported41% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Inkling minus Kimi K3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecInklingKimi K3Edge
Context window1M tokens1.0M tokensKimi K3
Max output—1.0M tokens—
Input price / 1M$1 ↗$3 ↗Inkling
Output price / 1M$4.05 ↗$15 ↗Inkling
Cached input / 1M$0.17 ↗$0.3 ↗Inkling
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$5.05$18Inkling
Modalitiestext · vision · audiotext · visionInkling
Released2026-07-152026-07-16—
Cited benchmark scores3987—
More matchups

Kimi K3 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Inkling and Kimi K3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.