CorX Labs

Kimi K2 Thinking

by Moonshot AI · China
Open weightsReasoningTool callingMixture of experts262K context

A trillion-parameter open-weight reasoning model that holds tool-use chains across hundreds of sequential calls.


Specification

The numbers

Maker
Moonshot AI
Released
2025-11
Parameters
1T total / 32B active
Architecture
MoE
Context window
262,144 tokens
Max output
131,072 tokens
Input
Text
Output
Text
Reasoning
Yes
Tool calling
Yes
Knowledge cutoff
Not reported
Licence
Modified MIT
Availability
Not reported
Weights
Downloadable

Cost

Price per million tokens

Input
$0.60 / M tokens
Output
$2.50 / M tokens
Blended 3:1
$1.07

This model has open weights, so there is no first-party price. The figures above are a representative third-party hosting rate — you can also run it yourself for the cost of the hardware.

Published scores

Benchmarks

Figures published by Moonshot AI or taken from a public leaderboard. Row last checked 2026-08.

Kimi K2 Thinking by capability category, with its rank among models reporting the same tests.
CategoryScore Rank
ReasoningGPQA Diamond84.5%7 of 83 reporting the same tests
MathsAIME 202594.5%4 of 56 reporting the same tests
CodingSWE-bench Verified, LiveCodeBench77.2%1 of 4 reporting the same tests
KnowledgeBreadth of factual recall under exam conditionsNot reported
MultimodalReading charts, diagrams and photographsNot reported
Instruction followingObeying an exact, checkable formatNot reported
Human preferenceWhich answer people pick, blindNot reported

A category averages every benchmark in it that Kimi K2 Thinking reports. The rank counts only models that report the same tests, so it never compares an average over three benchmarks against an average over one.

Every reported test

MMLU-ProKnowledgeNot reportedNo figure published
GPQA DiamondReasoning84.5%Rank 7 of 83 models reporting
AIME 2025Maths94.5%Rank 4 of 56 models reporting
MATH-500MathsNot reportedNo figure published
SWE-bench VerifiedCoding71.3%Rank 9 of 33 models reporting
SWE-bench ProCodingNot reportedNo figure published
Terminal-Bench 2.1CodingNot reportedNo figure published
Frontier-Bench v0.1ReasoningNot reportedNo figure published
Terminal-Bench 4.0CodingNot reportedNo figure published
LiveCodeBenchCoding83.1%Rank 1 of 11 models reporting
HumanEvalCodingNot reportedNo figure published
MMMUMultimodalNot reportedNo figure published
IFEvalInstruction followingNot reportedNo figure published
LMArena EloHuman preferenceNot reportedNo figure published

Where these numbers come from

Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.

A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.

Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.

Same maker

Other models from Moonshot AI

AI models with context window, price per million tokens and published benchmark scores. Sortable by any column.
Kimi K3Moonshot AIMoonshot AI1M$3.00$15.00Kimi K3 License
Kimi K2 InstructMoonshot AIMoonshot AI131K$0.60$2.5081.1%65.8%Modified MIT
Kimi-Dev-72BMoonshot AIMoonshot AI131K$0.290$1.1560.4%Modified MIT