Gemini 2.5 Pro
by Google DeepMind · United StatesThe first Gemini with reasoning on by default, and a long run at the top of the public preference leaderboards.
Specification
The numbers
- Maker
- Google DeepMind
- Released
- 2025-03
- Parameters
- Not reported
- Architecture
- Not reported
- Context window
- 1,048,576 tokens
- Max output
- 65,536 tokens
- Input
- Text, Image, Audio, Video
- Output
- Text
- Reasoning
- Yes
- Tool calling
- Yes
- Knowledge cutoff
- 2025-01
- Licence
- Proprietary
- Availability
- Not reported
- Weights
- Not released
Cost
Price per million tokens
- Input
- $1.25 / M tokens
- Output
- $10.00 / M tokens
- Cached input
- $0.31
- Blended 3:1
- $3.44
Standard first-party API rate, excluding batch discounts. Reasoning models bill thinking tokens as output, so cost per answer can far exceed the cost per token suggests.
Published scores
Benchmarks
Figures published by Google DeepMind or taken from a public leaderboard. Row last checked 2026-08.
| Category | Score | Rank |
|---|---|---|
| ReasoningGPQA Diamond | 84.0% | 8 of 83 reporting the same tests |
| MathsAIME 2025 | 86.7% | 23 of 56 reporting the same tests |
| CodingSWE-bench Verified | 63.8% | 21 of 33 reporting the same tests |
| KnowledgeBreadth of factual recall under exam conditions | Not reported | — |
| MultimodalMMMU | 81.7% | 4 of 39 reporting the same tests |
| Instruction followingObeying an exact, checkable format | Not reported | — |
| Human preferenceLMArena Elo | 1439 | 1 of 24 reporting the same tests |
A category averages every benchmark in it that Gemini 2.5 Pro reports. The rank counts only models that report the same tests, so it never compares an average over three benchmarks against an average over one.
Every reported test
| MMLU-ProKnowledge | Not reported | No figure published |
|---|---|---|
| GPQA DiamondReasoning | Rank 8 of 83 models reporting | |
| AIME 2025Maths | Rank 23 of 56 models reporting | |
| MATH-500Maths | Not reported | No figure published |
| SWE-bench VerifiedCoding | Rank 21 of 33 models reporting | |
| SWE-bench ProCoding | Not reported | No figure published |
| Terminal-Bench 2.1Coding | Not reported | No figure published |
| Frontier-Bench v0.1Reasoning | Not reported | No figure published |
| Terminal-Bench 4.0Coding | Not reported | No figure published |
| LiveCodeBenchCoding | Not reported | No figure published |
| HumanEvalCoding | Not reported | No figure published |
| MMMUMultimodal | Rank 4 of 39 models reporting | |
| IFEvalInstruction following | Not reported | No figure published |
| LMArena EloHuman preference | Rank 1 of 24 models reporting |
Where these numbers come from
Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.
A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.
Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.
Head to head
Gemini 2.5 Pro compared
Same maker
Other models from Google DeepMind
| Gemini 3 ProGoogle DeepMind | Google DeepMind | 1M | $2.00 | $12.00 | — | 91.9% | 95% | 76.2% | Proprietary |
| Gemini 2.5 Flash-LiteGoogle DeepMind | Google DeepMind | 1M | $0.10 | $0.40 | — | 64.6% | 49.8% | — | Proprietary |
| Gemini 2.5 FlashGoogle DeepMind | Google DeepMind | 1M | $0.30 | $2.50 | — | 78.3% | 78% | — | Proprietary |
| Gemma 3 27BGoogle DeepMind | Google DeepMind | 131K | $0.10 | $0.20 | 67.5% | 42.4% | — | — | Gemma Terms of Use |
| Gemma 3 12BGoogle DeepMind | Google DeepMind | 131K | $0.05 | $0.10 | 60.6% | 34.9% | — | — | Gemma Terms of Use |
| Gemma 3 4BGoogle DeepMind | Google DeepMind | 131K | $0.02 | $0.04 | 43.6% | — | — | — | Gemma Terms of Use |
| Gemini 2.0 Flash-LiteGoogle DeepMind | Google DeepMind | 1M | $0.075 | $0.30 | 71.6% | 51.5% | — | — | Proprietary |
| Gemini 2.0 FlashGoogle DeepMind | Google DeepMind | 1M | $0.10 | $0.40 | 77.6% | 62.1% | — | — | Proprietary |
| Gemma 2 27BGoogle DeepMind | Google DeepMind | 8.2K | $0.27 | $0.27 | 56% | — | — | — | Gemma Terms of Use |
| Gemma 2 9BGoogle DeepMind | Google DeepMind | 8.2K | $0.06 | $0.06 | 45% | — | — | — | Gemma Terms of Use |
| Gemini 1.5 ProGoogle DeepMind | Google DeepMind | 2.1M | $1.25 | $5.00 | 75.8% | 59.1% | — | — | Proprietary |
| Gemini 1.5 FlashGoogle DeepMind | Google DeepMind | 1M | $0.075 | $0.30 | 67.3% | 51% | — | — | Proprietary |