Benchmarks
Compare 159 AI models, side by side.
Context windows, prices and published benchmark scores for every major model from 42 labs — searchable, sortable, and honest about what has not been measured.
The CorX Labs benchmark index tracks 159 language models from 42 companies, including OpenAI, Anthropic, Google DeepMind, Meta, Mistral, DeepSeek, Alibaba and xAI. For each model it records the context window, the price per million input and output tokens, the licence, and published scores on 14 standard evaluations — MMLU-Pro, GPQA Diamond, AIME, SWE-bench Verified and others. Pick any two to four models to see them column by column.
Where these numbers come from
Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.
A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.
Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.
The index
Every model in one table
Search by name or maker, filter by capability, and sort by any column. Tick two or more rows to compare them properly.
Showing 159 of 159 models
| Compare | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3 ProGoogle DeepMind | Google DeepMind | 1M | $2.00 | $12.00 | — | 91.9% | 95% | 76.2% | Proprietary | |
| Grok 4xAI | xAI | 256K | $3.00 | $15.00 | — | 87.5% | 94% | 72% | Proprietary | |
| Claude Opus 4.5Anthropic | Anthropic | 200K | $5.00 | $25.00 | — | 87% | 96% | 80.9% | Proprietary | |
| GPT-5OpenAI | OpenAI | 400K | $1.25 | $10.00 | — | 85.7% | 94.6% | 74.9% | Proprietary | |
| Grok 4 FastxAI | xAI | 2M | $0.20 | $0.50 | — | 85.7% | 92% | — | Proprietary | |
| Qwen3-MaxAlibaba Qwen | Alibaba Qwen | 262K | $1.20 | $6.00 | — | 85.4% | 92.3% | 69.6% | Proprietary | |
| Kimi K2 ThinkingMoonshot AI | Moonshot AI | 262K | $0.60 | $2.50 | — | 84.5% | 94.5% | 71.3% | Modified MIT | |
| Gemini 2.5 ProGoogle DeepMind | Google DeepMind | 1M | $1.25 | $10.00 | — | 84% | 86.7% | 63.8% | Proprietary | |
| Claude Sonnet 4.5Anthropic | Anthropic | 1M | $3.00 | $15.00 | — | 83.4% | 87% | 77.2% | Proprietary | |
| o3OpenAI | OpenAI | 200K | $2.00 | $8.00 | — | 83.3% | 88.9% | 69.1% | Proprietary | |
| GLM-4.6Z.ai (Zhipu) | Z.ai (Zhipu) | 205K | $0.60 | $2.20 | — | 82.9% | 93.9% | 68% | MIT | |
| GPT-5 miniOpenAI | OpenAI | 400K | $0.25 | $2.00 | — | 82.3% | 91.1% | 71% | Proprietary | |
| o4-miniOpenAI | OpenAI | 200K | $1.10 | $4.40 | — | 81.4% | 92.7% | 68.1% | Proprietary | |
| Claude Opus 4.1Anthropic | Anthropic | 200K | $15.00 | $75.00 | — | 80.9% | 78% | 74.5% | Proprietary | |
| DeepSeek-V3.1DeepSeek | DeepSeek | 131K | $0.280 | $1.14 | 84.8% | 80.1% | 88.4% | 66% | MIT | |
| gpt-oss-120bOpenAI | OpenAI | 131K | $0.10 | $0.50 | 80.9% | 80.1% | 92.5% | — | Apache 2.0 | |
| DeepSeek-V3.2DeepSeek | DeepSeek | 164K | $0.280 | $0.42 | 85% | 79.9% | 89.3% | — | MIT | |
| o3-miniOpenAI | OpenAI | 200K | $1.10 | $4.40 | — | 79.7% | 87.3% | 49.3% | Proprietary | |
| GLM-4.5Z.ai (Zhipu) | Z.ai (Zhipu) | 131K | $0.60 | $2.20 | — | 79.1% | 91% | 64.2% | MIT | |
| Gemini 2.5 FlashGoogle DeepMind | Google DeepMind | 1M | $0.30 | $2.50 | — | 78.3% | 78% | — | Proprietary | |
| MiniMax-M2MiniMax | MiniMax | 205K | $0.30 | $1.20 | — | 78% | 78% | 69.4% | MIT | |
| o1OpenAI | OpenAI | 200K | $15.00 | $60.00 | — | 78% | 79.2% | 48.9% | Proprietary | |
| Llama-3.1-Nemotron-Ultra-253BNVIDIA | NVIDIA | 131K | $0.60 | $1.80 | — | 76% | 80.1% | — | NVIDIA Open Model | |
| Claude Sonnet 4Anthropic | Anthropic | 200K | $3.00 | $15.00 | — | 75.4% | 70.5% | 72.7% | Proprietary | |
| Grok 3xAI | xAI | 131K | $3.00 | $15.00 | — | 75.4% | 83.9% | — | Proprietary | |
| GLM-4.5-AirZ.ai (Zhipu) | Z.ai (Zhipu) | 131K | $0.20 | $1.10 | — | 75% | 89.4% | 57.6% | MIT | |
| Claude Haiku 4.5Anthropic | Anthropic | 200K | $1.00 | $5.00 | — | 73% | 77% | 73.3% | Proprietary | |
| DeepSeek-R1DeepSeek | DeepSeek | 131K | $0.550 | $2.19 | 84% | 71.5% | 79.8% | 49.2% | MIT | |
| gpt-oss-20bOpenAI | OpenAI | 131K | $0.05 | $0.20 | 73.2% | 71.5% | 90% | — | Apache 2.0 | |
| Seed-OSS-36BByteDance Seed | ByteDance Seed | 524K | $0.15 | $0.60 | — | 71.4% | 91.7% | — | Apache 2.0 | |
| GPT-5 nanoOpenAI | OpenAI | 400K | $0.05 | $0.40 | — | 71.2% | 85.2% | — | Proprietary | |
| Hunyuan-A13BTencent | Tencent | 262K | $0.30 | $1.00 | — | 71.2% | 87.3% | — | Tencent Hunyuan Community | |
| Qwen3-235B-A22BAlibaba Qwen | Alibaba Qwen | 131K | $0.20 | $0.60 | 68.2% | 71.1% | 85.7% | — | Apache 2.0 | |
| Magistral MediumMistral AI | Mistral AI | 41K | $2.00 | $5.00 | — | 70.8% | 73.6% | — | Proprietary | |
| Hermes 4 405BNous Research | Nous Research | 131K | $1.00 | $3.00 | — | 70.5% | 78.1% | — | Llama 3.1 Community | |
| Llama 4 MaverickMeta AI | Meta AI | 1M | $0.22 | $0.85 | 80.5% | 69.8% | — | — | Llama 4 Community | |
| Phi-4-reasoning-plusMicrosoft | Microsoft | 33K | $0.070 | $0.35 | — | 69.3% | 78% | — | MIT | |
| Magistral SmallMistral AI | Mistral AI | 41K | $0.50 | $1.50 | — | 68.2% | 70.7% | — | Apache 2.0 | |
| Claude 3.7 SonnetAnthropic | Anthropic | 200K | $3.00 | $15.00 | — | 68% | 61.3% | 62.3% | Proprietary | |
| Sonar Reasoning ProPerplexity | Perplexity | 127K | $2.00 | $8.00 | — | 68% | — | — | Proprietary | |
| Qwen3-32BAlibaba Qwen | Alibaba Qwen | 131K | $0.10 | $0.30 | — | 66.8% | 81.4% | — | Apache 2.0 | |
| EXAONE 4.0 32BLG AI Research | LG AI Research | 131K | $0.15 | $0.45 | — | 66.7% | 85.3% | — | EXAONE AI Model License | |
| Llama-3.3-Nemotron-Super-49BNVIDIA | NVIDIA | 131K | $0.13 | $0.40 | — | 66.7% | 67.5% | — | NVIDIA Open Model | |
| GPT-4.1OpenAI | OpenAI | 1M | $2.00 | $8.00 | 80.5% | 66.3% | — | 54.6% | Proprietary | |
| Grok 3 MinixAI | xAI | 131K | $0.30 | $0.50 | — | 66.2% | 90.8% | — | Proprietary | |
| Qwen3-30B-A3BAlibaba Qwen | Alibaba Qwen | 131K | $0.08 | $0.290 | — | 65.8% | 80.4% | — | Apache 2.0 | |
| QwQ-32BAlibaba Qwen | Alibaba Qwen | 131K | $0.15 | $0.20 | — | 65.2% | 79.5% | — | Apache 2.0 | |
| Claude 3.5 SonnetAnthropic | Anthropic | 200K | $3.00 | $15.00 | 78% | 65% | — | 49% | Proprietary | |
| GPT-4.1 miniOpenAI | OpenAI | 1M | $0.40 | $1.60 | — | 65% | — | 23.6% | Proprietary | |
| Gemini 2.5 Flash-LiteGoogle DeepMind | Google DeepMind | 1M | $0.10 | $0.40 | — | 64.6% | 49.8% | — | Proprietary | |
| Nemotron Nano 9B v2NVIDIA | NVIDIA | 131K | $0.04 | $0.16 | — | 64% | 72.1% | — | NVIDIA Open Model | |
| Qwen3-14BAlibaba Qwen | Alibaba Qwen | 131K | $0.06 | $0.24 | — | 64% | 79.3% | — | Apache 2.0 | |
| Mistral Medium 3Mistral AI | Mistral AI | 131K | $0.40 | $2.00 | 76% | 62.4% | — | — | Proprietary | |
| Gemini 2.0 FlashGoogle DeepMind | Google DeepMind | 1M | $0.10 | $0.40 | 77.6% | 62.1% | — | — | Proprietary | |
| DeepSeek-R1-Distill-Qwen-32BDeepSeek | DeepSeek | 131K | $0.12 | $0.18 | — | 62.1% | 72.6% | — | MIT | |
| Qwen3-8BAlibaba Qwen | Alibaba Qwen | 131K | $0.035 | $0.138 | — | 62% | 76% | — | Apache 2.0 | |
| o1-miniOpenAI | OpenAI | 128K | $1.10 | $4.40 | — | 60% | 63.6% | — | Proprietary | |
| Amazon Nova PremierAmazon | Amazon | 1M | $2.50 | $12.50 | 80% | 59.9% | — | — | Proprietary | |
| DeepSeek-V3DeepSeek | DeepSeek | 131K | $0.27 | $1.10 | 75.9% | 59.1% | — | — | MIT | |
| Gemini 1.5 ProGoogle DeepMind | Google DeepMind | 2.1M | $1.25 | $5.00 | 75.8% | 59.1% | — | — | Proprietary | |
| Llama 4 ScoutMeta AI | Meta AI | 10M | $0.11 | $0.34 | 74.3% | 57.2% | — | — | Llama 4 Community | |
| Phi-4Microsoft | Microsoft | 16K | $0.070 | $0.140 | 70.4% | 56.1% | — | — | MIT | |
| Grok 2xAI | xAI | 131K | $2.00 | $10.00 | 75.5% | 56% | — | — | Grok 2 Community | |
| Qwen3-4BAlibaba Qwen | Alibaba Qwen | 131K | $0.02 | $0.06 | — | 55.9% | 73.8% | — | Apache 2.0 | |
| GPT-4oOpenAI | OpenAI | 128K | $2.50 | $10.00 | 74.7% | 53.6% | — | — | Proprietary | |
| Gemini 2.0 Flash-LiteGoogle DeepMind | Google DeepMind | 1M | $0.075 | $0.30 | 71.6% | 51.5% | — | — | Proprietary | |
| Reka Flash 3Reka AI | Reka AI | 33K | $0.10 | $0.30 | — | 51.2% | 65% | — | Apache 2.0 | |
| Llama 3.1 405BMeta AI | Meta AI | 131K | $3.50 | $3.50 | 73.3% | 51.1% | — | — | Llama 3.1 Community | |
| Gemini 1.5 FlashGoogle DeepMind | Google DeepMind | 1M | $0.075 | $0.30 | 67.3% | 51% | — | — | Proprietary | |
| Llama 3.3 70BMeta AI | Meta AI | 131K | $0.23 | $0.40 | 68.9% | 50.5% | — | — | Llama 3.3 Community | |
| Claude 3 OpusAnthropic | Anthropic | 200K | $15.00 | $75.00 | 68.5% | 50.4% | — | — | Proprietary | |
| GPT-4.1 nanoOpenAI | OpenAI | 1M | $0.10 | $0.40 | — | 50.3% | — | — | Proprietary | |
| Qwen2.5-72BAlibaba Qwen | Alibaba Qwen | 131K | $0.35 | $0.40 | 71.1% | 49% | — | — | Qwen License | |
| Mistral Large 2Mistral AI | Mistral AI | 131K | $2.00 | $6.00 | 69.9% | 48% | — | — | Mistral Research | |
| GPT-4 TurboOpenAI | OpenAI | 128K | $10.00 | $30.00 | 63.7% | 48% | — | — | Proprietary | |
| Amazon Nova ProAmazon | Amazon | 300K | $0.80 | $3.20 | 75.5% | 46.9% | — | — | Proprietary | |
| Llama 3.1 70BMeta AI | Meta AI | 131K | $0.12 | $0.30 | 66.4% | 46.7% | — | — | Llama 3.1 Community | |
| Mistral Small 3.2 24BMistral AI | Mistral AI | 131K | $0.10 | $0.30 | 69.1% | 46.1% | — | — | Apache 2.0 | |
| Gemma 3 27BGoogle DeepMind | Google DeepMind | 131K | $0.10 | $0.20 | 67.5% | 42.4% | — | — | Gemma Terms of Use | |
| Claude 3.5 HaikuAnthropic | Anthropic | 200K | $0.80 | $4.00 | 65% | 41.6% | — | 40.6% | Proprietary | |
| GPT-4o miniOpenAI | OpenAI | 128K | $0.15 | $0.60 | 63.1% | 40.2% | — | — | Proprietary | |
| Gemma 3 12BGoogle DeepMind | Google DeepMind | 131K | $0.05 | $0.10 | 60.6% | 34.9% | — | — | Gemma Terms of Use | |
| Llama 3.1 8BMeta AI | Meta AI | 131K | $0.03 | $0.05 | 48.3% | 32.8% | — | — | Llama 3.1 Community | |
| Kimi K2 InstructMoonshot AI | Moonshot AI | 131K | $0.60 | $2.50 | 81.1% | — | — | 65.8% | Modified MIT | |
| ERNIE 4.5 300B-A47BBaidu | Baidu | 131K | $0.280 | $1.10 | 74% | — | — | — | Apache 2.0 | |
| Qwen2.5-32BAlibaba Qwen | Alibaba Qwen | 131K | $0.08 | $0.20 | 69% | — | — | — | Apache 2.0 | |
| Command ACohere | Cohere | 256K | $2.50 | $10.00 | 68% | — | — | — | CC-BY-NC | |
| Llama 3.2 90B VisionMeta AI | Meta AI | 131K | $0.35 | $0.40 | 68% | — | — | — | Llama 3.2 Community | |
| Solar Pro 2Upstage | Upstage | 66K | $0.50 | $0.50 | 66% | — | — | — | Proprietary | |
| Hunyuan-LargeTencent | Tencent | 262K | $0.50 | $1.50 | 60.2% | — | — | — | Tencent Hunyuan Community | |
| DeepSeek-Coder-V2DeepSeek | DeepSeek | 131K | $0.140 | $0.280 | 60% | — | — | — | DeepSeek License | |
| Jamba 1.6 LargeAI21 Labs | AI21 Labs | 256K | $2.00 | $8.00 | 60% | — | — | — | Jamba Open Model | |
| Yi-Large01.AI | 01.AI | 33K | $3.00 | $3.00 | 60% | — | — | — | Proprietary | |
| dots.llm1RedNote (Xiaohongshu) | RedNote (Xiaohongshu) | 33K | $0.20 | $0.60 | 60% | — | — | — | MIT | |
| Amazon Nova LiteAmazon | Amazon | 300K | $0.06 | $0.24 | 59% | — | — | — | Proprietary | |
| Falcon-H1 34BTII Falcon | TII Falcon | 262K | $0.15 | $0.45 | 58% | — | — | — | Falcon LLM License | |
| Qwen2.5-7BAlibaba Qwen | Alibaba Qwen | 131K | $0.025 | $0.05 | 56.3% | — | — | — | Apache 2.0 | |
| Command R+Cohere | Cohere | 128K | $2.50 | $10.00 | 56% | — | — | — | CC-BY-NC | |
| Gemma 2 27BGoogle DeepMind | Google DeepMind | 8.2K | $0.27 | $0.27 | 56% | — | — | — | Gemma Terms of Use | |
| Mixtral 8x22BMistral AI | Mistral AI | 66K | $0.90 | $0.90 | 56% | — | — | — | Apache 2.0 | |
| Reka CoreReka AI | Reka AI | 128K | $2.00 | $2.00 | 55% | — | — | — | Proprietary | |
| Phi-3.5-MoEMicrosoft | Microsoft | 131K | $0.08 | $0.16 | 54% | — | — | — | MIT | |
| Amazon Nova MicroAmazon | Amazon | 128K | $0.035 | $0.140 | 51.6% | — | — | — | Proprietary | |
| Mistral NeMo 12BMistral AI | Mistral AI | 131K | $0.03 | $0.070 | 50% | — | — | — | Apache 2.0 | |
| Yi-1.5-34B01.AI | 01.AI | 33K | $0.15 | $0.15 | 48% | — | — | — | Apache 2.0 | |
| Llama 3.2 11B VisionMeta AI | Meta AI | 131K | $0.055 | $0.055 | 47% | — | — | — | Llama 3.2 Community | |
| Ministral 8BMistral AI | Mistral AI | 131K | $0.10 | $0.10 | 47% | — | — | — | Mistral Research | |
| OLMo 2 32BAllen Institute | Allen Institute | 4.1K | $0.20 | $0.40 | 47% | — | — | — | Apache 2.0 | |
| DBRX InstructDatabricks | Databricks | 33K | $0.75 | $2.25 | 45% | — | — | — | Databricks Open Model | |
| GLM-4-9BZ.ai (Zhipu) | Z.ai (Zhipu) | 131K | $0.03 | $0.06 | 45% | — | — | — | GLM License | |
| Gemma 2 9BGoogle DeepMind | Google DeepMind | 8.2K | $0.06 | $0.06 | 45% | — | — | — | Gemma Terms of Use | |
| Granite 3.3 8BIBM | IBM | 131K | $0.03 | $0.06 | 45% | — | — | — | Apache 2.0 | |
| MiniCPM4 8BOpenBMB | OpenBMB | 33K | $0.02 | $0.05 | 45% | — | — | — | Apache 2.0 | |
| Falcon 3 10BTII Falcon | TII Falcon | 33K | $0.05 | $0.10 | 44% | — | — | — | Falcon LLM License | |
| Gemma 3 4BGoogle DeepMind | Google DeepMind | 131K | $0.02 | $0.04 | 43.6% | — | — | — | Gemma Terms of Use | |
| Jamba 1.6 MiniAI21 Labs | AI21 Labs | 256K | $0.20 | $0.40 | 43% | — | — | — | Jamba Open Model | |
| Command R7BCohere | Cohere | 128K | $0.037 | $0.15 | 42% | — | — | — | CC-BY-NC | |
| LFM2-8B-A1BLiquid AI | Liquid AI | 33K | $0.02 | $0.05 | 40% | — | — | — | LFM Open | |
| Snowflake ArcticSnowflake | Snowflake | 4.1K | $0.60 | $1.80 | 40% | — | — | — | Apache 2.0 | |
| GPT-3.5 TurboOpenAI | OpenAI | 16K | $0.50 | $1.50 | 38% | — | — | — | Proprietary | |
| OLMo 2 13BAllen Institute | Allen Institute | 4.1K | $0.10 | $0.20 | 35% | — | — | — | Apache 2.0 | |
| Llama 3.2 3BMeta AI | Meta AI | 131K | $0.015 | $0.025 | 33% | — | — | — | Llama 3.2 Community | |
| Falcon 180BTII Falcon | TII Falcon | 2K | $1.80 | $1.80 | 30% | — | — | — | Falcon 180B TII License | |
| Mistral 7BMistral AI | Mistral AI | 33K | $0.025 | $0.025 | 30% | — | — | — | Apache 2.0 | |
| Llama 3.2 1BMeta AI | Meta AI | 131K | $0.01 | $0.02 | 22% | — | — | — | Llama 3.2 Community | |
| OpenELM 3BApple | Apple | 2K | $0.01 | $0.02 | 20% | — | — | — | Apple Sample Code License | |
| Claude Fable 5Anthropic | Anthropic | 1M | $10.00 | $50.00 | — | — | — | — | Proprietary | |
| Claude Fable 5.1Anthropic | Anthropic | 1M | $10.00 | $50.00 | — | — | — | — | Proprietary | |
| Claude Mythos 5Anthropic | Anthropic | 1M | $10.00 | $50.00 | — | — | — | — | Proprietary | |
| Claude Mythos 5.1Anthropic | Anthropic | 1M | $10.00 | $50.00 | — | — | — | — | Proprietary | |
| Claude Opus 5Anthropic | Anthropic | 1M | $5.00 | $25.00 | — | — | — | — | Proprietary | |
| Claude Sonnet 5Anthropic | Anthropic | 1M | $2.00 | $10.00 | — | — | — | — | Proprietary | |
| Codestral 25.08Mistral AI | Mistral AI | 262K | $0.30 | $0.90 | — | — | — | — | Mistral AI Non-Production | |
| CorX1.5CorX Labs | CorX Labs | 1K | — | — | — | — | — | — | Apache 2.0 | |
| CorX3.8-27BCorX Labs | CorX Labs | 33K | — | — | — | — | — | — | Apache 2.0 | |
| DeepSeek-R1-Distill-Llama-8BDeepSeek | DeepSeek | 131K | $0.04 | $0.04 | — | — | 50.4% | — | MIT | |
| Devstral MediumMistral AI | Mistral AI | 131K | $0.40 | $2.00 | — | — | — | 61.6% | Proprietary | |
| ERNIE X1Baidu | Baidu | 131K | $0.280 | $1.10 | — | — | — | — | Proprietary | |
| GLM-5.3Z.ai (Zhipu) | Z.ai (Zhipu) | 1M | $1.40 | $4.40 | — | — | — | — | GLM-5.3 License | |
| Grok Code Fast 1xAI | xAI | 256K | $0.20 | $1.50 | — | — | — | 70.8% | Proprietary | |
| HyperCLOVA X SEED 14BNaver | Naver | 33K | $0.08 | $0.16 | — | — | — | — | HyperCLOVA X SEED License | |
| Kimi K3Moonshot AI | Moonshot AI | 1M | $3.00 | $15.00 | — | — | — | — | Kimi K3 License | |
| Kimi-Dev-72BMoonshot AI | Moonshot AI | 131K | $0.290 | $1.15 | — | — | — | 60.4% | Modified MIT | |
| Ling-1TInclusionAI (Ant) | InclusionAI (Ant) | 131K | $0.50 | $2.00 | — | — | 70.4% | — | MIT | |
| Mercury CoderInception Labs | Inception Labs | 33K | $0.25 | $1.00 | — | — | — | — | Proprietary | |
| MiniMax-M1MiniMax | MiniMax | 1M | $0.40 | $2.10 | — | — | 86% | 56% | Apache 2.0 | |
| Molmo 72BAllen Institute | Allen Institute | 4.1K | $0.35 | $0.40 | — | — | — | — | Apache 2.0 | |
| Palmyra X5Writer | Writer | 1M | $0.60 | $6.00 | — | — | — | — | Proprietary | |
| Phi-4-multimodalMicrosoft | Microsoft | 131K | $0.05 | $0.10 | — | — | — | — | MIT | |
| Pixtral LargeMistral AI | Mistral AI | 131K | $2.00 | $6.00 | — | — | — | — | Mistral Research | |
| Qwen2.5-Coder-32BAlibaba Qwen | Alibaba Qwen | 131K | $0.070 | $0.16 | — | — | — | — | Apache 2.0 | |
| Qwen2.5-VL-72BAlibaba Qwen | Alibaba Qwen | 131K | $0.40 | $0.40 | — | — | — | — | Qwen License | |
| Qwen3-Coder-480B-A35BAlibaba Qwen | Alibaba Qwen | 262K | $0.30 | $1.20 | — | — | — | 69.6% | Apache 2.0 | |
| R1-1776Perplexity | Perplexity | 131K | $2.00 | $8.00 | — | — | 78% | — | MIT | |
| Sarvam-MSarvam AI | Sarvam AI | 33K | $0.10 | $0.30 | — | — | — | — | Apache 2.0 | |
| Sonar ProPerplexity | Perplexity | 200K | $3.00 | $15.00 | — | — | — | — | Proprietary | |
| Step-3StepFun | StepFun | 66K | $0.30 | $1.20 | — | — | — | — | Apache 2.0 | |
| TriStream-SVSCorX LabsSinging voice synthesis | CorX Labs | — | — | — | — | — | — | — | Apache 2.0 | |
| xLAM-2-70BSalesforce AI | Salesforce AI | 131K | $0.30 | $0.60 | — | — | — | — | CC-BY-NC |
The tests
What each benchmark actually measures
A score is only useful if you know what it was measuring. These are the 14 evaluations tracked here, and what each one does and does not tell you.
MMLU-Pro
12,000 reasoning-heavy multiple-choice questions across 14 academic subjects, with ten options instead of four. The harder successor to MMLU.
Read the testGPQA Diamond
198 graduate-level physics, chemistry and biology questions written to be Google-proof. PhD holders in the matching field score about 65%.
Read the testAIME 2025
The American Invitational Mathematics Examination — 15 problems, integer answers, no partial credit. A standard test of multi-step maths reasoning.
Read the testMATH-500
500 competition maths problems sampled from the MATH benchmark, graded on the final answer.
Read the testSWE-bench Verified
500 human-validated GitHub issues from real Python repositories. The model must produce a patch that makes the project's own tests pass.
Read the testSWE-bench Pro
A harder, contamination-resistant successor to SWE-bench Verified, drawn from commercial and copyleft repositories that were never public training data.
Read the testTerminal-Bench 2.1
The 2.1 revision of the terminal agent benchmark. Scores on it are not comparable with Terminal-Bench 4.0 — the task set changed.
Read the testFrontier-Bench v0.1
Novel problems built to resist memorisation, scored on whether the model gets anywhere at all. Absolute numbers are low by design.
Read the testTerminal-Bench 4.0
End-to-end tasks in a real terminal — install, build, debug, run — scored on whether the machine ends up in the required state.
Read the testLiveCodeBench
Competitive-programming problems collected after each model's training cutoff, so contamination cannot inflate the score.
Read the testHumanEval
164 short Python functions written from a docstring. Saturated at the frontier — kept here for continuity with older models.
Read the testMMMU
College-level questions that require reading charts, diagrams, tables and photographs alongside the text.
Read the testIFEval
Verifiable instructions — word counts, formats, forbidden words — checked by a program rather than a judge model.
Read the testLMArena Elo
Elo rating from blind pairwise votes by the public on LMArena. Measures what people prefer, not what is correct.
Read the testHead to head
The comparisons people actually make
By maker
42 labs
Method, and what to distrust
This index is a collection of published figures, not an independent evaluation. That distinction matters more than it sounds, so here is exactly what was and was not done.
What is collected
For every model: the context window and maximum output length, the input and output modalities, the licence and whether weights are downloadable, the first-party API price per million tokens, the architecture where the lab has disclosed it, and the release month. For scores: whatever the lab published in its model card, system card, technical report or launch post, plus LMArena Elo where the model has been rated.
What is not done
No evaluation was re-run. No score was estimated, interpolated from a sibling model, or carried over from a previous version. Where a lab has not published a figure the cell says Not reported and stays empty, even when that leaves a gap in an otherwise full row.
Why self-reported scores are slippery
- The scaffold moves the number. SWE-bench Verified in particular is a measure of a whole agent — retrieval, retries, test execution — not of a model alone. Two labs reporting the same benchmark may be running very different harnesses.
- Attempts vary. A score taken at pass@1 and one taken with majority voting over many samples are not comparable, and the difference is often larger than the gap between two models.
- Contamination. Older benchmarks leak into training data over time. HumanEval is effectively saturated; LiveCodeBench exists precisely because it collects problems published after a model's cutoff.
- Reasoning budgets. A model with adjustable thinking can post a much higher score at a much higher cost per answer. The price column does not capture that, because tokens spent thinking are billed as output.
Prices
Prices are the standard first-party rate per million tokens, excluding batch discounts and cached-input rates unless noted on the model's own page. For open-weight models there is no first-party price, so the figure shown is a representative third-party hosting rate and is marked as such — the weights themselves are free to download.
Corrections
Every model records the month its row was last verified. If a figure is wrong or has been superseded, send a correction with a link to the source and it will be updated.
Trademarks
Company marks are shown to identify each lab's own models. All trademarks belong to their respective owners; CorX Labs is not affiliated with, endorsed by, or sponsored by any of the companies listed. Logo files are from the lobe-icons (MIT) and simple-icons (CC0) sets.
Questions
Common questions
What is the best AI model right now?
There is no single answer, which is why this page is a table rather than a ranking. The frontier models — GPT-5, Claude Opus 4.5, Gemini 3 Pro and Grok 4 — trade places depending on the test: reasoning benchmarks like GPQA Diamond, agentic coding benchmarks like SWE-bench Verified, and human preference on LMArena all pick different winners. Sort the table by the column that matches the work you are actually doing.
Which AI model is cheapest?
Among capable models, the open-weight ones hosted by third parties are usually cheapest — DeepSeek, Qwen, GLM and gpt-oss all sit far below the frontier proprietary models. Sort by In / M or Out / M to see the current order. Note that reasoning models generate many more output tokens than their price per token suggests, so a cheap reasoning model can cost more per answer than an expensive non-reasoning one.
What does open weights mean?
The lab has published the trained parameters, so you can download the model and run it on your own hardware. It does not necessarily mean the training data or code is public, and it does not always mean unrestricted commercial use — the Licence column records the actual terms, which range from Apache 2.0 and MIT through to non-commercial and custom community licences.
Did CorX Labs run these benchmarks?
No. Every score here is a published figure from the lab that built the model or from a public leaderboard, collected and put in one table. Most benchmark numbers in this industry are self-reported, and the scaffolding around a model can move a score by more than the difference between two models. Use them to build a shortlist, then test the shortlist on your own task.
How often is this updated?
Each model row records the month it was last checked against its sources. New models are added as they are released. If you spot a figure that is out of date or wrong, tell us and it will be corrected.
Can I compare more than two models?
Yes. Tick the boxes in the leaderboard and press Compare, or open the comparison tool and add up to four models side by side.
CorX Labs
We build models too.
CorX3.8-27B is Jamaica's first large open-weight LLM — a 27B Jamaican Patois assistant with open weights under Apache 2.0. It is in this index like everything else.