CorX Labs

GPT-4.1

by OpenAI · United States
ProprietaryVisionTool calling1M context

A million-token non-reasoning model built for long-context retrieval and instruction following.


Specification

The numbers

Maker
OpenAI
Released
2025-04
Parameters
Not reported
Architecture
Not reported
Context window
1,047,576 tokens
Max output
32,768 tokens
Input
Text, Image
Output
Text
Reasoning
No
Tool calling
Yes
Knowledge cutoff
2024-06
Licence
Proprietary
Availability
Not reported
Weights
Not released

Cost

Price per million tokens

Input
$2.00 / M tokens
Output
$8.00 / M tokens
Cached input
$0.50
Blended 3:1
$3.50

Standard first-party API rate, excluding batch discounts. Reasoning models bill thinking tokens as output, so cost per answer can far exceed the cost per token suggests.

Published scores

Benchmarks

Figures published by OpenAI or taken from a public leaderboard. Row last checked 2026-08.

GPT-4.1 by capability category, with its rank among models reporting the same tests.
CategoryScore Rank
ReasoningGPQA Diamond66.3%44 of 83 reporting the same tests
MathsCompetition mathematics, graded on the final answerNot reported
CodingSWE-bench Verified54.6%27 of 33 reporting the same tests
KnowledgeMMLU-Pro80.5%6 of 77 reporting the same tests
MultimodalMMMU74.8%10 of 39 reporting the same tests
Instruction followingIFEval87.4%9 of 15 reporting the same tests
Human preferenceWhich answer people pick, blindNot reported

A category averages every benchmark in it that GPT-4.1 reports. The rank counts only models that report the same tests, so it never compares an average over three benchmarks against an average over one.

Every reported test

MMLU-ProKnowledge80.5%Rank 6 of 77 models reporting
GPQA DiamondReasoning66.3%Rank 44 of 83 models reporting
AIME 2025MathsNot reportedNo figure published
MATH-500MathsNot reportedNo figure published
SWE-bench VerifiedCoding54.6%Rank 27 of 33 models reporting
SWE-bench ProCodingNot reportedNo figure published
Terminal-Bench 2.1CodingNot reportedNo figure published
Frontier-Bench v0.1ReasoningNot reportedNo figure published
Terminal-Bench 4.0CodingNot reportedNo figure published
LiveCodeBenchCodingNot reportedNo figure published
HumanEvalCodingNot reportedNo figure published
MMMUMultimodal74.8%Rank 10 of 39 models reporting
IFEvalInstruction following87.4%Rank 9 of 15 models reporting
LMArena EloHuman preferenceNot reportedNo figure published

Where these numbers come from

Every score on this page is a published figure, taken from the model's own card, system card, technical report or release post, or from a public leaderboard. CorX Labs did not run these evaluations. Most are self-reported by the lab that built the model, which means they were produced under that lab's own choice of prompt, scaffold and number of attempts — so treat them as a starting point for a shortlist, not as a settled ranking.

A score someone other than the model's maker measured is marked Independent and names its measurer. Those are the stronger numbers on this page — an outside harness has no reason to flatter anyone — and there are not many of them.

Where a figure has not been published, the cell reads Not reported rather than an estimate. Nothing here is inferred, interpolated or guessed. Each model records the month its row was last checked. Full method and caveats.

Same maker

Other models from OpenAI

AI models with context window, price per million tokens and published benchmark scores. Sortable by any column.
GPT-5OpenAIOpenAI400K$1.25$10.0085.7%94.6%74.9%Proprietary
GPT-5 miniOpenAIOpenAI400K$0.25$2.0082.3%91.1%71%Proprietary
GPT-5 nanoOpenAIOpenAI400K$0.05$0.4071.2%85.2%Proprietary
gpt-oss-120bOpenAIOpenAI131K$0.10$0.5080.9%80.1%92.5%Apache 2.0
gpt-oss-20bOpenAIOpenAI131K$0.05$0.2073.2%71.5%90%Apache 2.0
o3OpenAIOpenAI200K$2.00$8.0083.3%88.9%69.1%Proprietary
o4-miniOpenAIOpenAI200K$1.10$4.4081.4%92.7%68.1%Proprietary
GPT-4.1 miniOpenAIOpenAI1M$0.40$1.6065%23.6%Proprietary
GPT-4.1 nanoOpenAIOpenAI1M$0.10$0.4050.3%Proprietary
o3-miniOpenAIOpenAI200K$1.10$4.4079.7%87.3%49.3%Proprietary
o1OpenAIOpenAI200K$15.00$60.0078%79.2%48.9%Proprietary
o1-miniOpenAIOpenAI128K$1.10$4.4060%63.6%Proprietary
GPT-4o miniOpenAIOpenAI128K$0.15$0.6063.1%40.2%Proprietary
GPT-4oOpenAIOpenAI128K$2.50$10.0074.7%53.6%Proprietary
GPT-4 TurboOpenAIOpenAI128K$10.00$30.0063.7%48%Proprietary
GPT-3.5 TurboOpenAIOpenAI16K$0.50$1.5038%Proprietary