0

Claude 4 Sonnet vs GPT-4.1

Head to head on 19 shared benchmarks, with price, context window, and release dates.

Benchmark scores

59

evals · sort by any column, ⤢ to expand fullscreen

Model
Claude 4 Sonnet

Anthropic

40.7%38.0%54.7%19.8%1480elo67.5%68.3%4.3%45.4%44.9%93.4%83.7%62.5%37.3%35.6%74.6%69.6%27.3%52.3%-----22.9%47.8%--58.4-64.8%77.5%54.6%44.3%72.9%70.5%69.0%-----33.9%-72.4%86.7%-43.9%6.186.1835.0%-58.3%-0.5984.0%84.0%--55.5%
GPT-4.1

OpenAI

43.7%34.7%63.1%12.4%1417elo60.9%66.6%4.2%43.0%45.7%91.3%80.6%65.9%38.1%31.1%39.6%75.1%13.6%47.1%0.0%52.4%26.7%61.5%1.64--83.7%1.55-5.5%-------70.8%0.490.4932.1%32.1%-69.4%--50.5%----8.57-10.0%---92.0%92.0%48.0%
2 / 2 models
AttributeClaude 4 Sonnet

Anthropic

GPT-4.1

OpenAI

DescriptionClaude 4 Sonnet is an AI model from Anthropic.GPT-4.1 is an AI model from OpenAI.
FamilyClaudeGPT
Params--
Open weightsclosedclosed
Licenseproprietaryproprietary
Released22 May 202514 Apr 2025
Context200,0001,047,576
Input $/1M$3.00$2.00
Output $/1M$15.00$8.00
Cache read $/1M-$0.500
Cache write $/1M--
Speed0 tok/s0 tok/s
Latency0.00 s0.00 s
Modalitiestext, imageimage, text, file
Providers--
PublisherAnthropicOpenAI
Reported scores

FAQ

Is Claude 4 Sonnet or GPT-4.1 better?
Across the 19 benchmarks both models report on Sophon, Claude 4 Sonnet leads on 13 of them. Which one is "better" depends on the benchmark - the table above shows every shared score.
How do Claude 4 Sonnet and GPT-4.1 compare on GPQA Diamond?
Claude 4 Sonnet scores 68.3% and GPT-4.1 scores 66.6% on GPQA Diamond.
Which is cheaper, Claude 4 Sonnet or GPT-4.1?
GPT-4.1 is cheaper at $8.00 per million output tokens against $15.00 for Claude 4 Sonnet - about 1.9x.
When were Claude 4 Sonnet and GPT-4.1 released?
Claude 4 Sonnet was released 22 May 2025 by Anthropic. GPT-4.1 was released 14 Apr 2025 by OpenAI.

Comparing something else? Build your own side-by-side.