0

Claude 4 Sonnet vs GPT-4o

Head to head on 20 shared benchmarks, with price, context window, and release dates.

Benchmark scores

53

evals · sort by any column, ⤢ to expand fullscreen

Model
Claude 4 Sonnet

Anthropic

40.7%38.0%54.7%19.8%68.3%4.3%45.4%64.8%77.5%44.3%72.9%44.9%93.4%83.7%62.5%37.3%74.6%69.6%27.3%52.3%---22.9%47.8%-1480elo58.4-67.5%--54.6%70.5%69.0%33.9%-72.4%86.7%43.9%-6.186.1835.0%-58.3%35.6%--0.5984.0%84.0%-55.5%
GPT-4o

OpenAI

15.0%6.0%45.9%6.1%54.3%2.4%34.3%55.1%46.1%70.1%49.0%30.9%75.9%74.8%57.4%33.3%21.6%74.5%8.3%25.1%18.2%380.0%--0.33--0.3%-96.10.0%----10.8%---54.0%---5.71--61.20.98---32.5%33.4%
2 / 2 models
AttributeClaude 4 Sonnet

Anthropic

GPT-4o

OpenAI

DescriptionClaude 4 Sonnet is an AI model from Anthropic.GPT-4o is an AI model from OpenAI.
FamilyClaudeGPT
Params--
Open weightsclosedclosed
Licenseproprietaryproprietary
Released22 May 202520 Nov 2024
Context200,000128,000
Input $/1M$3.00$2.50
Output $/1M$15.00$10.00
Cache read $/1M-$1.250
Cache write $/1M--
Speed0 tok/s0 tok/s
Latency0.00 s0.00 s
Modalitiestext, imagetext, image, file
Providers--
PublisherAnthropicOpenAI
Reported scores
Taubench61.2 pass^1
GSM8K96.1 Accuracy
ScholarSearch5.71 Social Sciences & Humanities (%)
AIME202438 pass@1

FAQ

Is Claude 4 Sonnet or GPT-4o better?
Across the 20 benchmarks both models report on Sophon, Claude 4 Sonnet leads on 18 of them. Which one is "better" depends on the benchmark - the table above shows every shared score.
How do Claude 4 Sonnet and GPT-4o compare on GPQA Diamond?
Claude 4 Sonnet scores 68.3% and GPT-4o scores 54.3% on GPQA Diamond.
Which is cheaper, Claude 4 Sonnet or GPT-4o?
GPT-4o is cheaper at $10.00 per million output tokens against $15.00 for Claude 4 Sonnet - about 1.5x.
When were Claude 4 Sonnet and GPT-4o released?
Claude 4 Sonnet was released 22 May 2025 by Anthropic. GPT-4o was released 20 Nov 2024 by OpenAI.

Comparing something else? Build your own side-by-side.