0

Claude Sonnet 4.5 vs gpt-oss-120b

Head to head on 23 shared benchmarks, with price, context window, and release dates.

Benchmark scores

90

evals · sort by any column, ⤢ to expand fullscreen

Model
gpt-oss-120b

OpenAI

66.7%58.2%959elo35.867.2%5.9%58.3%51.0%60.2%38.8%50.3%48.6%68.9%39.2%70.7%1.621.5377.5%36.0%26.0%71.6%5.3%45.0%--41.8%-----38.9%7.6%---13.0%-----------57.6-----------80.0%--------------97.9%---------65.6%----49.6%
Claude Sonnet 4.5

Anthropic

37.0%60.8%1675elo71.472.7%7.2%42.7%70.7%80.4%57.0%53.4%76.5%79.3%77.6%59.0%0.750.7586.0%42.8%74.8%73.3%28.8%70.5%0.0%0.0%-1000.0%0.0%0.0%5.0%--0.0%0.0%0.0%-0.0%0.0%0.0%0.0%15.2%0.0%0.0%83.9%0.0%0.0%0.0%-0.0%0.0%87.5%83.8%0.0%0.0%0.0%040.6%84.5%0-64.0%0.0%0.0%62.9%0.0%0.0%19.0%0.0%0.0%0.0%0.0%0.0%32.9%0.0%-00.0%0.0%77.2%67.0%0.0%0050.34-022.6%0.0%1.4927.5%
2 / 2 models
AttributeClaude Sonnet 4.5

Anthropic

gpt-oss-120b

OpenAI

Descriptionanthropic/claude-sonnet-4.5 is an AI model.GPT-Oss 120b is an AI model from OpenAI, released with open weights.
FamilyClaudeGPT
ParamsUndisclosed (closed)-
Open weightsclosedopen
LicenseProprietaryapache-2.0
Released29 Sep 20255 Aug 2025
Context1,000,000131,072
Input $/1M$3.00$0.15
Output $/1M$15.00$0.60
Cache read $/1M$0.300$0.030
Cache write $/1M$3.750-
Speed0 tok/s186 tok/s
Latency0.00 s0.50 s
Modalitiestext, image, filetext
Providersanthropic, aws-bedrock, vertex-ai-
PublisherAnthropicOpenAI
Reported scores
LiveBench - Math79.3% (LiveBench)
LiveBench - Coding80.4% (LiveBench)
LiveBench - Language76.5% (LiveBench)
LiveBench - Data Analysis57.0% (LiveBench)
GPQA Diamond67.2% (AA)
LiveCodeBench70.7% (AA)
MMLU-Pro77.5% (AA)
SciCode36.0% (AA)

FAQ

Is Claude Sonnet 4.5 or gpt-oss-120b better?
Across the 23 benchmarks both models report on Sophon, Claude Sonnet 4.5 leads on 18 of them. Which one is "better" depends on the benchmark - the table above shows every shared score.
How do Claude Sonnet 4.5 and gpt-oss-120b compare on GPQA Diamond?
Claude Sonnet 4.5 scores 72.7% and gpt-oss-120b scores 67.2% on GPQA Diamond.
Which is cheaper, Claude Sonnet 4.5 or gpt-oss-120b?
gpt-oss-120b is cheaper at $0.60 per million output tokens against $15.00 for Claude Sonnet 4.5 - about 25.0x.
When were Claude Sonnet 4.5 and gpt-oss-120b released?
Claude Sonnet 4.5 was released 29 Sep 2025 by Anthropic. gpt-oss-120b was released 5 Aug 2025 by OpenAI.

Comparing something else? Build your own side-by-side.