0

Claude Opus 4.5 vs gpt-oss-120b

Head to head on 21 shared benchmarks, with price, context window, and release dates.

Benchmark scores

51

evals · sort by any column, ⤢ to expand fullscreen

Model
Claude Opus 4.5

Anthropic

62.7%61.3%1683elo73.181.0%13.2%43.0%78.1%79.7%74.4%62.5%81.3%90.4%80.1%73.8%88.9%47.0%79.2%74.3%40.9%86.3%---16.7%-35.6%6.1%20.7%84.7%12.13-83.8%83.5%--45.2%83.2%-68.7%100.0%37.436.0%77.5%45.0%45.3-70.7%61.8-20.6%62.2%
gpt-oss-120b

OpenAI

66.7%58.2%959elo35.867.2%5.9%58.3%51.0%60.2%38.8%50.3%48.6%68.9%39.2%70.7%77.5%36.0%26.0%71.6%5.3%45.0%41.8%38.9%7.6%-13.0%-----57.6--1.621.53--80.0%-------97.9%--65.6%-49.6%
2 / 2 models
AttributeClaude Opus 4.5

Anthropic

gpt-oss-120b

OpenAI

DescriptionClaude Opus 4.5 is an AI model from Anthropic.GPT-Oss 120b is an AI model from OpenAI, released with open weights.
FamilyClaudeGPT
Params--
Open weightsclosedopen
Licenseproprietaryapache-2.0
Released24 Nov 20255 Aug 2025
Context200,000131,072
Input $/1M$5.00$0.15
Output $/1M$25.00$0.60
Cache read $/1M$0.500$0.030
Cache write $/1M$6.250-
Speed0 tok/s186 tok/s
Latency0.00 s0.50 s
Modalitiesfile, image, texttext
Providers--
PublisherAnthropicOpenAI
Reported scores
GPQA Diamond67.2% (AA)
LiveCodeBench70.7% (AA)
MMLU-Pro77.5% (AA)
SciCode36.0% (AA)

FAQ

Is Claude Opus 4.5 or gpt-oss-120b better?
Across the 21 benchmarks both models report on Sophon, Claude Opus 4.5 leads on 19 of them. Which one is "better" depends on the benchmark - the table above shows every shared score.
How do Claude Opus 4.5 and gpt-oss-120b compare on GPQA Diamond?
Claude Opus 4.5 scores 81.0% and gpt-oss-120b scores 67.2% on GPQA Diamond.
Which is cheaper, Claude Opus 4.5 or gpt-oss-120b?
gpt-oss-120b is cheaper at $0.60 per million output tokens against $25.00 for Claude Opus 4.5 - about 41.7x.
When were Claude Opus 4.5 and gpt-oss-120b released?
Claude Opus 4.5 was released 24 Nov 2025 by Anthropic. gpt-oss-120b was released 5 Aug 2025 by OpenAI.

Comparing something else? Build your own side-by-side.