0

Terminal-Bench Science

70 command-line research tasks across five scientific fields, run as an agent in a terminal.

Domain
science
Published
Sep 2026
Notable for
Scientific agentic-terminal benchmark reported by both OpenAI and Anthropic in 2026.

Cite

Notes

Only stored in your browser.

Attribution

Leaderboard scores
OpenAI
Attribution policy →

Top score 64.6% by GPT-6 Astra - 1 model reporting (1 frontier)

Top models

1
Terminal-Bench ScienceBar chart with 1 bar. Highest value: GPT-6 Astra at 64.6.
1 model

FAQ

What is Terminal-Bench Science?
70 command-line research tasks across five scientific fields, run as an agent in a terminal.
What is the current top score on Terminal-Bench Science?
The top reported score is 64.6% by GPT-6 Astra, across 1 model reporting (1 from frontier labs).