Terminal-Bench Science
70 command-line research tasks across five scientific fields, run as an agent in a terminal.
- Domain
- science
- Published
- Sep 2026
- Notable for
- Scientific agentic-terminal benchmark reported by both OpenAI and Anthropic in 2026.
Cite
Notes
Only stored in your browser.
Top score 64.6% by GPT-6 Astra - 1 model reporting (1 frontier)
Top models
1FAQ
- What is Terminal-Bench Science?
- 70 command-line research tasks across five scientific fields, run as an agent in a terminal.
- What is the current top score on Terminal-Bench Science?
- The top reported score is 64.6% by GPT-6 Astra, across 1 model reporting (1 from frontier labs).