DeepSWE
Agentic software-engineering benchmark; the v1.1 split runs 113 tasks and is scored on end-to-end task completion.
- Domain
- coding
- Published
- Sep 2026
- Notable for
- Headline agentic-coding benchmark at the GPT-6 Astra launch.
Cite
Notes
Only stored in your browser.
Top score 74.1% by GPT-6 Astra - 1 model reporting (1 frontier)
Top models
1FAQ
- What is DeepSWE?
- Agentic software-engineering benchmark; the v1.1 split runs 113 tasks and is scored on end-to-end task completion.
- What is the current top score on DeepSWE?
- The top reported score is 74.1% by GPT-6 Astra, across 1 model reporting (1 from frontier labs).