0

DeepSWE

Agentic software-engineering benchmark; the v1.1 split runs 113 tasks and is scored on end-to-end task completion.

Domain
coding
Published
Sep 2026
Notable for
Headline agentic-coding benchmark at the GPT-6 Astra launch.

Cite

Notes

Only stored in your browser.

Attribution

Leaderboard scores
OpenAI
Attribution policy →

Top score 74.1% by GPT-6 Astra - 1 model reporting (1 frontier)

Top models

1
DeepSWEBar chart with 1 bar. Highest value: GPT-6 Astra at 74.1.
1 model

FAQ

What is DeepSWE?
Agentic software-engineering benchmark; the v1.1 split runs 113 tasks and is scored on end-to-end task completion.
What is the current top score on DeepSWE?
The top reported score is 74.1% by GPT-6 Astra, across 1 model reporting (1 from frontier labs).