s1-deepresearch
S1-DeepResearch closed-ended deep-research QA: the agent answers in chat (bring your own web search) and is graded by a binary LLM judge with [CORRECT]/[INCORRECT] semantics. Only the verifiable "Closed-ended Multi-hop Resolution" rows with a non-empty gold answer are kept.
Taskset
- Source:
ScienceOne-AI/S1-DeepResearch-15k - Size: the closed-ended verifiable subset of ~15,000 trajectory rows (open-ended exploration rows have no gradable gold and are dropped)
Changelog
- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.
- 2026-07-31: Moved the semantic-grading prompt into a packaged, environment-owned reference judge for
verifiers>=0.2.2.dev65; saved configs now carry the judge ID instead of a checkout-specific prompt path. - 2026-07-02: Initial v1 taskset.