0

S1 Deepresearch

Fresh

S1-DeepResearch closed-ended deep-research QA; the agent answers in chat (bring your own web search), graded by a binary LLM judge.

Type
RL Env
Tags
V1
Runtime
single-turn
License
unknown
Size
v0.1.0
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

s1-deepresearch

S1-DeepResearch closed-ended deep-research QA: the agent answers in chat (bring your own web search) and is graded by a binary LLM judge with [CORRECT]/[INCORRECT] semantics. Only the verifiable "Closed-ended Multi-hop Resolution" rows with a non-empty gold answer are kept.

Taskset

  • Source: ScienceOne-AI/S1-DeepResearch-15k
  • Size: the closed-ended verifiable subset of ~15,000 trajectory rows (open-ended exploration rows have no gradable gold and are dropped)

Changelog

  • 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.
  • 2026-07-31: Moved the semantic-grading prompt into a packaged, environment-owned reference judge for verifiers>=0.2.2.dev65; saved configs now carry the judge ID instead of a checkout-specific prompt path.
  • 2026-07-02: Initial v1 taskset.