i3-science
Science tasks solved by an agent in a sandbox. Each task poses a science question with a boxed final answer; scored with math-verify against the gold answer, with an LLM-judge fallback when math-verify can't confirm it.
Taskset
- Source: PrimeIntellect/INTELLECT-3-RL
- Size: 29,307 tasks
Changelog
- 2026-09-03: Restore default solver network access by reverting the
network_allow=[]default-deny policy introduced in #780; training rollouts need outbound network. - 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.
- 2026-06-25: Initial v1 taskset.