Prime is a team.
Cite
Notes
Only stored in your browser.
Prime Forge 166 synthetic AutomationBench tasks with V1 tool contract and PR #764 fixes.
Sandboxed Python-REPL harness for training models to manage their own context window across turns.
Chess environment where an agent plays as White against configurable opponents (random, LLM, or Stockfish)
Zapier AutomationBench business workflows with task-scoped, structured API discovery and execution.
Agentic antiviral design against Bundibugyo ebolavirus GP entry — sandboxed Muni/OnePot tool use with onepot-CORE plausibility reward
Respond to a USPTO office action against a simulated examiner: multi-turn patent prosecution with retrieval tools, a per-rejection examiner state m...
V1 Taskset/Harness environment training LangChain deep-agents on Wikispeedia navigation
PMPP CUDA evaluation environment with local and FastAPI eval modes
Multi-turn visual click calibration tasks with click and computer tool formats across pixel and normalized coordinate schemas.
GSM8K environment
Unified backdoor-ifeval env: difficulty, aggregation, no-v check, inoculation, group monitors
Reverse text character by character.
Multimodal aim training environment where agents click targets in images. Demonstrates visual reasoning with coordinate-based responses.
Reward hacking with deterministic IF constraints
Cross-repo code-search tasks over prime-rl, verifiers, vllm, pytorch
Backdoor-ifeval env with group-level reward monitors for within-batch advantage variance
Backdoor-ifeval env for inoculation experiments (pre-no-v version)
OpenRCA root cause analysis benchmark environment for Verifiers (ICLR 2025)
BrowserEnv demo for web browsing tasks using Browserbase
Multi-turn web-search QA environment with Exa-style benchmark support
Tic-Tac-Toe environment with configurable text or image board state and random opponent
Just another swe grep environment
τ²-bench with custom synthetic domains (library, fitness_gym, tech_support, telecom, cloud_incident_response, daily_planner, ev_charging_support)
Stateful tool-based environment for constrained meeting scheduling
WebVoyager browser benchmark with filtered dataset (600 tasks from sites without anti-bot protection)