Prime Community is a team.
Cite
Notes
Only stored in your browser.
Multi-turn text-based environment for evaluating agents on the Spiral-Bench dataset.
TransformerPuzzles by Sasha Rush
AgentHarm environment to evaluate agentic reasoning and safety
Data Agent Benchmark for Multi-step Reasoning benchmark
Hanabi game
Big Bench + BBH implementation
Verifiers port of the minif2f benchmark
LLM Training Puzzles by Sasha Rush
GitHub MCP environment
text-based muli-turn fruit box game environment
ARC-AGI 1 + 2 with tool calling (Abstract and Reasoning Corpus)
BALROG benchmark integration for verifiers: unified RL evaluation across game environments.
ART-E: a tool-using email research RL environment for Verifiers
SciCode evaluation environment
BixBench scientific reasoning evaluation environment
Mastermind multi-turn game environment for Verifiers
Humanity's Last Examination (HLE) benchmark environment for Prime Community Environments
MCP Universe environment for evaluating LLMs in wide range of tasks with MCP server
GPU puzzles environment by Sasha Rush using modal sandboxes
BackendBench environment for LLM kernel benchmarking
Polars DataFrame manipulation environment for training and evaluation
LegalBench environment for legal reasoning tasks
ENV for self-grading for LLM Writer Style.
Agentic RAG over Sherlock Holmes short stories for literary Q&A
MMLU evaluator for multi-subject multiple-choice reasoning.
ClockBench: multimodal clock reading and reasoning benchmark implemented for verifiers.
A multi-turn RL environment for formal theorem proving in Lean 4, where models alternate between reasoning, sketching proof code, and receiving ver...
Environment for the game Wiki Race
Classic Infocom interactive fiction games (Zork, Enchanter, etc.) for evaluating LLM reasoning, planning, and world modeling
Benchmarking model performance on SWE Bench in the Mini SWE Agent harness.
Evaluates sycophantic behavior in LLMs across four tasks from Sharma et al. (ICLR 2024).
A realistic virtual EHR environment to benchmark medical LLM agents on clinical tasks.
AndroidWorld benchmark for evaluating autonomous agents on real Android apps with 116 tasks across 20 apps
Test model's ability to correctly click on target UI
Multi-turn environment for testing coding abilities across multiple programming languages using Exercism exercises
Benchmark for agent robustness against prompt injection attacks in tool-use scenarios
AidanBench multi-turn environment for Verifiers
Verifiers environment for BrowseComp-Plus Deep-Research Agent Benchmark. Controlled agent/retriever evaluation on the fixed human-verified corpus.
Future House Aviary wrapper for verifiers - Scientific reasoning environments with tools
Word puzzle game where players find groups of 4 words sharing a common theme
Multi-turn Text-to-SQL environment with interactive database feedback following SkyRL-SQL methodology
τ-bench: Tool-Agent-User benchmark for conversational agents in customer service domains with user simulation
Codebase search environment for Triton GPU programming library - tests agent's ability to navigate and answer questions about the Triton codebase u...
Vision-SR1 environment (train+eval) using original graders