IFBench
Frontier
Instruction-following benchmark measuring adherence to multi-step constraints.
- Publisher
- Allen Institute for AI
- Published
- Jun 2025
- Canonical
- github.com/allenai/IFBench
Cite
Notes
Only stored in your browser.
Top score 82.9% by MiniMax M3 (batch) - 326 models reporting (68 frontier)
Score history
326Top models
326Related tools
3Implementations, trainers, datasets and scaffolds linked to this eval.
FAQ
- What is IFBench?
- Instruction-following benchmark measuring adherence to multi-step constraints.
- What is the current top score on IFBench?
- The top reported score is 82.9% by MiniMax M3 (batch), across 326 models reporting (68 from frontier labs).
- How can a model improve its IFBench score?
- Tools linked to IFBench on Sophon include Ifbench RL Env (Community), Ifbench RL Env (Prime Intellect), Ifbench RL Env (Community) - RL environments, datasets, and scaffolds that target this eval.
