LMArena
LMArena is a crowdsourced AI benchmark platform run by the LMSYS team. Models are compared by human preference votes; ratings use the Bradley-Terry system and are reported as Elo-like scores with 95% CIs.
- Type
- research-lab
- Website
- lmarena.ai
- GitHub
- github.com/lmarena
Cite
Notes
Only stored in your browser.
Evals
0
Tools
0
Models
0
Papers
0
Boards
5
People
0
Leaderboards
5LMArena - Search FactualityCrowdsourced search factuality model ratings from Arena. Elo-style scores computed from pairwise human preference votes.LMArena - Text FactualityCrowdsourced text factuality model ratings from Arena. Elo-style scores computed from pairwise human preference votes.LMArena - Creative WritingHuman-preference Elo on creative-writing prompts. The only large board where writing quality is judged by people rather than an LLM judge.LMArena - Writing and LiteratureHuman-preference Elo on prompts from writing, literature and language work - the closest public proxy for article and long-form copy quality.LMArena - Instruction FollowingHuman-preference Elo on prompts that carry explicit constraints - does the model actually do what the brief asked.