Cite
Notes
Only stored in your browser.
Attribution
RL/eval environment for ghostwriting in a client's voice, scored by deterministic rule checks plus an LLM judge.
Multi-turn contract negotiation environment with verifiable reward function
Run OSU-NLP QUEST's released RL tasks + rubric-tree rewards without their withheld cached corpus (live-fetch reproducibility bridge)
Mind2Web web-action prediction with swappable binary vs shaped GRPO reward
Multi-term contract negotiation with a verifiable, Pareto-frontier reward (no judge)