Cite
Notes
Only stored in your browser.
Attribution
Toy single-turn RL environment: maximize the letter 'a', stay short.
RL env: train a judge to spot planted writing violations; reward = deterministic span-F1 vs labels-by-construction (no LLM in the loop)
RL env: rewrite messy drafts into plain English; reward = within-group clarity rank x content-preservation gate (LLM judges)