Cite
Notes
Only stored in your browser.
Attribution
Reward Hacking Sprint env: routing/calibration-as-action. A cheap base model (Llama-3.2-1B) answers or ESCALATES to a simulated frontier oracle; re...
Reward-hacking env: persuasion / oversight-evasion. A cheap student answers gold-verifiable but subjective-seeming MCQs (TruthfulQA); a frozen LLM ...
Trains an RL router (Nemotron 3 Ultra / Claude Opus 4.7 / GPT-5.5) offline against a precomputed oracle table built from real, harness-verified SWE...