Human interventions are a common source of supervision in autonomous systems during deployment. Many existing approaches are based on avoiding interventions, yet the consequences of this objective are not well understood. We develop a geometric perspective on intervention learning that characterizes intervention avoidance as constraining policies to a face of the occupancy measure polytope. This view reveals that the effectiveness of intervention learning depends on the informativeness of the intervention strategy: highly informative interventions uniquely determine the solution, while weak interventions leave a large set of feasible policies, many of which are suboptimal. Motivated by this under-specification, we define Robust Intervention Learning (RIL) as the problem of learning policies that perform well under varying levels of intervention informativeness. From the geometric formulation, we derive Residual Intervention Fine-Tuning (RIFT), which combines interventions with a prior policy to select among feasible solutions. We show that RIFT provides provable improvement over the prior and corresponds to solving a constrained reinforcement learning problem with an induced reward. Empirically, RIFT yields consistent policy improvement across a range of intervention settings, particularly when interventions are sparse or weakly informative. These results highlight the importance of accounting for intervention informativeness and suggest a principled path toward robust learning from human feedback.
The Geometry of Learning to Avoid Interventions
Human interventions are a common source of supervision in autonomous systems during deployment. Many existing approaches are based on avoiding interventions, yet the consequences of this objective are not well understood.
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-SA-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2602.03825CC-BY-SA-4.0
- TL;DR
- Semantic Scholar