0

Responsiveness Verification: Will Predictions Change? How Much? How Often?

Machine learning models are often used in applications where their inputs change due to routine interactions, strategic manipulation, or noise. In such settings, models can undermine safety as these changes lead them to predict over regions of input space they have not seen.

Preview
Year
2025
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2507.02169CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Machine learning models are often used in applications where their inputs change due to routine interactions, strategic manipulation, or noise. In such settings, models can undermine safety as these changes lead them to predict over regions of input space they have not seen. We propose to address these challenges by measuring responsiveness---the probability that a model output changes when its inputs change under an interaction model. We develop algorithms to estimate responsiveness for any machine learning model, and a framework to specify broad classes of interaction models. We pair these algorithms with statistical guarantees that support practical validation. We demonstrate how our tools can promote safety and reliability across domains by detecting preclusion in recidivism prediction, estimating the cost of gaming in content moderation, and testing the robustness of benchmarks for LLMs.