We study policy-based reinforcement learning under the μ-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution μ, in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon H is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a \exp(Ω(H)) sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as \exp(Θ(\sqrt H)).
The Sample Complexity of Policy Learning with Mu-Resets
We study policy-based reinforcement learning under the $μ$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $μ$, in addition to the starting distribution.
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2608.07772CC-BY-4.0
- TL;DR
- Semantic Scholar