0

The Sample Complexity of Policy Learning with Mu-Resets

We study policy-based reinforcement learning under the $μ$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $μ$, in addition to the starting distribution.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2608.07772CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

We study policy-based reinforcement learning under the μ-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution μ, in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon H is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a \exp(Ω(H)) sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as \exp(Θ(\sqrt H)).