We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the W_\infty contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular M-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes α_t=c(t+1)^{-a} with a\in(1/2,1), the leading last-iterate fluctuation is of order \widetilde O\bigl(T^{-a/2}/\sqrt{1-γ}\bigr) and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order m^{-1} in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.
A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms.
- Preview

- Year
- 2026
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2608.27313ARXIV-DEFAULT
- TL;DR
- Semantic Scholar