0

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with…

Preview
Year
2026
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2608.12753ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards. Players cannot communicate during learning but may agree on a protocol a priori. For Problems A and B we propose mQ-learning and mQ-learning-intervals, achieving \tilde{O}(\sqrt{H^4 S A_{joint}, T}) regret, where H is the horizon, S the state count, T = KH the total steps, and A_{joint} = \prod_{i=1}^M |A_i| the joint action space across M players. For Problem C we give mEXC and mEXC-Bellman, two-phase explore-then-commit algorithms with regret \tilde{O}(H (S A_{joint})^{1/3} T^{2/3}). Against the centralized joint-action benchmark, decentralized learning under information asymmetry matches the single-agent Q-learning rate of \cite{jin2018q} up to logarithmic factors. Because A_{joint} grows exponentially in M, the bounds are most meaningful for small M or small per-player action sets.