We study online learning for a seller that jointly chooses per-period inventory positions and a uniform price, then fulfills realized demand through a downstream allocation. The main difficulty is not only demand learning: the price shifts demand and reshapes the transportation LP, making the population objective globally non-convex and non-smooth. To solve this problem, we propose OCSAA, an algorithm that exploits demand observations through counterfactual translation and proposes joint (price, inventory) decisions through lower-confidence optimism. OCSAA admits a polynomial-time additive-accuracy implementation for rational-polytope inventory sets. We prove a high-probability \widetilde O(\sqrt T) regret guarantee and establish a matching-in-T information-theoretic lower bound. Our results illustrate an effective integration of statistical learning methodologies with complex operations research problems.
Online Pricing and Allocation with Demand Learning and Fulfillment Cost
We study online learning for a seller that jointly chooses per-period inventory positions and a uniform price, then fulfills realized demand through a downstream allocation.
- Year
- 2025
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2501.18049ARXIV-DEFAULT
- TL;DR
- Semantic Scholar