We study repeated contextual procurement auctions in which producers have private costs and the platform must learn context-dependent product values from bandit feedback. The objective is welfare rather than revenue or a virtual-cost surrogate: regret is the total surplus loss relative to the full-information efficient procurement rule. We first show that the natural UCB allocation rule attains \tilde O(\sqrt{ngT}) welfare regret under truthful bids, but its adaptive bid-dependent learning path does not by itself give a truthfulness guarantee. To obtain exact incentives, we design a bid-independent explore-then-commit mechanism with empirical critical payments; it is dominant-strategy truthful and has \tilde O((ng)^{1/3}T^{2/3}) regret. We then introduce frozen-payment UCB, which estimates payments in an initial bid-independent exploration phase, freezes those payment estimates, and continues adaptive UCB allocation learning afterwards. Under a smoothed truthful-path margin condition, this mechanism gives a regret-incentive tradeoff: the near-UCB tuning attains \tilde O(\sqrt{ngT}) welfare regret, while the average per-round gain from any fixed deviation is at most \tilde O(T^{-1/4}) for fixed n,g. A matching lower bound shows that this frozen-payment frontier is unavoidable.
Contextual Procurement Auctions with Bandit Learning
We study repeated contextual procurement auctions in which producers have private costs and the platform must learn context-dependent product values from bandit feedback.
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2607.05813CC-BY-4.0
- TL;DR
- Semantic Scholar