0

TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes

In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2608.13057CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below n^* \approx 156--168 tokens, HBM weight streaming dominates---cost attaches to activated replicas, not tokens; above it, grouped GEMM rounds tokens to 128-tile M-tiles, so splitting an expert adds padded compute. A max-affine profile t=\max(a+bG,,c+βN) captures both regimes. Realistic decode batches hold hot experts in the linear regime and cold in the flat simultaneously; recorded batches show proxy dispatches differ by 1.4--1.6\times in modeled block time (p95 up to 1.7\times), and which proxy wins flips with the regime. We formalize per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and present TEMPO, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count collection into one in-graph kernel. Anchored by an 8-GPU Testbed A microbenchmark, TEMPO stays within 1% of the best fixed baseline everywhere and wins by up to 15.5% where regimes mix. End-to-end on Testbed B, Qwen3-235B (inside the win region) gains 4--6% throughput and cuts p99 latency by \sim 15.6%; DeepSeek-V3 (outside, communication-dominated) shows only mechanism cost. A phase diagram, not a universal win, is the claim: it predicts both outcomes before deployment.