Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
3 Sep 2026
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be.