FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
20 Aug 2026
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase.
200.1/h