HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
5 Oct 2026
Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult.