New Block Sparse Attention Achieves Log-Linear Complexity
| Source: HF Papers | Original article
Researchers propose a block sparse attention method that reduces self‑attention complexity from quadratic to log‑linear, addressing the bottleneck of block selection in long‑context language models.
A new paper titled **“Block Sparse Attention with Log‑Linear Complexity”** introduces **PISA (Pyramid Sparse Attention)**, a trainable block‑sparse attention mechanism that cuts the cost of selecting attention blocks from quadratic to log‑linear time.
Current large‑language models struggle with long contexts because the standard self‑attention matrix grows quadratically with sequence length. Block‑sparse attention mitigates this by limiting computation to a subset of blocks, but the selection step itself still requires scoring every query‑block pair, preserving the quadratic bottleneck. PISA replaces that step with a hierarchical, pyramid‑shaped process that identifies the most relevant blocks in **O(N log N)** time while remaining trainable.
The advance matters because it restores the theoretical efficiency gains of block‑sparse attention for both pre‑fill and generation phases. Existing trainable sparse methods such as NSA, MoBA and HiLS still incur quadratic costs during pre‑fill, and the training‑free HiP approach, though log‑linear, cannot be optimized end‑to‑end. By delivering a trainable, log‑linear selection stage, PISA promises to enable language models to handle much longer inputs without prohibitive memory or compute demands, opening the door to richer document‑level reasoning, more detailed code analysis, and extended conversational memory.
The paper joins a growing suite of efficient‑attention research, including the “Log‑linear Sparse Attention” (LLSA) framework that also leverages hierarchical structures. The next steps will be empirical: benchmarking PISA against LLSA, HiP and existing block‑sparse kernels, and integrating the method into open‑source libraries such as the MIT HAN Lab’s Block‑Sparse‑Attention repo. If PISA delivers the claimed speed‑up without sacrificing accuracy, model developers and cloud providers are likely to adopt it quickly, reshaping how future AI systems scale to ever‑longer contexts.
Sources
Back to AIPULSEN