Researchers unveil Sparse Feature Attention, a technique that slashes Transformer self‑attention costs by representing queries and keys as k‑sparse codes. By shifting sparsity from sequence length to feature dimension, the method preserves expressivity while cutting complexity from Θ(n²d) to Θ(n²k²/d).
The Signal
The FlashSFA kernel, an extension of FlashAttention, enables efficient large‑scale deployment.