MassAlloc Attention: Let Attention Allocate Its Own Compute
Abstract
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive retention of the work. MALA reduces low-contribution post-score computation. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens with tensor parallelism, MALA reduces forward and backward latency during training by 2.2x and 3.0x and decoding latency during inference by 1.6x relative to FullAttn. Across scaling-law training from 0.6B to 14B parameters, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.
Community
Sharing two recent explorations in attention design from our team.
We started with two straightforward questions: Does every attention head need to repeatedly attend to the entire causal history? Once attention scores have been computed, do regions with very little contribution still need the full subsequent computation?
We explored two approaches:
CoWindow Attention (CoWA): Let heads share the work of accessing history. Heads share local context and divide distant context into complementary windows. Each head attends sparsely, while the heads collectively cover the full causal history.
📄 arxiv.org/abs/2609.32704
MassAlloc Attention (MALA): Let attention allocate its own compute. MALA preserves full causal QK scoring, then uses attention’s own softmax statistics to reduce subsequent computation in low-contribution regions.
📄 arxiv.org/abs/2609.32712
Both approaches support training forward and backward passes, as well as inference prefill and decoding. In attention-operator benchmarks at 128K tokens on 8×H100 with TP=8, compared with FullAttn:
- CoWA: 7.4× forward, 8.6× backward, and 3.0× decoding speedups.
- MALA: 2.2× forward, 3.0× backward, and 1.6× decoding speedups.
We also conducted scaling experiments from 0.6B to 14B, alongside separate continued-training experiments at 32B. During 14B training with 32K context, CoWA and MALA reduced total training FLOPs by 28.5% and 23.1%, respectively, while maintaining performance comparable to FullAttn on the evaluated model capabilities.
From method design to kernel implementation to model training, our goal was to explore which attention computations can be eliminated—and how to turn those savings into practical gains in ML infrastructure.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CEDAR: Error-Bounded Residual Routing for Efficient Long-Context Attention (2026)
- TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2026)
- LoGo: Token-Level Dynamic Local-Global Attention (2026)
- On-Demand Attention: Language Models Know When to Recall (2026)
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models (2026)
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving (2026)
- CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.32712 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper