Post
37
Sharing two recent explorations in attention design from our team.
We started with two straightforward questions: Does every attention head need to repeatedly attend to the entire causal history? Once attention scores have been computed, do regions with very little contribution still need the full subsequent computation?
We explored two approaches:
CoWindow Attention (CoWA): Let heads share the work of accessing history. Heads share local context and divide distant context into complementary windows. Each head attends sparsely, while the heads collectively cover the full causal history.
CoWindow Attention: Full Causal Coverage Is a Collective Property (2609.32704)
MassAlloc Attention (MALA): Let attention allocate its own compute. MALA preserves full causal QK scoring, then uses attentionβs own softmax statistics to reduce subsequent computation in low-contribution regions.
MassAlloc Attention: Let Attention Allocate Its Own Compute (2609.32712)
Both approaches support training forward and backward passes, as well as inference prefill and decoding. In attention-operator benchmarks at 128K tokens on 8ΓH100 with TP=8, compared with FullAttn:
CoWA: 7.4Γ forward, 8.6Γ backward, and 3.0Γ decoding speedups.
MALA: 2.2Γ forward, 3.0Γ backward, and 1.6Γ decoding speedups.
We also conducted scaling experiments from 0.6B to 14B, alongside separate continued-training experiments at 32B. During 14B training with 32K context, CoWA and MALA reduced total training FLOPs by 28.5% and 23.1%, respectively, while maintaining performance comparable to FullAttn on the evaluated model capabilities.
From method design to kernel implementation to model training, our goal was to explore which attention computations can be eliminated, and how to turn those savings into practical gains in ML infrastructure.
We started with two straightforward questions: Does every attention head need to repeatedly attend to the entire causal history? Once attention scores have been computed, do regions with very little contribution still need the full subsequent computation?
We explored two approaches:
CoWindow Attention (CoWA): Let heads share the work of accessing history. Heads share local context and divide distant context into complementary windows. Each head attends sparsely, while the heads collectively cover the full causal history.
CoWindow Attention: Full Causal Coverage Is a Collective Property (2609.32704)
MassAlloc Attention (MALA): Let attention allocate its own compute. MALA preserves full causal QK scoring, then uses attentionβs own softmax statistics to reduce subsequent computation in low-contribution regions.
MassAlloc Attention: Let Attention Allocate Its Own Compute (2609.32712)
Both approaches support training forward and backward passes, as well as inference prefill and decoding. In attention-operator benchmarks at 128K tokens on 8ΓH100 with TP=8, compared with FullAttn:
CoWA: 7.4Γ forward, 8.6Γ backward, and 3.0Γ decoding speedups.
MALA: 2.2Γ forward, 3.0Γ backward, and 1.6Γ decoding speedups.
We also conducted scaling experiments from 0.6B to 14B, alongside separate continued-training experiments at 32B. During 14B training with 32K context, CoWA and MALA reduced total training FLOPs by 28.5% and 23.1%, respectively, while maintaining performance comparable to FullAttn on the evaluated model capabilities.
From method design to kernel implementation to model training, our goal was to explore which attention computations can be eliminated, and how to turn those savings into practical gains in ML infrastructure.