Doge family of small language models.
Loser Cheems
JingzeShi
AI & ML interests
I like training small languge models.
Recent Activity
posted an update 1 day ago
Sharing two recent explorations in attention design from our team.
We started with two straightforward questions: Does every attention head need to repeatedly attend to the entire causal history? Once attention scores have been computed, do regions with very little contribution still need the full subsequent computation?
We explored two approaches:
CoWindow Attention (CoWA): Let heads share the work of accessing history. Heads share local context and divide distant context into complementary windows. Each head attends sparsely, while the heads collectively cover the full causal history.
https://huggingface.co/papers/2609.32704
MassAlloc Attention (MALA): Let attention allocate its own compute. MALA preserves full causal QK scoring, then uses attention’s own softmax statistics to reduce subsequent computation in low-contribution regions.
https://huggingface.co/papers/2609.32712
Both approaches support training forward and backward passes, as well as inference prefill and decoding. In attention-operator benchmarks at 128K tokens on 8×H100 with TP=8, compared with FullAttn:
CoWA: 7.4× forward, 8.6× backward, and 3.0× decoding speedups.
MALA: 2.2× forward, 3.0× backward, and 1.6× decoding speedups.
We also conducted scaling experiments from 0.6B to 14B, alongside separate continued-training experiments at 32B. During 14B training with 32K context, CoWA and MALA reduced total training FLOPs by 28.5% and 23.1%, respectively, while maintaining performance comparable to FullAttn on the evaluated model capabilities.
From method design to kernel implementation to model training, our goal was to explore which attention computations can be eliminated, and how to turn those savings into practical gains in ML infrastructure. upvoted a paper 1 day ago
CoWindow Attention: Full Causal Coverage Is a Collective Property upvoted a paper 1 day ago
MassAlloc Attention: Let Attention Allocate Its Own Compute