Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published 3 days ago • 14
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Paper • 2602.10718 • Published Apr 28
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression Paper • 2601.01204 • Published Jan 3
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published 3 days ago • 14
Massive Activations in Hybrid Linear Attention Models Collection Official PAS, ISP, gating-ablation, and scale-study checkpoints for Massive Activations in Hybrid Linear Attention Models. • 10 items • Updated 1 day ago
startlux-models/gdn-gatedfa-340m-isp-hybrid-3to1-10b Text Generation • 0.4B • Updated 1 day ago • 176
startlux-models/gdn-gatedfa-340m-isp-hybrid-3to1-10b Text Generation • 0.4B • Updated 1 day ago • 176
startlux-models/gdn-nooutgate-340m-isp-hybrid-3to1-10b Text Generation • 0.4B • Updated 1 day ago • 178