Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 3 days ago • 1
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 3 days ago • 1