Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block Quantization Paper • 2403.12422 • Published Mar 19, 2024 • 2
TokenRouter: Efficient Serving System for Token-Level LLM Routing Paper • 2610.12242 • Published 1 day ago • 86
Gemstone Models Collection Our 22 open source Gemstone models for scaling laws range from 50M to 2B parameters, spanning 11 widths from 256 to 3072 and 18 depths from 3 to 80. • 66 items • Updated 4 days ago • 11
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention Paper • 2509.24006 • Published Sep 28, 2025 • 119