view article Article Assisted Generation: a new direction toward low-latency text generation joaogante • May 11, 2023 • 82
Gated Recurrent Transformer (GRT) Collection Pre-trained checkpoints for the Gated Recurrent Transformer (GRT). arXiv:2608.15062 • 5 items • Updated Sep 1 • 2
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published Sep 3 • 115
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published Sep 4 • 26
Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs Paper • 2507.07996 • Published Jul 10, 2025 • 36
view article Article Vision Language Models (Better, faster, stronger) +3 merve, sergiopaniego, ariG23498, pcuenq, andito • May 12, 2025 • 619