Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published Sep 3 • 114
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published Sep 4 • 26
Gated Recurrent Transformer (GRT) Collection Pre-trained checkpoints for the Gated Recurrent Transformer (GRT). arXiv:2608.15062 • 5 items • Updated Sep 1 • 2
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published Aug 26 • 12
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published Aug 26 • 12 • 5
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published Aug 26 • 12 • 5
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published Aug 26 • 12