view article Article Efficient LLM Pretraining: Packed Sequences and Masked Attention sirluk • Oct 7, 2024 • 74
view article Article Re-understanding KL Approximation from an RL-for-LLM Lens: Notes on “Approximating KL Divergence” NormalUhr • Aug 11, 2025 • 14
view article Article There is no such thing as a tokenizer-free lunch catherinearnett • Sep 25, 2025 • 102