MassAlloc Attention: Let Attention Allocate Its Own Compute Paper • 2609.32712 • Published 6 days ago • 74
CoWindow Attention: Full Causal Coverage Is a Collective Property Paper • 2609.32704 • Published 6 days ago • 66
Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models Paper • 2506.11116 • Published Jun 9, 2025 • 5
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Paper • 2506.07463 • Published Jun 9, 2025 • 12
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models Paper • 2410.18505 • Published Oct 24, 2024 • 11
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data Paper • 2410.18558 • Published Oct 24, 2024 • 19
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Paper • 2408.06567 • Published Aug 13, 2024 • 2