Smaller Models, Better Rejects: Preference Distillation Scaling Paper • 2609.38987 • Published 8 days ago • 28
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning Paper • 2609.36820 • Published 9 days ago • 39
DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation Paper • 2609.33711 • Published 11 days ago • 19
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Paper • 2609.35954 • Published 10 days ago • 50
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 12 days ago • 324
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents Paper • 2609.33642 • Published 11 days ago • 32
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers Paper • 2609.32814 • Published 12 days ago • 22
Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models Paper • 2609.26637 • Published 16 days ago • 26
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 15 days ago • 56
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World Paper • 2609.23038 • Published 19 days ago • 63
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization Paper • 2609.15991 • Published 20 days ago • 17
MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads Paper • 2609.09206 • Published Sep 5 • 11
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 21 days ago • 139