Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior Paper • 2609.39827 • Published 5 days ago • 11
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning Paper • 2609.36820 • Published 6 days ago • 24
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 5 days ago • 50
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models Paper • 2609.32264 • Published 9 days ago • 45
DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation Paper • 2609.33711 • Published 8 days ago • 12
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation Paper • 2609.39687 • Published 5 days ago • 15
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 6 days ago • 465
Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation Paper • 2610.02148 • Published 4 days ago • 19
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics Paper • 2609.35259 • Published 7 days ago • 167
Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation Paper • 2609.38660 • Published 6 days ago • 35
Persona Dosing: Calibrated Activation Steering for Graded Trait Control Paper • 2609.36388 • Published 7 days ago • 40
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It Paper • 2609.36585 • Published 6 days ago • 63
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 4 days ago • 49
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization Paper • 2609.32993 • Published 9 days ago • 42
Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning Paper • 2609.37915 • Published 6 days ago • 7
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning Paper • 2610.02181 • Published 4 days ago • 14
Smaller Models, Better Rejects: Preference Distillation Scaling Paper • 2609.38987 • Published 5 days ago • 15
Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens Paper • 2610.01939 • Published 4 days ago • 34
RPTune: Learned Context Curation for LLM Catalog Search Paper • 2610.00964 • Published 4 days ago • 20