The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 3 days ago • 271
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 7 days ago • 154
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 15 days ago • 138
What Does Privileged Information Add to On-Policy Self-Distillation? Paper • 2609.20612 • Published 15 days ago • 36
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 22 days ago • 706
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 18 days ago • 33
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents Paper • 2609.01281 • Published about 1 month ago • 15
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation Paper • 2608.05879 • Published Aug 29 • 9
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience Paper • 2609.03241 • Published 29 days ago • 54
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published 29 days ago • 71
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published Aug 31 • 42
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 29 days ago • 143
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published about 1 month ago • 222
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 29 days ago • 76
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models Paper • 2609.02886 • Published 30 days ago • 118