Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published 3 days ago • 272
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Paper • 2608.02580 • Published 11 days ago • 24
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 143
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 30 days ago • 212
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation Paper • 2607.03819 • Published Jul 4 • 10
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 107
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 274
Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper • 2605.09998 • Published May 11 • 19
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 238
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning Paper • 2605.03314 • Published May 6 • 4
Leveraging Verifier-Based Reinforcement Learning in Image Editing Paper • 2604.27505 • Published Apr 30 • 59
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Paper • 2604.19667 • Published Apr 21 • 23
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Paper • 2604.06628 • Published Apr 8 • 330
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver Paper • 2604.08377 • Published Apr 9 • 294
Locally Confident, Globally Stuck: The Quality-Exploration Dilemma in Diffusion Language Models Paper • 2604.00375 • Published Apr 1 • 6