ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Paper • 2609.09076 • Published 16 days ago • 24
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 16 days ago • 82
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 28 days ago • 155
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Paper • 2608.20281 • Published Aug 20 • 14
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 152
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models Paper • 2608.10538 • Published Aug 11 • 17
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Paper • 2608.08722 • Published Aug 9 • 8
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Paper • 2608.06060 • Published Aug 6 • 41
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published Aug 6 • 23