What Does Privileged Information Add to On-Policy Self-Distillation? Paper • 2609.20612 • Published 19 days ago • 36
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 20 days ago • 84
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 29 days ago • 376
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 27 days ago • 45
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published Aug 27 • 155
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement Paper • 2608.11951 • Published Aug 12 • 9
Mitigating Gender Bias in English to Romanian Machine Translation Paper • 2608.08606 • Published Aug 9 • 9
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Paper • 2608.10692 • Published Aug 11 • 13
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Paper • 2608.01481 • Published Aug 2 • 75
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Paper • 2608.02392 • Published Aug 3 • 16
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published Aug 3 • 161
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 97