When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 16 days ago • 110
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 26 days ago • 376
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Paper • 2609.09076 • Published 25 days ago • 24
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 25 days ago • 84
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 30 days ago • 85
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Paper • 2608.08975 • Published Aug 10 • 48
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Paper • 2608.10915 • Published Aug 11 • 195