When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 18 days ago • 110
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 27 days ago • 327
Steering Geometry: Validating Human Value Geometry in LLM Steering Space Paper • 2609.06289 • Published 30 days ago • 31
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published Aug 27 • 155
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Paper • 2607.27637 • Published Aug 1 • 7
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Paper • 2608.05369 • Published Aug 5 • 27
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Paper • 2608.03812 • Published Aug 4 • 29
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published Aug 3 • 161
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published Jul 30 • 37