STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Paper • 2609.38169 • Published 11 days ago • 117
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? Paper • 2610.07557 • Published 4 days ago • 57
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments? Paper • 2610.08215 • Published 2 days ago • 110
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published about 1 month ago • 656
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published Aug 27 • 57
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published Aug 26 • 151