VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 5 days ago • 48
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Paper • 2609.39071 • Published 6 days ago • 50
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models Paper • 2610.03665 • Published 4 days ago • 51
Native Action-Prior Learning from Videos for World Action Models Paper • 2610.03391 • Published 4 days ago • 77
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 5 days ago • 258
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 7 days ago • 101
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 7 days ago • 136
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 10 days ago • 323