Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 10 days ago • 79
HyQuant: Hybrid-Precision Quantization for LLM Attention Paper • 2608.27875 • Published 29 days ago • 30
Memory as Plans: World-Action Modeling with Memory-Grounded Planning Paper • 2609.11561 • Published 16 days ago • 41
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published 18 days ago • 71
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs Paper • 2609.08345 • Published 18 days ago • 24
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance Paper • 2609.02373 • Published 24 days ago • 14
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 30 days ago • 155
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 265
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published Aug 4 • 36
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents Paper • 2608.03509 • Published Aug 4 • 24
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing Paper • 2607.29025 • Published Jul 31 • 18