WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published about 1 month ago • 139
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 16 days ago • 172
ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation Paper • 2609.00968 • Published 23 days ago • 5
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 110
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection Paper • 2608.06865 • Published Aug 7 • 11
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 178
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published Jul 22 • 111