James4Ever0/computer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcp Updated Aug 10, 2025 • 108 • 3
Marathoner: Ultra-Long-Horizon Autonomous Intelligence Paper • 2609.34378 • Published 5 days ago • 31
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 4 days ago • 122
JonusNattapong/Reinforcement-Learning-for-Gold-Trading-Model Reinforcement Learning • Updated Dec 23, 2025 • 96 • 20
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 10 days ago • 13
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models Paper • 2609.04355 • Published 15 days ago • 9
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 9 days ago • 13
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 14 days ago • 41
Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand Paper • 2609.17172 • Published 18 days ago • 5
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 14 days ago • 17
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents Paper • 2609.18779 • Published 17 days ago • 16