From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 8 days ago • 13
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 10 days ago • 28
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 8 days ago • 144
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 8 days ago • 32
Srishti280992/repro-exact-unlearning-in-reinforcement-learning-traces Traces • Updated Aug 3 • 76 • 3
JonusNattapong/Reinforcement-Learning-for-Gold-Trading-Model Reinforcement Learning • Updated Dec 23, 2025 • 103 • 14
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 9 days ago • 42
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 9 days ago • 41
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 9 days ago • 56
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 9 days ago • 52