Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 6 days ago • 10
ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models Paper • 2609.13231 • Published 23 days ago • 8
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 5 days ago • 43
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 7 days ago • 131
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 21 days ago • 113
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 7 days ago • 32
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 9 days ago • 28
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 7 days ago • 13
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 8 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 8 days ago • 42
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 8 days ago • 56
JonusNattapong/Reinforcement-Learning-for-Gold-Trading-Model Reinforcement Learning • Updated Dec 23, 2025 • 97 • 14
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 9 days ago • 79
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 9 days ago • 100