Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 7 days ago • 36
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 10 days ago • 17
ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models Paper • 2609.13231 • Published 27 days ago • 18
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 9 days ago • 50
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 11 days ago • 136
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 25 days ago • 119
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 11 days ago • 35
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 13 days ago • 30
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 11 days ago • 13
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 12 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 12 days ago • 44
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 12 days ago • 57
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 13 days ago • 82
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 13 days ago • 101
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 13 days ago • 47
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 19 days ago • 275
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 15 days ago • 214