Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 4 days ago • 15
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 12 days ago • 138
NeoHorse-1 Collection NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness • 3 items • Updated 20 days ago • 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 22 days ago • 327
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 30 days ago • 63
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Image-Text-to-Text • 305B • Updated 28 days ago • 981k • • 930
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities Paper • 2608.28122 • Published Aug 28 • 66