FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience Paper • 2609.03241 • Published Sep 3 • 54
UI-Mate Collection Open-weight CUA models and office-centric CUA benchmark • 4 items • Updated 18 days ago • 21
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations Paper • 2608.15930 • Published Aug 16 • 47
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published Jul 30 • 75
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Paper • 2607.18722 • Published Jul 21 • 36
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 126
On Data Engineering for Scaling LLM Terminal Capabilities Paper • 2602.21193 • Published Feb 24 • 104
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Execution-Free Reward Models Paper • 2602.17684 • Published Feb 4 • 24
LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition Paper • 2307.13269 • Published Jul 25, 2023 • 34
Language Models Can Learn from Verbal Feedback Without Scalar Rewards Paper • 2509.22638 • Published Sep 26, 2025 • 70
cwm Collection Collection for Code World Model, an agentic coding model from FAIR. • 3 items • Updated Sep 24, 2025 • 20
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Paper • 2506.20512 • Published Jun 25, 2025 • 49
Optimizing Anytime Reasoning via Budget Relative Policy Optimization Paper • 2505.13438 • Published May 19, 2025 • 36
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis Paper • 2505.13227 • Published May 19, 2025 • 46