-
Efficient RL Training for LLMs with Experience Replay
Paper • 2604.08706 • Published • 23 -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Paper • 2605.30789 • Published • 26 -
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Paper • 2607.18722 • Published • 35
Chintu Kumar
chang2394
AI & ML interests
None yet
Recent Activity
updated a collection 1 day ago
Off policy/entropy updated a collection about 1 month ago
Off policy/entropy updated a collection 7 months ago
Tool useOrganizations
None yet