-
Efficient RL Training for LLMs with Experience Replay
Paper • 2604.08706 • Published • 23 -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Paper • 2605.30789 • Published • 26 -
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Paper • 2607.18722 • Published • 35
Chintu Kumar
chang2394
AI & ML interests
None yet
Recent Activity
updated a collection 2 days ago
Off policy/entropy updated a collection about 1 month ago
Off policy/entropy updated a collection 7 months ago
Tool useOrganizations
None yet
Inference improvements
Attention
LLM memory
RL Advantage
Eval
Tool use
-
Provable Benefits of In-Tool Learning for Large Language Models
Paper • 2508.20755 • Published • 11 -
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Paper • 2508.20453 • Published • 63 -
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
Paper • 2508.20931 • Published • 16 -
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Paper • 2509.08755 • Published • 56
Off policy/entropy
-
Efficient RL Training for LLMs with Experience Replay
Paper • 2604.08706 • Published • 23 -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Paper • 2605.30789 • Published • 26 -
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Paper • 2607.18722 • Published • 35
RL Advantage
Inference improvements
Eval
Attention
Tool use
-
Provable Benefits of In-Tool Learning for Large Language Models
Paper • 2508.20755 • Published • 11 -
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Paper • 2508.20453 • Published • 63 -
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
Paper • 2508.20931 • Published • 16 -
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Paper • 2509.08755 • Published • 56
LLM memory