InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 9
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 9
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning Paper • 2504.04524 • Published Apr 6, 2025 • 1
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 9
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Paper • 2603.25040 • Published Mar 26 • 132
Xuerui2312/DeepSeek-R1-Distill-Qwen-7B-TRPA-DeepScaleR-verl0326 Text Generation • 8B • Updated Jun 20, 2025 • 26 • 1
Xuerui2312/DeepSeek-R1-Distill-Qwen-7B-TRPA-DeepScaleR-verl0326 Text Generation • 8B • Updated Jun 20, 2025 • 26 • 1
Xuerui2312/DeepSeek-R1-Distill-Qwen-7B-Rollout64-32k-AIME2024-AIME2025-GPQA Updated Jun 11, 2025 • 14