Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning Paper • 2504.04524 • Published Apr 6, 2025 • 1
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 9
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Paper • 2603.25040 • Published Mar 26 • 132