PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection Paper • 2607.04690 • Published 24 days ago • 3
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
timaeus/rl-lm-pythia70m-formality-pos-beta0-grpo-nostd-gs4-tp1-tk0-pt80000-steerDotL6c64s8-seed20 Updated 26 days ago • 1 • 1
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents Paper • 2605.25624 • Published May 25 • 35