Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published Jul 8 • 88
timaeus/rl-lm-pythia70m-formality-neg-beta0-grpo-nostd-gs4-tp1-tk0-pt80000-steerSRL12c8000sm0.004-seed1 Updated Jul 8 • 1
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Convex Low-resource Accent-Robust Language Detection in Speech Recognition Paper • 2605.23235 • Published May 22 • 6
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433