StudentSim: Training LLM-based Student Simulators Paper • 2609.01591 • Published about 1 month ago • 494
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 8 days ago • 23
Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models Paper • 2609.26637 • Published 10 days ago • 23
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation Paper • 2609.24981 • Published 11 days ago • 71
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 11 days ago • 55
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 15 days ago • 43
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 16 days ago • 67
Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid? Paper • 2609.14803 • Published 19 days ago • 12