MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following Paper • 2605.03858 • Published May 5 • 1
Towards Understanding Multimodal Fine-Tuning: Spatial Features Paper • 2602.08713 • Published Feb 6 • 1
EVCL: Elastic Variational Continual Learning with Weight Consolidation Paper • 2406.15972 • Published Jun 23, 2024 • 1
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought Paper • 2403.05518 • Published Mar 8, 2024 • 3
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks Paper • 2605.01417 • Published May 2 • 2
Reasoning Fine-Tuning Induces Persistent Latent Policy States Paper • 2607.18532 • Published 18 days ago • 1
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published 9 days ago • 6
Constitutional Midtraining Collection Data and models for the paper 'Constitutional Midtraining: Content Presence Drives Alignment Gains' . • 17 items • Updated 10 days ago • 3
Lie Detection Collection Did you lie? Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms • 9 items • Updated Jun 10 • 6
view article Article NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI nvidia • Jan 5 • 64
Physical AI Collection Collection of open, commercial-grade datasets for physical AI developers • 57 items • Updated 4 days ago • 178
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards Paper • 2511.07403 • Published Nov 10, 2025 • 14
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
view article Article Smol2Operator: Post-Training GUI Agents for Computer Use +3 A-Mahla, merve, sergiopaniego, reach-vb, lewtun • Sep 23, 2025 • 138
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning Paper • 2504.16656 • Published Apr 23, 2025 • 59