MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following Paper • 2605.03858 • Published May 5 • 1
Towards Understanding Multimodal Fine-Tuning: Spatial Features Paper • 2602.08713 • Published Feb 6 • 1
EVCL: Elastic Variational Continual Learning with Weight Consolidation Paper • 2406.15972 • Published Jun 23, 2024 • 1
Reasoning Fine-Tuning Induces Persistent Latent Policy States Paper • 2607.18532 • Published 18 days ago • 1
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published 9 days ago • 6
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks Paper • 2605.01417 • Published May 2 • 2
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought Paper • 2403.05518 • Published Mar 8, 2024 • 3
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published 9 days ago • 6
SpatialThinker Collection This collection consists of SpatialThinker 3B and 7B model checkpoints, and STVQA-7K, a Spatial VQA dataset used for training the models. • 5 items • Updated Jul 2 • 1