Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models Paper • 2610.03665 • Published 4 days ago • 22
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models Paper • 2610.03665 • Published 4 days ago • 22
HuRo: Robotizing Human Videos for Scalable VLA Pretraining Paper • 2609.10706 • Published 18 days ago • 29
CoTEVer: Chain of Thought Prompting Annotation Toolkit for Explanation Verification Paper • 2303.03628 • Published Mar 7, 2023 • 2
Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization Paper • 2209.00930 • Published Sep 2, 2022 • 2
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models Paper • 2406.05761 • Published Jun 9, 2024 • 3
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up Paper • 2609.13443 • Published 25 days ago • 13
Meta-Reinforcement Learning with Self-Reflection for Agentic Search Paper • 2603.11327 • Published Mar 11 • 11
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains Paper • 2507.06187 • Published Jul 8, 2025
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists Paper • 2605.20668 • Published May 20 • 14
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization Paper • 2605.26457 • Published May 26 • 6
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts Paper • 2606.02404 • Published Jun 1 • 58
A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls Paper • 2412.01340 • Published Dec 2, 2024
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists Paper • 2605.20668 • Published May 20 • 14
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts Paper • 2606.02404 • Published Jun 1 • 58
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts Paper • 2606.02404 • Published Jun 1 • 58
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation Paper • 2505.24456 • Published May 30, 2025