Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent Tokens Paper • 2602.10229 • Published Feb 10 • 5
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering Paper • 2605.29648 • Published May 28 • 8
SPA: Towards A Computational Friendly Cloud-Base and On-Devices Collaboration Seq2seq Personalized Generation Paper • 2403.07088 • Published Mar 11, 2024 • 1
ERA-CoT: Improving Chain-of-Thought through Entity Relationship Analysis Paper • 2403.06932 • Published Mar 11, 2024 • 1
RA-ISF: Learning to Answer and Understand from Retrieval Augmentation via Iterative Self-Feedback Paper • 2403.06840 • Published Mar 11, 2024 • 1
MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models Paper • 2402.12851 • Published Feb 20, 2024 • 2
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? Paper • 2608.21833 • Published Aug 22 • 18
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published Aug 18 • 138
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published Aug 3 • 25
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? Paper • 2606.17861 • Published Jun 16 • 60
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? Paper • 2606.17861 • Published Jun 16 • 60
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks Paper • 2502.14848 • Published Feb 20, 2025 • 1
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 43
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning Paper • 2604.16029 • Published Apr 17 • 22
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper • 2512.04324 • Published Dec 3, 2025 • 160
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning Paper • 2506.01710 • Published Jun 2, 2025 • 3
S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models Paper • 2310.15147 • Published Oct 23, 2023 • 2