GigaAM Multilingual: Foundation Model for Underrepresented Languages Paper • 2607.10371 • Published 17 days ago • 33
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors Paper • 2606.32029 • Published 28 days ago • 14
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills? Paper • 2606.01993 • Published Jun 1 • 15
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth Paper • 2605.25052 • Published May 24 • 14
DCAgent3/dev_set_v2_rl__24GPU_base_excl_timeouts__exp_rpt_pymethods2test_large__GLM_4_7_c2148a8d Viewer • Updated May 27 • 296 • 15 • 1
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering Paper • 2605.17526 • Published May 17 • 7