view article Article MiniMax Goes Sparse: Decoding M3's Attention from a Single Diagram AtlasCloud-AI • May 29 • 11
IntelligenceLab/Long-Horizon-Terminal-Bench Viewer • Updated about 6 hours ago • 46 • 3.51k • 120
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 19 days ago • 76
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 10 days ago • 137
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 46
WARP: Weight-Space Analysis for Recovering Training Data Portfolios Paper • 2607.01686 • Published 26 days ago • 10
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 6 days ago • 31
AutoIndex: Learning Representation Programs for Retrieval Paper • 2607.18603 • Published 7 days ago • 10
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction Paper • 2605.11354 • Published May 12 • 2