DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 53
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 53
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 9 days ago • 63
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 14 days ago • 137
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Paper • 2605.22012 • Published May 21 • 46
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Paper • 2605.07593 • Published May 8 • 1
Towards Next-Generation LLM Training: From the Data-Centric Perspective Paper • 2603.14712 • Published Mar 16
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Paper • 2604.04707 • Published Apr 6 • 203
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 9 days ago • 63
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 14 days ago • 137
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System Paper • 2407.06027 • Published Jul 8, 2024 • 10
KeyVideoLLM: Towards Large-scale Video Keyframe Selection Paper • 2407.03104 • Published Jul 3, 2024 • 1
Synth-Empathy: Towards High-Quality Synthetic Empathy Data Paper • 2407.21669 • Published Jul 31, 2024
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs Paper • 2408.01122 • Published Aug 2, 2024
MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark Paper • 2408.07543 • Published Aug 14, 2024
SynthVLM: High-Efficiency and High-Quality Synthetic Data for Vision Language Models Paper • 2407.20756 • Published Jul 30, 2024 • 1
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search Paper • 2409.17972 • Published Sep 26, 2024
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction Paper • 2410.21169 • Published Oct 28, 2024 • 30
MC-LLaVA: Multi-Concept Personalized Vision-Language Model Paper • 2411.11706 • Published Nov 18, 2024 • 1