Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision Paper β’ 2602.03103 β’ Published Feb 3
ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning Paper β’ 2512.04555 β’ Published Dec 4, 2025 β’ 1
ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning Paper β’ 2512.04555 β’ Published Dec 4, 2025 β’ 1
Running Agents 287 Infinite Dataset Hub βΎ 287 Search and save datasets generated with a LLM in real time
IndicXTREME Collection IndicXTREME is a human-supervised benchmark of 9 diverse NLU tasks across 20 languages, featuring 105 evaluation sets in total. β’ 8 items β’ Updated Oct 23, 2024 β’ 2
Unveiling the Multi-Annotation Process: Examining the Influence of Annotation Quantity and Instance Difficulty on Model Performance Paper β’ 2310.14572 β’ Published Oct 23, 2023 β’ 1
Unveiling the Multi-Annotation Process: Examining the Influence of Annotation Quantity and Instance Difficulty on Model Performance Paper β’ 2310.14572 β’ Published Oct 23, 2023 β’ 1
Model Hubs and Beyond: Analyzing Model Popularity, Performance, and Documentation Paper β’ 2503.15222 β’ Published Mar 19, 2025 β’ 1
Airavata Evaluation Suite Collection A collection of benchmarks used for evaluation of Airavata, an Hindi instruction-tuned model on top of Sarvam's OpenHathi base model. β’ 20 items β’ Updated Mar 2 β’ 10