Mellum: Production-Grade in-IDE Contextual Code Completion with Multi-File Project Understanding Paper • 2510.05788 • Published Oct 7, 2025 • 4
Confidence and Stability of Global and Pairwise Scores in NLP Evaluation Paper • 2507.01633 • Published Jul 2, 2025
IMDB-WIKI-SbS: An Evaluation Dataset for Crowdsourced Pairwise Comparisons Paper • 2110.14990 • Published Oct 28, 2021
Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning Paper • 2402.06619 • Published Feb 9, 2024 • 57
Leveraging Contextual Information for Effective Entity Salience Detection Paper • 2309.07990 • Published Sep 14, 2023 • 8
Leveraging Contextual Information for Effective Entity Salience Detection Paper • 2309.07990 • Published Sep 14, 2023 • 8
Watset: Local-Global Graph Clustering with Applications in Sense and Frame Induction Paper • 1808.06696 • Published Aug 20, 2018
Best Prompts for Text-to-Image Models and How to Find Them Paper • 2209.11711 • Published Sep 23, 2022 • 3
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics Paper • 2102.01672 • Published Feb 2, 2021 • 1
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies Paper • 2305.16958 • Published May 26, 2023 • 2
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models Paper • 2206.04615 • Published Jun 9, 2022 • 6
CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription Paper • 2107.01091 • Published Jul 2, 2021