Cluster & Tune: Boost Cold Start Performance in Text Classification Paper • 2203.10581 • Published Mar 20, 2022
Label Sleuth: From Unlabeled Text to a Classifier in a Few Hours Paper • 2208.01483 • Published Oct 31, 2022
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Paper • 2602.16763 • Published Jun 29
Automatically Extracting Challenge Sets for Non local Phenomena in Neural Machine Translation Paper • 1909.06814 • Published Sep 25, 2019
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs Paper • 2408.10646 • Published Aug 20, 2024
Mediators in Determining what Processing BERT Performs First Paper • 2104.06400 • Published Apr 13, 2021
Lossless and Near-Lossless Compression for Foundation Models Paper • 2404.15198 • Published Apr 5, 2024
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness Paper • 2502.19412 • Published Feb 17
PreQuEL: Quality Estimation of Machine Translation Outputs in Advance Paper • 2205.09178 • Published Dec 4, 2022
SERRANT: a syntactic classifier for English Grammatical Error Types Paper • 2104.02310 • Published Apr 7, 2021
Reinforcement Learning with Large Action Spaces for Neural Machine Translation Paper • 2210.03053 • Published Oct 6, 2022
GrASP: A Library for Extracting and Exploring Human-Interpretable Textual Patterns Paper • 2104.03958 • Published Jun 16, 2022
Enhancing the Transformer Decoder with Transition-based Syntax Paper • 2101.12640 • Published Oct 31, 2022
ComSum: Commit Messages Summarization and Meaning Preservation Paper • 2108.10763 • Published Aug 23, 2021
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Paper • 2605.29512 • Published May 28
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models Paper • 2601.15812 • Published Feb 17