LEDGER: A Long-Context Benchmark of Corporate Annual Reports for Grounded Financial Retrieval and Extraction Paper • 2606.13100 • Published Jun 11 • 2
LEDGER Collection A Long-Context Benchmark of Corporate Annual Reports for Grounded Financial Retrieval and Extraction • 4 items • Updated 9 days ago • 1
LEDGER Collection A Long-Context Benchmark of Corporate Annual Reports for Grounded Financial Retrieval and Extraction • 4 items • Updated 9 days ago • 6
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation Paper • 2604.09497 • Published Apr 10 • 29
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate Paper • 2509.04492 • Published Sep 1, 2025 • 10
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate Paper • 2509.04492 • Published Sep 1, 2025 • 10
MLM vs CLM Collection Research material on research about pre-training encoders, with extensive comparison on masked language modeling paradigm vs causal langage modeling. • 5 items • Updated Dec 1, 2025