FilBench: Can LLMs Understand and Generate Filipino? Paper • 2508.03523 • Published Aug 5, 2025 • 1
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming Paper • 2601.11332 • Published Jan 16 • 1
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 6 days ago • 56
Multilinguality at the Edge: Developing Language Models for the Global South Paper • 2604.21637 • Published Apr 23
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation Paper • 2604.11290 • Published Apr 13 • 4
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation Paper • 2604.11290 • Published Apr 13 • 4
FilBench Eval Collection FilBench-Eval is an Open LLM Evaluation Suite for Philippine Languages. The eval runner is integrated with HuggingFace's lighteval. • 5 items • Updated Jan 14 • 1
FilBench: Can LLMs Understand and Generate Filipino? Paper • 2508.03523 • Published Aug 5, 2025 • 1
FilBench Eval Collection FilBench-Eval is an Open LLM Evaluation Suite for Philippine Languages. The eval runner is integrated with HuggingFace's lighteval. • 5 items • Updated Jan 14 • 1
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability Paper • 2506.01789 • Published Jun 2, 2025 • 15
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation Paper • 2505.24456 • Published May 30, 2025
Universal Dependencies for Tagalog Collection Models and dependency parsers for Tagalog using the UD_NewsCrawl dataset • 8 items • Updated May 29, 2025