Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction Paper • 2605.06191 • Published May 7
CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions Paper • 2604.05435 • Published Apr 7
Scalable multilingual PII annotation for responsible AI in LLMs Paper • 2510.06250 • Published Oct 9, 2025
Human + AI for Accelerating Ad Localization Evaluation Paper • 2509.12543 • Published Oct 6, 2025
LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents Paper • 2510.03418 • Published Oct 10, 2025
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents Paper • 2601.08988 • Published Jan 13
Running Agents Healthcare Document-Grounded QA Benchmark (Extending GDP.pdf) 📚 Document-work benchmark for healthcare persona