philomath-1209/gpt2-reward_model_hh-rlhf Reinforcement Learning • 0.1B • Updated about 2 hours ago • 41
philomath-1209/gpt2-reward_model_hh-rlhf Reinforcement Learning • 0.1B • Updated about 2 hours ago • 41
philomath-1209/english-to-hindi-high-quality-training-data Viewer • Updated May 24, 2025 • 128k • 16 • 1
philomath-1209/english-to-hindi-high-quality-training-data Viewer • Updated May 24, 2025 • 128k • 16 • 1
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Paper • 2501.18362 • Published Jan 30, 2025 • 26
Isotonic/deberta-v3-base_finetuned_ai4privacy_v2 Token Classification • 0.2B • Updated Sep 13, 2024 • 34.5k • • 23
philomath-1209/programming-language-identification Text Classification • 83.5M • Updated Feb 2, 2024 • 7.7k • • 14