# rei-v1: evaluation Accuracy means agreeing with the teacher ensemble's answer, on research questions never seen in training. Temperatures: choice 1.55, noul 1.8. Trained on 74855 items, 3.0 epochs, 10478s. **Test: 76.7% accuracy** (majority baseline 42.8%), calibration error (ECE) 0.031. Validation: 79.0%. ## Test by task | | items | accuracy | majority baseline | soft CE | |---|---|---|---|---| | claim_check | 1085 | 81.4% | 33.9% | 0.537 | | custom | 421 | 72.0% | 15.0% | 0.674 | | needs_fresh | 181 | 82.9% | 60.2% | 0.375 | | next_step | 177 | 76.3% | 56.5% | 0.686 | | relevance | 1041 | 68.5% | 48.2% | 0.781 | | same_info | 180 | 88.3% | 86.1% | 0.323 | | search_type | 173 | 76.9% | 31.8% | 0.767 | | source_type | 347 | 80.1% | 18.2% | 0.647 | | worth_opening | 363 | 79.6% | 78.5% | 0.467 | Score tasks, mean absolute error in scale points: custom 0.50, relevance 0.33 Noul tasks, AUC (how well P(true) ranks true above false; 0.5 is chance): custom 0.83, needs_fresh 0.91, same_info 0.83, worth_opening 0.76 ## Test by language | | items | accuracy | majority baseline | soft CE | |---|---|---|---|---| | ar | 166 | 79.5% | 43.4% | 0.602 | | de | 254 | 80.3% | 50.4% | 0.563 | | en | 374 | 77.0% | 39.0% | 0.627 | | es | 322 | 79.8% | 44.4% | 0.555 | | fr | 275 | 80.4% | 44.7% | 0.550 | | hi | 140 | 75.7% | 43.6% | 0.654 | | id | 402 | 74.4% | 42.8% | 0.661 | | is | 474 | 69.6% | 42.0% | 0.764 | | ja | 468 | 78.4% | 41.0% | 0.575 | | ko | 301 | 75.7% | 42.9% | 0.608 | | pt | 306 | 75.5% | 43.8% | 0.635 | | ru | 212 | 75.0% | 42.9% | 0.631 | | zh | 274 | 80.7% | 40.1% | 0.536 | ## Test by profile | | items | accuracy | majority baseline | soft CE | |---|---|---|---|---| | academic | 742 | 77.8% | 41.6% | 0.591 | | anime | 976 | 78.4% | 38.7% | 0.570 | | general | 1486 | 74.4% | 44.5% | 0.684 | | programming | 764 | 78.0% | 45.9% | 0.580 |