moderation / eval /moderation_eval.md
bogdanraduta's picture
Upload eval/moderation_eval.md with huggingface_hub
e9d248b verified
|
Raw History Blame Contribute Delete
4 kB

moderation evaluation

Threshold 0.91 (calibrated on the validation split, objective macro_f1). The config default of 0.5 scored 0.866 against 0.904 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold.

Per language

Language Support P R F1 Notes
bg Bulgarian 52 1.000 1.000 1.000
cs Czech 60 0.983 0.950 0.966
da Danish 51 1.000 0.980 0.990
de German 54 1.000 0.981 0.991
el Greek 57 1.000 0.982 0.991
en English 62 0.984 0.984 0.984
es Spanish 52 1.000 0.923 0.960
et Estonian 59 1.000 0.966 0.983
fi Finnish 50 1.000 0.960 0.980
fr French 60 1.000 0.950 0.974
ga Irish 50 0.913 0.840 0.875
hr Croatian 60 0.967 0.983 0.975
hu Hungarian 63 0.969 0.984 0.976
it Italian 61 0.984 1.000 0.992
lt Lithuanian 52 1.000 0.981 0.990
lv Latvian 55 1.000 0.945 0.972
mt Maltese 48 0.872 0.854 0.863 not in base model pretraining
nl Dutch 56 0.982 0.982 0.982
pl Polish 61 1.000 0.984 0.992
pt Portuguese 52 0.980 0.962 0.971
ro Romanian 63 1.000 0.952 0.976
sk Slovak 48 0.979 0.979 0.979
sl Slovenian 59 1.000 0.983 0.991
sv Swedish 52 0.962 0.981 0.971
tr Turkish 60 0.937 0.983 0.959
az Azerbaijani 57 0.981 0.930 0.955

The base-model note is a fact about pretraining, not a cause of the score beside it. nsfw Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this.

Per register

Register Support P R F1 FPR
cyber_intrusion 119 1.000 0.958 0.979 0.000
cyber_intrusion_near_miss 0 0.000 0.000 0.000 0.000
defamation 111 1.000 1.000 1.000 0.000
defamation_near_miss 0 0.000 0.000 0.000 0.000
election_integrity 116 1.000 0.974 0.987 0.000
election_integrity_near_miss 0 0.000 0.000 0.000 0.000
extremism 130 1.000 0.908 0.952 0.000
extremism_near_miss 0 0.000 0.000 0.000 0.068
fraud_deception 124 1.000 0.952 0.975 0.000
fraud_deception_near_miss 0 0.000 0.000 0.000 0.051
hate_incitement 106 1.000 0.991 0.995 0.000
hate_incitement_near_miss 0 0.000 0.000 0.000 0.000
illicit_drugs 141 1.000 0.957 0.978 0.000
illicit_drugs_near_miss 0 0.000 0.000 0.000 0.000
mundane_account_access 0 0.000 0.000 0.000 0.024
mundane_informational 0 0.000 0.000 0.000 0.005
mundane_operational 0 0.000 0.000 0.000 0.000
mundane_transactional 0 0.000 0.000 0.000 0.000
property_crime 129 1.000 0.953 0.976 0.000
property_crime_near_miss 0 0.000 0.000 0.000 0.074
self_harm 99 1.000 0.949 0.974 0.000
self_harm_near_miss 0 0.000 0.000 0.000 0.111
sexual_exploitation 119 1.000 0.983 0.992 0.000
sexual_exploitation_near_miss 0 0.000 0.000 0.000 0.019
violent_facilitation 127 1.000 0.961 0.980 0.000
violent_facilitation_near_miss 0 0.000 0.000 0.000 0.018
weapons_cbrn 133 1.000 0.977 0.989 0.000
weapons_cbrn_near_miss 0 0.000 0.000 0.000 0.000

Known weaknesses

The three weakest languages by F1: mt at 0.863, ga at 0.875, az at 0.955.

These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table.