|
Download eval/injection_eval.md from flowxai/injection: direct link, hf CLI and curl.
- Browser
- Download file 3.64 kB
-
https://huggingface.co/flowxai/injection/resolve/main/eval/injection_eval.md
- Command line
-
hf download hf://flowxai/injection/eval/injection_eval.md
-
curl -L -o injection_eval.md https://huggingface.co/flowxai/injection/resolve/main/eval/injection_eval.md
3.64 kB
| # injection evaluation | |
| **Threshold 0.02 (calibrated on the validation split, objective recall_at_fpr).** The config default of 0.5 scored 0.986 against 0.991 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold. | |
| > the chosen threshold 0.02 is the lowest value in the sweep, so the scores are compressed toward zero and this is underfitting rather than a calibration success. Check average precision before trusting the F1. | |
| ## Per language | |
| | Language | Support | P | R | F1 | Notes | | |
| |---|---|---|---|---|---| | |
| | `bg` Bulgarian | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `cs` Czech | 42 | 0.955 | 1.000 | 0.977 | | | |
| | `da` Danish | 41 | 0.976 | 1.000 | 0.988 | | | |
| | `de` German | 42 | 1.000 | 1.000 | 1.000 | | | |
| | `el` Greek | 42 | 1.000 | 1.000 | 1.000 | | | |
| | `en` English | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `es` Spanish | 40 | 0.976 | 1.000 | 0.988 | | | |
| | `et` Estonian | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `fi` Finnish | 42 | 1.000 | 1.000 | 1.000 | | | |
| | `fr` French | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `ga` Irish | 42 | 0.976 | 0.976 | 0.976 | | | |
| | `hr` Croatian | 42 | 0.977 | 1.000 | 0.988 | | | |
| | `hu` Hungarian | 42 | 0.955 | 1.000 | 0.977 | | | |
| | `it` Italian | 42 | 1.000 | 1.000 | 1.000 | | | |
| | `lt` Lithuanian | 42 | 1.000 | 1.000 | 1.000 | | | |
| | `lv` Latvian | 41 | 0.976 | 1.000 | 0.988 | | | |
| | `mt` Maltese | 42 | 0.804 | 0.976 | 0.882 | not in base model pretraining | | |
| | `nl` Dutch | 41 | 1.000 | 0.976 | 0.988 | | | |
| | `pl` Polish | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `pt` Portuguese | 42 | 0.977 | 1.000 | 0.988 | | | |
| | `ro` Romanian | 42 | 0.977 | 1.000 | 0.988 | | | |
| | `sk` Slovak | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `sl` Slovenian | 42 | 0.977 | 1.000 | 0.988 | | | |
| | `sv` Swedish | 41 | 1.000 | 1.000 | 1.000 | | | |
| | `tr` Turkish | 42 | 1.000 | 1.000 | 1.000 | | | |
| | `az` Azerbaijani | 42 | 1.000 | 1.000 | 1.000 | | | |
| The base-model note is a fact about pretraining, not a cause of the score beside it. `nsfw` Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this. | |
| ## Per register | |
| | Register | Support | P | R | F1 | FPR | | |
| |---|---|---|---|---|---| | |
| | `encoded_payload` | 155 | 1.000 | 1.000 | 1.000 | 0.000 | | |
| | `hypothetical_framing` | 156 | 1.000 | 0.994 | 0.997 | 0.000 | | |
| | `ignore_instructions` | 155 | 1.000 | 1.000 | 1.000 | 0.000 | | |
| | `instruction_in_document` | 155 | 1.000 | 0.994 | 0.997 | 0.000 | | |
| | `instruction_in_tool_output` | 148 | 1.000 | 0.993 | 0.997 | 0.000 | | |
| | `meta_question` | 0 | 0.000 | 0.000 | 0.000 | 0.012 | | |
| | `mundane_account_access` | 0 | 0.000 | 0.000 | 0.000 | 0.000 | | |
| | `mundane_informational` | 0 | 0.000 | 0.000 | 0.000 | 0.000 | | |
| | `mundane_operational` | 0 | 0.000 | 0.000 | 0.000 | 0.005 | | |
| | `mundane_transactional` | 0 | 0.000 | 0.000 | 0.000 | 0.005 | | |
| | `ordinary_instruction` | 0 | 0.000 | 0.000 | 0.000 | 0.000 | | |
| | `ordinary_question` | 0 | 0.000 | 0.000 | 0.000 | 0.006 | | |
| | `persona_override` | 155 | 1.000 | 1.000 | 1.000 | 0.000 | | |
| | `quoted_attack` | 0 | 0.000 | 0.000 | 0.000 | 0.009 | | |
| | `roleplay_benign` | 0 | 0.000 | 0.000 | 0.000 | 0.012 | | |
| | `security_discussion` | 0 | 0.000 | 0.000 | 0.000 | 0.009 | | |
| | `system_prompt_extraction` | 156 | 1.000 | 1.000 | 1.000 | 0.000 | | |
| | `technical_identifiers` | 0 | 0.000 | 0.000 | 0.000 | 0.006 | | |
| | `technical_payload` | 0 | 0.000 | 0.000 | 0.000 | 0.006 | | |
| ## Known weaknesses | |
| The three weakest languages by F1: `mt` at 0.882, `ga` at 0.976, `cs` at 0.977. | |
| These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table. |