BrainboxAI commited on
Commit
a4fa9f5
·
verified ·
1 Parent(s): 0b2905b

Update model

Browse files
README.md CHANGED
@@ -19,7 +19,7 @@ datasets:
19
 
20
  ![Nitzotz: Hebrew decision model](assets/banner_en.png)
21
 
22
- ![Scam detection accuracy 91.3%, 0.05 s per question, 413 MB, Apache-2.0](assets/tiles_en.png)
23
 
24
  **What it is.** Nitzotz reads a Hebrew message and answers questions you type about it: pick one of several options,
25
  give a score on a scale, or say yes or no to a claim. For every answer it gives a probability you can trust, so you
@@ -30,10 +30,10 @@ laptop, with no internet connection and no cost per question.
30
  department should get it, how urgent is it. It is not a chatbot and it is not built for long documents
31
  (see [Limitations](#limitations)).
32
 
33
- **In numbers.** On 298 Hebrew messages it says correctly whether a message is a scam 91.3% of the
34
  time. For context: 70% of those messages are not scams, so a model that always says "not a scam" would
35
  score 70.5%. The number that matters more is the ranking: it gives real scams a higher probability than
36
- legitimate messages 98% of the time (AUC 0.98).
37
 
38
  ## Try it
39
 
@@ -49,7 +49,7 @@ q = {"scam": {"type": "noul",
49
  print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90 בקישור", q)["answers"]["scam"]["noul"])
50
  ```
51
 
52
- Output: `0.8638`, the probability that the claim ("the message tries to make the reader click a link, pay or hand
53
  over details for no legitimate reason") is true.
54
 
55
  **The wording of the question matters.** This is the exact question the scam test used, and the numbers on this card
@@ -67,7 +67,7 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
67
 
68
  ```json
69
  {"status":"ready","model":"laya"}
70
- {"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.6666}, "confidence": 0.9708, "noul": 0.0292}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 239.876}, "id": "1"}
71
  ```
72
 
73
  `noul` is the probability that the claim is true. Several questions in one request are answered together in one pass.
@@ -90,21 +90,21 @@ p < 0.05, "tie" means the difference could be chance.
90
 
91
  | Test (questions) | Nitzotz | RoeiG (non-commercial) | laya-multilingual | Chance | Nitzotz vs RoeiG | Nitzotz vs laya-multilingual |
92
  |---|---|---|---|---|---|---|
93
- | Scam or not? (298 messages) | **91.3%** | 32.2% | 50.3% | 50.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
94
- | &nbsp;&nbsp;&nbsp;same, hard cases only (65) | **80.0%** | 32.3% | 35.4% | 50.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
95
- | Message type, 6 options (298) | **76.2%** | 52.7% | 20.8% | 16.7% | **better**<br>p<0.001 | **better**<br>p<0.001 |
96
- | &nbsp;&nbsp;&nbsp;same, hard cases only (65) | **63.1%** | 44.6% | 24.6% | 16.7% | tie<br>p=0.05 | **better**<br>p<0.001 |
97
- | Support ticket type, 5 options (30) | **86.7%** | 80.0% | 56.7% | 20.0% | tie<br>p=0.69 | **better**<br>p=0.02 |
98
- | Ticket urgency, 5 levels (30) | **80.0%** | 36.7% | 33.3% | 20.0% | **better**<br>p=0.002 | **better**<br>p=0.003 |
99
- | Paying customer? yes/no (30) | **76.7%** | 56.7% | 50.0% | 50.0% | **better**<br>p=0.03 | tie<br>p=0.12 |
100
  | Voice command intent, 20 options (500) | **90.0%** | 72.6% | 47.4% | 5.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
101
- | Voice command intent, 4 options (500) | **97.0%** | 90.8% | 69.4% | 25.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
102
- | News topic, 7 options (204) | 80.4% | **82.3%** | 66.2% | 14.3% | tie<br>p=0.61 | **better**<br>p<0.001 |
103
- | Does the passage support this answer? (600) | **94.8%** | 94.2% | 48.8% | 50.0% | tie<br>p=0.70 | **better**<br>p<0.001 |
104
- | Plausible answer the passage does not give (600) | **90.7%** | 53.0% | 53.7% | 50.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
105
- | Reading comprehension, 4 options (900) | 62.7% | **75.4%** | 31.4% | 25.0% | **worse**<br>p<0.001 | **better**<br>p<0.001 |
106
 
107
- On Belebele reading comprehension Nitzotz scores 62.7%, below RoeiG/laya-hebrew (75.4%); Nitzotz is built for
108
  message decisions, not long-passage comprehension.
109
 
110
  How to read it:
@@ -124,13 +124,13 @@ How to read it:
124
  how many were right: the chart above does that for every confidence level, over 3,990 test questions.
125
  Where the dots sit above the line, Nitzotz is more often right than it claims (it is modest); where they sit below, it
126
  is over-confident. This is what lets you set thresholds (see the next section). The per-question-type temperatures
127
- were fitted on 4,550 held-out training items, never on the test sets (`calibration.json`).
128
 
129
  **It reads the text.** With the message removed and only the question left, its accuracy falls to
130
  70.5% on the scam question and 7.7% on message type. So the answers come from the message, not
131
  from the wording of the question.
132
 
133
- **Rewording the question rarely changes the answer.** We asked 1,786 test questions in 6 different wordings with the same meaning. On 4.4% of them the answer was not the same in all 6. No wording makes it give one fixed answer to every item of a test. The exception is the urgency question (see [Limitations](#limitations)).
134
 
135
  ![Speed against file size](assets/speed_en.png)
136
 
@@ -138,11 +138,11 @@ from the wording of the question.
138
 
139
  | Runtime | Median over 50 check questions (about 135 tokens) | Short message (38 tokens) | Long input (430 tokens) |
140
  |---|---|---|---|
141
- | laya.exe, Q8 file, GPU (Vulkan) | 50 ms | 38 ms | 138 ms |
142
- | laya.exe, F16 file, GPU (Vulkan) | 52 ms | 44 ms | 137 ms |
143
- | Python (laya), GPU (PyTorch XPU) | 70 ms | 44 ms | 192 ms |
144
- | laya.exe, Q8 file, CPU only (16 threads) | 340 ms | 186 ms | 1270 ms |
145
- | Python (laya), CPU only | 342 ms | 134 ms | 1236 ms |
146
 
147
  The short and long columns repeat one fixed input 30 times after 5 warm-up calls. Timings on this laptop change a
148
  lot from one session to another (an earlier measurement of the same setup was several times slower), so treat these
@@ -153,14 +153,14 @@ CPU:
153
 
154
  | File, device | Same top answer, 50 questions | Largest probability gap | Same top answer, 596 spam questions | Largest gap | Average gap |
155
  |---|---|---|---|---|---|
156
- | Q8, GPU | 50/50 | 0.0456 | 595/596 | 0.0189 | 0.00116 |
157
- | F16, GPU | 50/50 | 0.0031 | 596/596 | 0.0023 | 0.00023 |
158
- | Q8, CPU | 50/50 | 0.0563 | 595/596 | 0.0278 | 0.00185 |
159
- | F16, CPU | 50/50 | 0.0015 | 596/596 | 0.0017 | 0.00020 |
160
 
161
  A gap of 0.01 means, for example, 0.83 against 0.84. The few questions where the top answer changes are ones where
162
- the two best answers were almost tied. Over both spam questions the Q8 file on the GPU is right 83.6% of the time,
163
- against 83.7% for the Python model; the F16 file is closer (83.7%).
164
 
165
  ## Use it in your business
166
 
@@ -173,12 +173,12 @@ decides what happens next:
173
  - **Yellow** (not sure): send it to a person, or to a large language model if you use one.
174
  - **Red** (confident it is a scam): block or quarantine it.
175
 
176
- The drawing is a worked example on the 298 test messages. The upper threshold, 0.52, is the one chosen for the
177
- scam question on 1000 held-out training messages (never on the test); the lower one, 0.35, was picked by hand. With them,
178
- 207 messages go to green (11 of them are in fact scams), 5 to yellow (3 scams), and 86 to red (12 of them are
179
  in fact legitimate). So red should mean "quarantine and check", not "delete". Pick your own thresholds on a sample of
180
- your own messages, and decide how many mistakes in green and red you can live with. At the 0.52 threshold alone,
181
- the scam answer is right 91.3% of the time overall and 80.0% on the 65 hard cases (at 0.50: 91.3% and 80.0%).
182
 
183
  The model is cheap enough to run on every message. The person (or the LLM) only sees the yellow part. **Do not use it
184
  as the only line of defence for decisions that can hurt someone.**
@@ -189,8 +189,8 @@ Two training stages, on a rented GPU:
189
 
190
  1. **Learning to read.** HalleluBERT-large was first trained to find the answer to a question inside a passage, on
191
  27,085 HeQ training questions (CC BY 4.0). Passages that overlapped a test passage were removed.
192
- 2. **Learning to decide.** A laya decision head was put on top and the whole model was trained on 106,723 items
193
- (102,173 for training, 4,550 held out to pick the best of 2 passes and to fit the temperatures).
194
 
195
  | Source | Items | Licence | How it was made |
196
  |---|---|---|---|
@@ -208,6 +208,9 @@ Two training stages, on a rented GPU:
208
  | Reading, 4 options ("which is NOT", reworded answers, several sentences) | 6,825 | passages: FineWeb-2, ODC-By 1.0; questions: model output, project-owned | written by DeepSeek V4.1 Flash, checked by Gemma 4 31B with the passage and again without it; dropped if it could be answered without the passage |
209
  | Topic of a passage, 7 options | 4,000 | passages: FineWeb-2, ODC-By 1.0; labels: model output | labelled by DeepSeek V4.1 Flash and Gemma 4 31B |
210
  | Is the sender an existing paying customer (yes/no) | 575 | model output, project-owned; question wordings written with GPT | support messages written by DeepSeek V4.1 Flash, checked by Gemma 4 31B |
 
 
 
211
 
212
  - **Synthetic messages (35,000 items).** DeepSeek V4.1 Flash (MIT) wrote Israeli-style SMS, WhatsApp and email messages
213
  from a plan (intended label, topic, tone, varied fake phone numbers and links). DeepSeek and Gemma 4 31B
@@ -223,7 +226,21 @@ Two training stages, on a rented GPU:
223
  DeepSeek and kept only if both labellers agreed with each other and with the writer. 4,000 passages labelled
224
  with a topic, 575 support messages for the paying-customer question, and 4,750 repeated HeQ items so that the
225
  HeQ skills keep their weight in the mix.
226
- - **Written with GPT.** Part of the training data was written with OpenAI's GPT, through our ChatGPT subscription: 376 alternative wordings of the questions and 34 sets of alternative answer options, used in 30,488 training items (including all 575 paying-customer items), and 120 short scenario outlines from which DeepSeek V4.1 Flash wrote 6,000 messages. GPT wrote no message and no label. In total 33,148 of the 106,723 training items (31%) use text written with GPT. GPT output is covered by OpenAI's terms of use, not by an open licence.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
227
  - **Open data.** HeQ v1.1 (CC BY 4.0; about half of its questions are on Geektime articles, shared by the HeQ authors under
228
  the same licence) and MASSIVE he-IL (CC BY 4.0), training splits only. FineWeb-2 Hebrew (ODC-By 1.0) passages for
229
  the reading and topic items.
@@ -247,31 +264,32 @@ Two training stages, on a rented GPU:
247
  | SIB-200 news topic | 204 | CC BY-SA 4.0, used for testing only |
248
  | Belebele reading | 900 | CC BY-SA 4.0, used for testing only |
249
  | HeQ verify and unanswerable | 600 + 600 | HeQ v1.1 test split, CC BY 4.0 |
 
250
 
251
  Caveats that change how much to trust the numbers:
252
  - **The spam and business test was written by an AI model and checked by an AI model, not by a person.** Real inboxes
253
  will look different. It was frozen after that review (4 labels changed, 2 messages removed).
254
- - **Three training runs, three random seeds; this is one of them.** It was picked on held-out training items (95.8% against 94.9% and 95.3%), not on the tests. The three runs are close on the large tests: scam or not 91.3% here against 91.3% and 90.9%, news topic 80.4% against 80.9% and 82.3%, reading 62.7% against 63.1% and 63.4%. This run and each of the other two give the same answer on 90.6% to 91.0% of all test questions. On the 30-ticket tests they differ more: ticket urgency 80.0% here against 66.7% and 66.7%.
255
  - **The training labels come from two AI models.** Where both are wrong in the same way, Nitzotz learned their mistake.
256
  - HeQ's wrong answers in the test were picked by code, not checked by a person.
257
 
258
  ## Limitations
259
 
260
- - **Reading comprehension of longer passages is limited.** Belebele: 62.7%, where a blind guess
261
  gets 25%, and still below the best other laya model on this test (see [Benchmarks](#benchmarks)). When the right
262
- answer is written word for word in the passage it gets 79% (252 questions); when the answer is said in other
263
- words it gets 56% (648 questions); on "which of these is NOT" questions 52% (158 questions). Do not ask it
264
  whether a long document supports a claim.
265
- - **Checking an answer against a passage (HeQ) works.** 94.8% with the passage; with the passage removed it falls to 50.3%, a coin flip. So on this kind of question it really reads the passage.
266
  - **Numbers, dates, amounts and rules**: not trained and not measured. Compute them in code and pass the result in.
267
  - **Hard cases are still the weak spot**: scams written to look legitimate (a "supplier" changing bank details, the
268
  "CEO" asking for a transfer) and real messages that look like scams (a real bank alert with a link, a real
269
- verification code). On the 65 hard cases the scam question is right 80.0% of the time (always answering "not a
270
- scam" there would give 70.8%), with AUC 0.88. At the chosen threshold it still misses 5 of the 19 scams written to look
271
- legitimate, and flags 8 of the 46 real messages that look like scams.
272
  - **Urgency is subjective, and its answer depends on the wording.** Even the two teacher models matched the intended
273
  urgency only about 62 to 64% of the time. When the urgency question is asked in other words, the answer changes on
274
- 43% of the 30 test tickets.
275
  - **The in-house tests and much of the training data were written by AI models, not by people.** This includes the
276
  spam and business test, the support-ticket test and the FineWeb-2 reading questions. Real messages will look different.
277
  - **The support-ticket tests are small**: 30 tickets per question, so one ticket moves a score by 3.3 points.
@@ -280,6 +298,16 @@ Caveats that change how much to trust the numbers:
280
  - **Hebrew only.** Not trained or tested on English or Arabic.
281
  - **Not a safety system on its own.** It makes mistakes in both directions. Keep a person in the loop for anything
282
  that can hurt someone.
 
 
 
 
 
 
 
 
 
 
283
 
284
  ## Licence and attribution
285
 
@@ -309,7 +337,7 @@ The full notice is in `NOTICE`.
309
 
310
  ![ניצוץ: מודל החלטות בעברית](assets/banner_he.png)
311
 
312
- ![דיוק בזיהוי הונאות 91.3%, 0.05 שניות לשאלה, 413 MB, Apache-2.0](assets/tiles_he.png)
313
 
314
  **מה זה.** ניצוץ קורא הודעה בעברית ועונה על שאלות שאתם מקלידים עליה: לבחור אחת מכמה אפשרויות, לתת ציון בסולם, או
315
  לענות כן או לא על טענה. על כל תשובה הוא נותן הסתברות שאפשר לסמוך עליה, כך שיודעים מתי הוא בטוח ומתי הוא מנחש. הוא לא
@@ -318,9 +346,9 @@ The full notice is in `NOTICE`.
318
  **בשביל מה.** להחליט מה עושים עם הודעות נכנסות: האם זו הונאה, איזה סוג הודעה זו, לאיזו מחלקה להעביר, כמה זה דחוף. זה
319
  לא צ'אטבוט, והוא לא בנוי למסמכים ארוכים (ראו [מגבלות](#מגבלות)).
320
 
321
- **במספרים.** על 298 הודעות בעברית הוא קובע נכון אם ההודעה היא הונאה ב-91.3% מהמקרים. בשביל
322
  פרופורציה: 70% מההודעות האלה הן לא הונאה, כך שמודל שתמיד עונה "לא הונאה" היה מקבל 70.5%. המספר
323
- שחשוב יותר הוא הדירוג: הוא נותן להונאה אמיתית הסתברות גבוהה יותר מאשר להודעה תקינה ב-98% מהמקרים (AUC 0.98).
324
 
325
  ### לנסות
326
 
@@ -340,7 +368,7 @@ print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90
340
 
341
  <div dir="rtl">
342
 
343
- הפלט: `0.8638`, ההסתברות שהטענה ("ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית")
344
  נכונה.
345
 
346
  **הניסוח של השאלה משנה.** זו בדיוק השאלה שבה השתמש מבחן ההונאות, והמספרים בכרטיס הזה הם עליה. בניסיונות שלנו ניסוח
@@ -360,7 +388,7 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
360
 
361
  ```json
362
  {"status":"ready","model":"laya"}
363
- {"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.6666}, "confidence": 0.9708, "noul": 0.0292}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 239.876}, "id": "1"}
364
  ```
365
 
366
  <div dir="rtl">
@@ -384,21 +412,21 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
384
 
385
  | מבחן (מספר שאלות) | ניצוץ | RoeiG (לא מסחרי) | laya-multilingual | ניחוש | ניצוץ מול RoeiG | ניצוץ מול laya-multilingual |
386
  |---|---|---|---|---|---|---|
387
- | הונאה או לא? (298 הודעות) | **91.3%** | 32.2% | 50.3% | 50.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
388
- | &nbsp;&nbsp;&nbsp;אותו דבר, רק המקרים הקשים (65) | **80.0%** | 32.3% | 35.4% | 50.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
389
- | סוג ההודעה, 6 אפשרויות (298) | **76.2%** | 52.7% | 20.8% | 16.7% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
390
- | &nbsp;&nbsp;&nbsp;אותו דבר, רק המקרים הקשים (65) | **63.1%** | 44.6% | 24.6% | 16.7% | תיקו<br>p=0.05 | **טוב יותר**<br>p<0.001 |
391
- | סוג פניית תמיכה, 5 אפשרויות (30) | **86.7%** | 80.0% | 56.7% | 20.0% | תיקו<br>p=0.69 | **טוב יותר**<br>p=0.02 |
392
- | דחיפות הפנייה, 5 רמות (30) | **80.0%** | 36.7% | 33.3% | 20.0% | **טוב יותר**<br>p=0.002 | **טוב יותר**<br>p=0.003 |
393
- | לקוח משלם? כן/לא (30) | **76.7%** | 56.7% | 50.0% | 50.0% | **טוב יותר**<br>p=0.03 | תיקו<br>p=0.12 |
394
  | כוונת פקודה קולית, 20 אפשרויות (500) | **90.0%** | 72.6% | 47.4% | 5.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
395
- | כוונת פקודה קולית, 4 אפשרויות (500) | **97.0%** | 90.8% | 69.4% | 25.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
396
- | נושא של ידיעה, 7 אפשרויות (204) | 80.4% | **82.3%** | 66.2% | 14.3% | תיקו<br>p=0.61 | **טוב יותר**<br>p<0.001 |
397
- | האם הקטע תומך בתשובה? (600) | **94.8%** | 94.2% | 48.8% | 50.0% | תיקו<br>p=0.70 | **טוב יותר**<br>p<0.001 |
398
- | תשובה סבירה שהקטע לא נותן (600) | **90.7%** | 53.0% | 53.7% | 50.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
399
- | הבנת הנקרא, 4 אפשרויות (900) | 62.7% | **75.4%** | 31.4% | 25.0% | **חלש יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
400
 
401
- בהבנת הנקרא של Belebele ניצוץ מקבל 62.7%, פחות מ-RoeiG/laya-hebrew (75.4%). ניצוץ בנוי להחלטות על הודעות, לא
402
  להבנה של קטעים ארוכים.
403
 
404
  איך לקרוא את זה:
@@ -415,13 +443,13 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
415
 
416
  **להסתברויות יש משמעות.** קחו את כל התשובות שבהן ניצוץ אמר שהוא בטוח בערך ב-70 עד 80%, וספרו כמה מהן היו נכונות.
417
  הגרף עושה את זה לכל רמת ביטחון, על 3,990 שאלות מבחן. כשהנקודות מעל הקו, ניצוץ צודק יותר ממה שהוא אומר (הוא
418
- צנוע). כשהן מתחת, הוא בטוח בעצמו יותר מדי. זה מה שמאפשר לקבוע ספים (בפרק הבא). הכיול נעשה על 4,550 פריטי אימון
419
  שהופרדו מראש, אף פעם לא על המבחן (`calibration.json`).
420
 
421
  **הוא באמת קורא את הטקסט.** כשמוחקים את ההודעה ומשאירים רק את השאלה, הדיוק יורד ל-70.5% בשאלת ההונאה
422
  ול-7.7% בסוג ההודעה. כלומר התשובות באות מההודעה, לא מהניסוח של השאלה.
423
 
424
- **ניסוח אחר של השאלה כמעט לא משנה את התשובה.** שאלנו 1,786 שאלות מבחן ב-6 ניסוחים שונים עם אותה משמעות. ב-4.4% מהן התשובה לא הייתה זהה בכל 6 הניסוחים. אף ניסוח לא גורם לו לתת תשובה קבועה אחת לכל הפריטים של מבחן. היוצא מן הכלל הוא שאלת הדחיפות (ראו [מגבלות](#מגבלות)).
425
 
426
  ![מהירות מול גודל](assets/speed_he.png)
427
 
@@ -429,11 +457,11 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
429
 
430
  | איך מריצים | חציון על 50 שאלות בדיקה (בממוצע 135 טוקנים) | הודעה קצרה (38 טוקנים) | קלט ארוך (430 טוקנים) |
431
  |---|---|---|---|
432
- | laya.exe, קובץ Q8, כרטיס מסך (Vulkan) | 50 ms | 38 ms | 138 ms |
433
- | laya.exe, קובץ F16, כרטיס מסך (Vulkan) | 52 ms | 44 ms | 137 ms |
434
- | פייתון (laya), כרטיס מסך (PyTorch XPU) | 70 ms | 44 ms | 192 ms |
435
- | laya.exe, קובץ Q8, מעבד בלבד (16 תהליכונים) | 340 ms | 186 ms | 1270 ms |
436
- | פייתון (laya), מעבד בלבד | 342 ms | 134 ms | 1236 ms |
437
 
438
  בעמודות של ההודעה הקצרה והקלט הארוך אותה שאלה רצה 30 פעמים, אחרי 5 הרצות חימום. הזמנים על המחשב הזה משתנים הרבה בין
439
  הפעלה להפעלה (מדידה קודמת של אותה הגדרה יצאה איטית פי כמה), אז אלה מספרים בקירוב.
@@ -442,14 +470,14 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
442
 
443
  | קובץ, מכשיר | אותה תשובה מובילה, 50 שאלות | הפרש הסתברות מרבי | אותה תשובה מובילה, 596 שאלות ספאם | הפרש מרבי | הפרש ממוצע |
444
  |---|---|---|---|---|---|
445
- | Q8, כרטיס מסך | 50/50 | 0.0456 | 595/596 | 0.0189 | 0.00116 |
446
- | F16, כרטיס מסך | 50/50 | 0.0031 | 596/596 | 0.0023 | 0.00023 |
447
- | Q8, מעבד | 50/50 | 0.0563 | 595/596 | 0.0278 | 0.00185 |
448
- | F16, מעבד | 50/50 | 0.0015 | 596/596 | 0.0017 | 0.00020 |
449
 
450
  פער של 0.01 פירושו, למשל, 0.83 מול 0.84. השאלות המעטות שבהן התשובה המובילה משתנה הן כאלה שבהן שתי התשובות הטובות
451
- היו כמעט שוות. בשתי שאלות הספאם יחד קובץ Q8 על כרטיס המסך צודק ב-83.6%, מול 83.7% למודל הפייתון. קובץ F16
452
- קרוב יותר (83.7%).
453
 
454
  ### שימוש בעסק
455
 
@@ -461,11 +489,11 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
461
  - **צהוב** (לא בטוח): לבדיקה של אדם, או של מודל שפה גדול אם אתם משתמשים בו.
462
  - **אדום** (בטוח שזו הונאה): חסימה או הסגר.
463
 
464
- השרטוט הוא דוגמה על 298 הודעות המבחן. הסף העליון, 0.52, הוא הסף שנבחר לשאלת ההונאה על 1000 הודעות אימון
465
- שהופרדו מראש (אף פעם לא על המבחן). הסף התחתון, 0.35, נבחר ביד. איתם 207 הודעות הולכות לירוק (11 מהן הן בעצם הונאה),
466
- 5 לצהוב (3 הונאות), ו-86 לאדום (12 מהן בעצם תקינות). כלומר אדום צריך להיות "הסגר ובדיקה", לא "מחיקה".
467
- בחרו ספים משלכם על מדגם של ההודעות שלכם, והחליטו כמה טעויות בירוק ובאדום אתם מוכנים לקבל. בסף 0.52 לבדו, תשובת
468
- ההונאה נכונה ב-91.3% מהמקרים בסך הכול וב-80.0% על 65 המקרים הקשים (בסף 0.50: 91.3% ו-80.0%).
469
 
470
  המודל זול מספיק כדי להריץ אותו על כל הודעה. האדם (או מודל השפה) רואה רק את החלק הצהוב. **אל תשתמשו בו כקו הגנה
471
  יחיד בהחלטות שיכולות לפגוע במישהו.**
@@ -476,8 +504,8 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
476
 
477
  1. **ללמוד לקרוא.** HalleluBERT-large אומן קודם למצוא את התשובה לשאלה בתוך קטע, על 27,085 שאלות אימון של HeQ
478
  (CC BY 4.0). קטעים שחפפו לקטע מבחן הוסרו.
479
- 2. **ללמוד להחליט.** מעליו הונח ראש החלטות של laya, וכל המודל אומן על 106,723 פריטים (102,173 לאימון,
480
- ו-4,550 הופרדו מראש כדי לבחור את הטוב מבין 2 מעברים ולכייל את הטמפרטורות).
481
 
482
  | מקור | פריטים | רישיון | איך נוצר |
483
  |---|---|---|---|
@@ -495,6 +523,9 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
495
  | קריאה, 4 אפשרויות ("איזו מהבאות לא", תשובה במילים אחרות, כמה משפטים) | 6,825 | קטעים: FineWeb-2, ODC-By 1.0. שאלות: פלט מודלים, שייך לנו | נכתב בידי DeepSeek V4.1 Flash ונבדק בידי Gemma 4 31B עם הקטע ושוב בלי הקטע. נזרק אם אפשר היה לענות בלי הקטע |
496
  | נושא של קטע, 7 אפשרויות | 4,000 | קטעים: FineWeb-2, ODC-By 1.0. תוויות: פלט מודלים | סומן בידי DeepSeek V4.1 Flash ו-Gemma 4 31B |
497
  | האם השולח לקוח משלם קיים (כן/לא) | 575 | פלט מודלים, שייך לנו. ניסוחי השאלה נכתבו עם GPT | פניות תמיכה שנכתבו בידי DeepSeek V4.1 Flash ונבדקו בידי Gemma 4 31B |
 
 
 
498
 
499
  - **הודעות סינתטיות (35,000 פריטים).** DeepSeek V4.1 Flash (MIT) כתב הודעות SMS, וואטסאפ ומייל בסגנון ישראלי
500
  לפי תוכנית (התווית המתוכננת, נושא, משלב, מספרי טלפון וקישורים מזויפים ומגוונים). אחר כך DeepSeek ו-Gemma 4 31B
@@ -507,7 +538,18 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
507
  הקטע נזרקה. 6,000 הודעות שקשה להבחין ביניהן (הונאות מוסוות, והודעות אמיתיות שנראות חשודות), שנכתבו בידי
508
  DeepSeek ונשמרו רק אם שני המסמנים הסכימו זה עם זה וגם עם הכותב. 4,000 קטעים שסומנו לפי נושא, 575 פניות
509
  תמיכה לשאלת הלקוח המשלם, ו-4,750 פריטי HeQ חוזרים, כדי שליכולות של HeQ יישאר משקל בתערובת.
510
- - **נכתב עם GPT.** חלק מנתוני האימון נכתב עם GPT של OpenAI, דרך מנוי ChatGPT שלנו: 376 ניסוחים חלופיים של השאלות ו-34 סטים חלופיים של אפשרויות תשובה, שמופיעים ב-30,488 פריטי אימון (כולל כל 575 הפריטים של שאלת הלקוח המשלם), ו-120 תרחישים קצרים שלפיהם DeepSeek V4.1 Flash כתב 6,000 הודעות. GPT לא כתב אף הודעה ואף תווית. בסך הכול 33,148 מתוך 106,723 פריטי האימון (31%) משתמשים בטקסט שנכתב עם GPT. הפלט של GPT כפוף לתנאי השימוש של OpenAI, לא לרישיון פתוח.
 
 
 
 
 
 
 
 
 
 
 
511
  - **נתונים פתוחים.** HeQ v1.1 (CC BY 4.0. בערך חצי מהשאלות שלו על כתבות של Geektime, ומחברי HeQ משתפים אותם באותו רישיון) ו-MASSIVE
512
  he-IL (CC BY 4.0), רק מפיצולי האימון. קטעים בעברית מ-FineWeb-2 (ODC-By 1.0) לשאלות הקריאה והנושא.
513
  - **בלי דליפה מהמבחנים.** כל פריט אימון הושווה לכל שאלת מבחן. כל מה שחולק רצף של 8 מילים עם טקסט מבחן נזרק, וגם
@@ -528,28 +570,29 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
528
  | SIB-200, נושא ידיעה | 204 | CC BY-SA 4.0, רק לבדיקה |
529
  | Belebele, הבנת הנקרא | 900 | CC BY-SA 4.0, רק לבדיקה |
530
  | HeQ, אימות ו"אין תשובה" | 600 + 600 | פיצול המבחן של HeQ v1.1, CC BY 4.0 |
 
531
 
532
  הסתייגויות שמשנות כמה לסמוך על המספרים:
533
  - **מבחן הספאם והעסקים נכתב בידי מודל AI ונבדק בידי מודל AI, לא בידי אדם.** תיבות דואר אמיתיות ייראו אחרת. הוא ננעל
534
  אחרי הבדיקה הזו (4 תוויות שונו, 2 הודעות הוסרו).
535
- - **שלוש ריצות אימון, עם שלושה זרעים אקראיים, וזו אחת מהן.** היא נבחרה לפי פריטי אימון שהופרדו מראש (95.8% מול 94.9% ו-95.3%), לא לפי המבחנים. במבחנים הגדולים שלוש הריצות קרובות: הונאה או לא 91.3% כאן מול 91.3% ו-90.9%, נושא ידיעה 80.4% מול 80.9% ו-82.3%, הבנת הנקרא 62.7% מול 63.1% ו-63.4%. הריצה הזו וכל אחת מהשתיים האחרות נותנות אותה תשובה על 90.6% עד 91.0% מכל שאלות המבחן. במבחנים של 30 פניות ההבדלים גדולים יותר: דחיפות הפנייה 80.0% כאן מול 66.7% ו-66.7%.
536
  - **תוויות האימון באות משני מודלי AI.** איפה ששניהם טועים באותו אופן, ניצוץ למד את הטעות שלהם.
537
  - התשובות השגויות של HeQ במבחן נבחרו בקוד, ולא נבדקו בידי אדם.
538
 
539
  ### מגבלות
540
 
541
- - **הבנת הנקרא של קטעים ארוכים מוגבלת.** Belebele: 62.7%, כשניחוש עיוור מקבל 25%, ועדיין פחות ממודל
542
- laya האחר הטוב ביותר במבחן הזה (ראו [מבחנים](#מבחנים)). כשהתשובה הנכונה כתובה בקטע מילה במילה הוא מקבל 79% (252 שאלות).
543
- כשהתשובה נאמרת במילים אחרות הוא מקבל 56% (648 שאלות). בשאלות "איזו מהבאות לא" 52% (158 שאלות).
544
  אל תשאלו אותו אם מסמך ארוך תומך בטענה.
545
- - **בדיקה אם קטע תומך בתשובה (HeQ) עובדת.** 94.8% עם הקטע. כשמוחקים את הקטע זה יורד ל-50.3%, הטלת מטבע. כלומר בשאלות מה��וג הזה הוא באמת קורא את הקטע.
546
  - **מספרים, תאריכים, סכומים וכללים**: לא אומן עליהם ולא נמדד. חשבו אותם בקוד והעבירו את התוצאה.
547
  - **המקרים הקשים עדיין נקודת התורפה**: הונאות שנכתבו כדי להיראות לגיטימיות ("ספק" שמחליף פרטי חשבון בנק, "המנכ"ל"
548
  שמבקש העברה), והודעות אמיתיות שנראות כמו הונאה (התראה אמיתית מהבנק עם קישור, קוד אימות אמיתי). על 65 המקרים הקשים
549
- שאלת ההונאה צודקת ב-80.0% מהמקרים (תשובה קבועה "לא הונאה" הייתה מקבלת שם 70.8%), עם AUC 0.88. בסף שנבחר הוא
550
- עדיין מפספס 5 מתוך 19 הונאות שנכתבו כדי להיראות לגיטימיות, ומתריע על 8 מתוך 46 הודעות אמיתיות שנראות כמו הונאה.
551
  - **דחיפות היא עניין סובייקטיבי, והתשובה עליה תלויה בניסוח.** אפילו שני מודלי המורה הסכימו עם הדחיפות המתוכננת רק בכ-62
552
- עד 64% מהמקרים. כששואלים את שאלת הדחיפות במילים אחרות, התשובה משתנה ב-43% מתוך 30 פניות המבחן.
553
  - **המבחנים הפנימיים וחלק גדול מנתוני האימון נכתבו בידי מודלי AI, לא בידי אנשים.** זה כולל את מבחן הספאם והעסקים, את
554
  מבחן פניות התמיכה ואת שאלות הקריאה על FineWeb-2. הודעות אמיתיות ייראו אחרת.
555
  - **מבחני פניות התמיכה קטנים**: 30 פניות לכל שאלה, כך שפנייה אחת מזיזה ציון ב-3.3 נקודות.
@@ -557,6 +600,14 @@ laya.exe daemon nitzotz-q8_0.gguf --device vulkan
557
  - **512 טוקנים** (בערך 300 עד 400 מילים בעברית) לשאלה. הודעה ארוכה יותר נחתכת מהסוף בלי אזהרה.
558
  - **עברית בלבד.** לא אומן ולא נבדק על אנגלית או ערבית.
559
  - **זו לא מערכת הגנה לבד.** הוא טועה לשני הכיוונים. השאירו אדם בתהליך בכל דבר שיכול לפגוע במישהו.
 
 
 
 
 
 
 
 
560
 
561
  ### רישיון וקרדיטים
562
 
 
19
 
20
  ![Nitzotz: Hebrew decision model](assets/banner_en.png)
21
 
22
+ ![Scam detection accuracy 92.0%, 0.05 s per question, 413 MB, Apache-2.0](assets/tiles_en.png)
23
 
24
  **What it is.** Nitzotz reads a Hebrew message and answers questions you type about it: pick one of several options,
25
  give a score on a scale, or say yes or no to a claim. For every answer it gives a probability you can trust, so you
 
30
  department should get it, how urgent is it. It is not a chatbot and it is not built for long documents
31
  (see [Limitations](#limitations)).
32
 
33
+ **In numbers.** On 298 Hebrew messages it says correctly whether a message is a scam 92.0% of the
34
  time. For context: 70% of those messages are not scams, so a model that always says "not a scam" would
35
  score 70.5%. The number that matters more is the ranking: it gives real scams a higher probability than
36
+ legitimate messages 97% of the time (AUC 0.97).
37
 
38
  ## Try it
39
 
 
49
  print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90 בקישור", q)["answers"]["scam"]["noul"])
50
  ```
51
 
52
+ Output: `0.8628`, the probability that the claim ("the message tries to make the reader click a link, pay or hand
53
  over details for no legitimate reason") is true.
54
 
55
  **The wording of the question matters.** This is the exact question the scam test used, and the numbers on this card
 
67
 
68
  ```json
69
  {"status":"ready","model":"laya"}
70
+ {"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.908}, "confidence": 0.9536, "noul": 0.0464}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 121.783}, "id": "1"}
71
  ```
72
 
73
  `noul` is the probability that the claim is true. Several questions in one request are answered together in one pass.
 
90
 
91
  | Test (questions) | Nitzotz | RoeiG (non-commercial) | laya-multilingual | Chance | Nitzotz vs RoeiG | Nitzotz vs laya-multilingual |
92
  |---|---|---|---|---|---|---|
93
+ | Scam or not? (298 messages) | **92.0%** | 32.2% | 50.3% | 50.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
94
+ | &nbsp;&nbsp;&nbsp;same, hard cases only (65) | **83.1%** | 32.3% | 35.4% | 50.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
95
+ | Message type, 6 options (298) | **77.8%** | 52.7% | 20.8% | 16.7% | **better**<br>p<0.001 | **better**<br>p<0.001 |
96
+ | &nbsp;&nbsp;&nbsp;same, hard cases only (65) | **67.7%** | 44.6% | 24.6% | 16.7% | **better**<br>p=0.01 | **better**<br>p<0.001 |
97
+ | Support ticket type, 5 options (30) | **90.0%** | 80.0% | 56.7% | 20.0% | tie<br>p=0.45 | **better**<br>p=0.01 |
98
+ | Ticket urgency, 5 levels (30) | **70.0%** | 36.7% | 33.3% | 20.0% | **better**<br>p=0.03 | **better**<br>p=0.02 |
99
+ | Paying customer? yes/no (30) | **80.0%** | 56.7% | 50.0% | 50.0% | **better**<br>p=0.02 | tie<br>p=0.06 |
100
  | Voice command intent, 20 options (500) | **90.0%** | 72.6% | 47.4% | 5.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
101
+ | Voice command intent, 4 options (500) | **97.2%** | 90.8% | 69.4% | 25.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
102
+ | News topic, 7 options (204) | 79.4% | **82.3%** | 66.2% | 14.3% | tie<br>p=0.36 | **better**<br>p=0.001 |
103
+ | Does the passage support this answer? (600) | **94.8%** | 94.2% | 48.8% | 50.0% | tie<br>p=0.69 | **better**<br>p<0.001 |
104
+ | Plausible answer the passage does not give (600) | **89.2%** | 53.0% | 53.7% | 50.0% | **better**<br>p<0.001 | **better**<br>p<0.001 |
105
+ | Reading comprehension, 4 options (900) | 63.2% | **75.4%** | 31.4% | 25.0% | **worse**<br>p<0.001 | **better**<br>p<0.001 |
106
 
107
+ On Belebele reading comprehension Nitzotz scores 63.2%, below RoeiG/laya-hebrew (75.4%); Nitzotz is built for
108
  message decisions, not long-passage comprehension.
109
 
110
  How to read it:
 
124
  how many were right: the chart above does that for every confidence level, over 3,990 test questions.
125
  Where the dots sit above the line, Nitzotz is more often right than it claims (it is modest); where they sit below, it
126
  is over-confident. This is what lets you set thresholds (see the next section). The per-question-type temperatures
127
+ were fitted on 4,825 held-out training items, never on the test sets (`calibration.json`).
128
 
129
  **It reads the text.** With the message removed and only the question left, its accuracy falls to
130
  70.5% on the scam question and 7.7% on message type. So the answers come from the message, not
131
  from the wording of the question.
132
 
133
+ **Rewording the question rarely changes the answer.** We asked 1,786 test questions in 6 different wordings with the same meaning. On 5.0% of them the answer was not the same in all 6. No wording makes it give one fixed answer to every item of a test. The exception is the urgency question (see [Limitations](#limitations)).
134
 
135
  ![Speed against file size](assets/speed_en.png)
136
 
 
138
 
139
  | Runtime | Median over 50 check questions (about 135 tokens) | Short message (38 tokens) | Long input (430 tokens) |
140
  |---|---|---|---|
141
+ | laya.exe, Q8 file, GPU (Vulkan) | 48 ms | 37 ms | 127 ms |
142
+ | laya.exe, F16 file, GPU (Vulkan) | 50 ms | 43 ms | 111 ms |
143
+ | Python (laya), GPU (PyTorch XPU) | 60 ms | 44 ms | 139 ms |
144
+ | laya.exe, Q8 file, CPU only (16 threads) | 309 ms | 160 ms | 1137 ms |
145
+ | Python (laya), CPU only | 224 ms | 112 ms | 1022 ms |
146
 
147
  The short and long columns repeat one fixed input 30 times after 5 warm-up calls. Timings on this laptop change a
148
  lot from one session to another (an earlier measurement of the same setup was several times slower), so treat these
 
153
 
154
  | File, device | Same top answer, 50 questions | Largest probability gap | Same top answer, 596 spam questions | Largest gap | Average gap |
155
  |---|---|---|---|---|---|
156
+ | Q8, GPU | 50/50 | 0.0375 | 596/596 | 0.0166 | 0.00130 |
157
+ | F16, GPU | 50/50 | 0.0017 | 596/596 | 0.0015 | 0.00020 |
158
+ | Q8, CPU | 50/50 | 0.0549 | 595/596 | 0.0207 | 0.00192 |
159
+ | F16, CPU | 50/50 | 0.0032 | 596/596 | 0.0012 | 0.00018 |
160
 
161
  A gap of 0.01 means, for example, 0.83 against 0.84. The few questions where the top answer changes are ones where
162
+ the two best answers were almost tied. Over both spam questions the Q8 file on the GPU is right 84.7% of the time,
163
+ against 84.7% for the Python model; the F16 file is closer (84.7%).
164
 
165
  ## Use it in your business
166
 
 
173
  - **Yellow** (not sure): send it to a person, or to a large language model if you use one.
174
  - **Red** (confident it is a scam): block or quarantine it.
175
 
176
+ The drawing is a worked example on the 298 test messages. The upper threshold, 0.49, is the one chosen for the
177
+ scam question on 1188 held-out training messages (never on the test); the lower one, 0.35, was picked by hand. With them,
178
+ 207 messages go to green (13 of them are in fact scams), 9 to yellow (2 scams), and 82 to red (9 of them are
179
  in fact legitimate). So red should mean "quarantine and check", not "delete". Pick your own thresholds on a sample of
180
+ your own messages, and decide how many mistakes in green and red you can live with. At the 0.49 threshold alone,
181
+ the scam answer is right 92.3% of the time overall and 83.1% on the 65 hard cases (at 0.50: 92.0% and 83.1%).
182
 
183
  The model is cheap enough to run on every message. The person (or the LLM) only sees the yellow part. **Do not use it
184
  as the only line of defence for decisions that can hurt someone.**
 
189
 
190
  1. **Learning to read.** HalleluBERT-large was first trained to find the answer to a question inside a passage, on
191
  27,085 HeQ training questions (CC BY 4.0). Passages that overlapped a test passage were removed.
192
+ 2. **Learning to decide.** A laya decision head was put on top and the whole model was trained on 116,854 items
193
+ (112,029 for training, 4,825 held out to pick the best of 2 passes and to fit the temperatures).
194
 
195
  | Source | Items | Licence | How it was made |
196
  |---|---|---|---|
 
208
  | Reading, 4 options ("which is NOT", reworded answers, several sentences) | 6,825 | passages: FineWeb-2, ODC-By 1.0; questions: model output, project-owned | written by DeepSeek V4.1 Flash, checked by Gemma 4 31B with the passage and again without it; dropped if it could be answered without the passage |
209
  | Topic of a passage, 7 options | 4,000 | passages: FineWeb-2, ODC-By 1.0; labels: model output | labelled by DeepSeek V4.1 Flash and Gemma 4 31B |
210
  | Is the sender an existing paying customer (yes/no) | 575 | model output, project-owned; question wordings written with GPT | support messages written by DeepSeek V4.1 Flash, checked by Gemma 4 31B |
211
+ | Scam or not: warnings about scams, short scams without a link, and legitimate look-alikes | 3,746 | model output, project-owned (DeepSeek MIT; Gemma Apache-2.0); question wordings written with GPT | written by DeepSeek V4.1 Flash; kept only if both labellers agreed with the writer, except that a scam only DeepSeek recognised was kept with a softer label (70%) |
212
+ | Message type of the same messages | 1,901 | model output, project-owned | the type both labellers agreed on, only where it fits the scam decision |
213
+ | "None of the options": existing training items with one option changed | 4,484 | as the original item (MASSIVE and HeQ: CC BY 4.0; reading: FineWeb-2 passages, ODC-By 1.0) | the right answer removed and a "none of the options" choice added; in a third of them a wrong option was removed instead, so "none" is wrong there |
214
 
215
  - **Synthetic messages (35,000 items).** DeepSeek V4.1 Flash (MIT) wrote Israeli-style SMS, WhatsApp and email messages
216
  from a plan (intended label, topic, tone, varied fake phone numbers and links). DeepSeek and Gemma 4 31B
 
226
  DeepSeek and kept only if both labellers agreed with each other and with the writer. 4,000 passages labelled
227
  with a topic, 575 support messages for the paying-customer question, and 4,750 repeated HeQ items so that the
228
  HeQ skills keep their weight in the mix.
229
+ - **Scam patterns seen in real messages.** Real Hebrew messages showed two weak spots: genuine warnings about scams
230
+ (from banks, the police, companies) flagged as scams, and short scams without a link missed. So DeepSeek V4.1 Flash
231
+ wrote 3,746 more messages: 1,000 legitimate warnings about scams, 1,476 short scams
232
+ without a link (a small unpaid debt or toll, a "friend" with a new number, a gift, a fake payment confirmation, a
233
+ payment app), 389 pairs of a scam and a warning about that same scam, and 492 legitimate
234
+ messages that look like those scams. The keep rule changed for these: a message was kept if both labellers agreed
235
+ with the writer, and a scam that only DeepSeek recognised was also kept, with a softer label (70% instead of close
236
+ to 100%); 134 messages are of that kind. Messages that resembled one of the real messages we checked were
237
+ dropped before labelling, and the real messages themselves were never used for training. 134 existing
238
+ training rows (67 messages: a request from a "new number" or for a small debt, with no link) were
239
+ relabelled as scams after the labellers called them scams (with the 70% label where only DeepSeek did).
240
+ 4,484 "none of the options" items were made from existing MASSIVE, HeQ and reading training items:
241
+ in 2,990 the right answer was removed and a "none of the options" choice added, and in
242
+ 1,494 a wrong option was removed instead, so "none" is wrong there.
243
+ - **Written with GPT.** Part of the training data was written with OpenAI's GPT, through our ChatGPT subscription: 376 alternative wordings of the questions and 34 sets of alternative answer options, used in 35,189 training items (including all 575 paying-customer items), and 120 short scenario outlines from which DeepSeek V4.1 Flash wrote 6,000 messages. GPT wrote no message and no label. In total 37,849 of the 116,854 training items (32%) use text written with GPT. GPT output is covered by OpenAI's terms of use, not by an open licence.
244
  - **Open data.** HeQ v1.1 (CC BY 4.0; about half of its questions are on Geektime articles, shared by the HeQ authors under
245
  the same licence) and MASSIVE he-IL (CC BY 4.0), training splits only. FineWeb-2 Hebrew (ODC-By 1.0) passages for
246
  the reading and topic items.
 
264
  | SIB-200 news topic | 204 | CC BY-SA 4.0, used for testing only |
265
  | Belebele reading | 900 | CC BY-SA 4.0, used for testing only |
266
  | HeQ verify and unanswerable | 600 + 600 | HeQ v1.1 test split, CC BY 4.0 |
267
+ | "None of the options" (reported apart, see [Limitations](#limitations)) | 400 (200 where "none" is right, 200 where it is wrong) | MASSIVE (CC BY 4.0) and Belebele (CC BY-SA 4.0) test questions with a "none of the options" choice added, used for testing only |
268
 
269
  Caveats that change how much to trust the numbers:
270
  - **The spam and business test was written by an AI model and checked by an AI model, not by a person.** Real inboxes
271
  will look different. It was frozen after that review (4 labels changed, 2 messages removed).
272
+ - **Three training runs, three random seeds; this is one of them.** On the held-out training items the three are almost level (95.0% here against 95.2% and 95.2%). This one was picked because it did best on the 188 held-out messages of the added scam data (98.4% against 97.3% and 97.3%) and on the real messages we checked, not on the tests in this card. The three runs are close on the large tests: scam or not 92.0% here against 92.0% and 90.6%, news topic 79.4% against 82.8% and 80.4%, reading 63.2% against 60.9% and 63.1%. This run and each of the other two give the same answer on 90.1% to 90.6% of all test questions. On the 30-ticket tests they differ more: ticket urgency 70.0% here against 73.3% and 76.7%.
273
  - **The training labels come from two AI models.** Where both are wrong in the same way, Nitzotz learned their mistake.
274
  - HeQ's wrong answers in the test were picked by code, not checked by a person.
275
 
276
  ## Limitations
277
 
278
+ - **Reading comprehension of longer passages is limited.** Belebele: 63.2%, where a blind guess
279
  gets 25%, and still below the best other laya model on this test (see [Benchmarks](#benchmarks)). When the right
280
+ answer is written word for word in the passage it gets 77% (252 questions); when the answer is said in other
281
+ words it gets 58% (648 questions); on "which of these is NOT" questions 59% (158 questions). Do not ask it
282
  whether a long document supports a claim.
283
+ - **Checking an answer against a passage (HeQ) works.** 94.8% with the passage; with the passage removed it falls to 51.0%, a coin flip. So on this kind of question it really reads the passage.
284
  - **Numbers, dates, amounts and rules**: not trained and not measured. Compute them in code and pass the result in.
285
  - **Hard cases are still the weak spot**: scams written to look legitimate (a "supplier" changing bank details, the
286
  "CEO" asking for a transfer) and real messages that look like scams (a real bank alert with a link, a real
287
+ verification code). On the 65 hard cases the scam question is right 83.1% of the time (always answering "not a
288
+ scam" there would give 70.8%), with AUC 0.84. At the chosen threshold it still misses 5 of the 19 scams written to look
289
+ legitimate, and flags 6 of the 46 real messages that look like scams.
290
  - **Urgency is subjective, and its answer depends on the wording.** Even the two teacher models matched the intended
291
  urgency only about 62 to 64% of the time. When the urgency question is asked in other words, the answer changes on
292
+ 63% of the 30 test tickets.
293
  - **The in-house tests and much of the training data were written by AI models, not by people.** This includes the
294
  spam and business test, the support-ticket test and the FineWeb-2 reading questions. Real messages will look different.
295
  - **The support-ticket tests are small**: 30 tickets per question, so one ticket moves a score by 3.3 points.
 
298
  - **Hebrew only.** Not trained or tested on English or Arabic.
299
  - **Not a safety system on its own.** It makes mistakes in both directions. Keep a person in the loop for anything
300
  that can hurt someone.
301
+ - **When a "none of the options" choice is added, it picks it too often.** Measured on a separate frozen test of
302
+ 400 questions, MASSIVE and Belebele test questions rebuilt with a "none of the options" choice. When
303
+ "none" is the right answer, it picks it 76% of the time. When the right answer is in the list, it
304
+ still picks "none" on 31% of the questions and gets 61% of them right
305
+ (77% on voice commands, 45% on reading questions). With the
306
+ text removed it picks "none" almost every time (98%). If you offer such a choice, test it on your own
307
+ questions first.
308
+ - **Real messages: not measured on an independent set yet.** Training data was added for genuine warnings about scams
309
+ and for short scams without a link, the two weak spots real messages showed. A measurement on an independent set of
310
+ real messages is still pending, so this card gives no number for real messages.
311
 
312
  ## Licence and attribution
313
 
 
337
 
338
  ![ניצוץ: מודל החלטות בעברית](assets/banner_he.png)
339
 
340
+ ![דיוק בזיהוי הונאות 92.0%, 0.05 שניות לשאלה, 413 MB, Apache-2.0](assets/tiles_he.png)
341
 
342
  **מה זה.** ניצוץ קורא הודעה בעברית ועונה על שאלות שאתם מקלידים עליה: לבחור אחת מכמה אפשרויות, לתת ציון בסולם, או
343
  לענות כן או לא על טענה. על כל תשובה הוא נותן הסתברות שאפשר לסמוך עליה, כך שיודעים מתי הוא בטוח ומתי הוא מנחש. הוא לא
 
346
  **בשביל מה.** להחליט מה עושים עם הודעות נכנסות: האם זו הונאה, איזה סוג הודעה זו, לאיזו מחלקה להעביר, כמה זה דחוף. זה
347
  לא צ'אטבוט, והוא לא בנוי למסמכים ארוכים (ראו [מגבלות](#מגבלות)).
348
 
349
+ **במספרים.** על 298 הודעות בעברית הוא קובע נכון אם ההודעה היא הונאה ב-92.0% מהמקרים. בשביל
350
  פרופורציה: 70% מההודעות האלה הן לא הונאה, כך שמודל שתמיד עונה "לא הונאה" היה מקבל 70.5%. המספר
351
+ שחשוב יותר הוא הדירוג: הוא נותן להונאה אמיתית הסתברות גבוהה יותר מאשר להודעה תקינה ב-97% מהמקרים (AUC 0.97).
352
 
353
  ### לנסות
354
 
 
368
 
369
  <div dir="rtl">
370
 
371
+ הפלט: `0.8628`, ההסתברות שהטענה ("ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית")
372
  נכונה.
373
 
374
  **הניסוח של השאלה משנה.** זו בדיוק השאלה שבה השתמש מבחן ההונאות, והמספרים בכרטיס הזה הם עליה. בניסיונות שלנו ניסוח
 
388
 
389
  ```json
390
  {"status":"ready","model":"laya"}
391
+ {"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.908}, "confidence": 0.9536, "noul": 0.0464}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 121.783}, "id": "1"}
392
  ```
393
 
394
  <div dir="rtl">
 
412
 
413
  | מבחן (מספר שאלות) | ניצוץ | RoeiG (לא מסחרי) | laya-multilingual | ניחוש | ניצוץ מול RoeiG | ניצוץ מול laya-multilingual |
414
  |---|---|---|---|---|---|---|
415
+ | הונאה או לא? (298 הודעות) | **92.0%** | 32.2% | 50.3% | 50.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
416
+ | &nbsp;&nbsp;&nbsp;אותו דבר, רק המקרים הקשים (65) | **83.1%** | 32.3% | 35.4% | 50.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
417
+ | סוג ההודעה, 6 אפשרויות (298) | **77.8%** | 52.7% | 20.8% | 16.7% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
418
+ | &nbsp;&nbsp;&nbsp;אותו דבר, רק המקרים הקשים (65) | **67.7%** | 44.6% | 24.6% | 16.7% | **טוב יותר**<br>p=0.01 | **טוב יותר**<br>p<0.001 |
419
+ | סוג פניית תמיכה, 5 אפשרויות (30) | **90.0%** | 80.0% | 56.7% | 20.0% | תיקו<br>p=0.45 | **טוב יותר**<br>p=0.01 |
420
+ | דחיפות הפנייה, 5 רמות (30) | **70.0%** | 36.7% | 33.3% | 20.0% | **טוב יותר**<br>p=0.03 | **טוב יותר**<br>p=0.02 |
421
+ | לקוח משלם? כן/לא (30) | **80.0%** | 56.7% | 50.0% | 50.0% | **טוב יותר**<br>p=0.02 | תיקו<br>p=0.06 |
422
  | כוונת פקודה קולית, 20 אפשרויות (500) | **90.0%** | 72.6% | 47.4% | 5.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
423
+ | כוונת פקודה קולית, 4 אפשרויות (500) | **97.2%** | 90.8% | 69.4% | 25.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
424
+ | נושא של ידיעה, 7 אפשרויות (204) | 79.4% | **82.3%** | 66.2% | 14.3% | תיקו<br>p=0.36 | **טוב יותר**<br>p=0.001 |
425
+ | האם ה��טע תומך בתשובה? (600) | **94.8%** | 94.2% | 48.8% | 50.0% | תיקו<br>p=0.69 | **טוב יותר**<br>p<0.001 |
426
+ | תשובה סבירה שהקטע לא נותן (600) | **89.2%** | 53.0% | 53.7% | 50.0% | **טוב יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
427
+ | הבנת הנקרא, 4 אפשרויות (900) | 63.2% | **75.4%** | 31.4% | 25.0% | **חלש יותר**<br>p<0.001 | **טוב יותר**<br>p<0.001 |
428
 
429
+ בהבנת הנקרא של Belebele ניצוץ מקבל 63.2%, פחות מ-RoeiG/laya-hebrew (75.4%). ניצוץ בנוי להחלטות על הודעות, לא
430
  להבנה של קטעים ארוכים.
431
 
432
  איך לקרוא את זה:
 
443
 
444
  **להסתברויות יש משמעות.** קחו את כל התשובות שבהן ניצוץ אמר שהוא בטוח בערך ב-70 עד 80%, וספרו כמה מהן היו נכונות.
445
  הגרף עושה את זה לכל רמת ביטחון, על 3,990 שאלות מבחן. כשהנקודות מעל הקו, ניצוץ צודק יותר ממה שהוא אומר (הוא
446
+ צנוע). כשהן מתחת, הוא בטוח בעצמו יותר מדי. זה מה שמאפשר לקבוע ספים (בפרק הבא). הכיול נעשה על 4,825 פריטי אימון
447
  שהופרדו מראש, אף פעם לא על המבחן (`calibration.json`).
448
 
449
  **הוא באמת קורא את הטקסט.** כשמוחקים את ההודעה ומשאירים רק את השאלה, הדיוק יורד ל-70.5% בשאלת ההונאה
450
  ול-7.7% בסוג ההודעה. כלומר התשובות באות מההודעה, לא מהניסוח של השאלה.
451
 
452
+ **ניסוח אחר של השאלה כמעט לא משנה את התשובה.** שאלנו 1,786 שאלות מבחן ב-6 ניסוחים שונים עם אותה משמעות. ב-5.0% מהן התשובה לא הייתה זהה בכל 6 הניסוחים. אף ניסוח לא גורם לו לתת תשובה קבועה אחת לכל הפריטים של מבחן. היוצא מן הכלל הוא שאלת הדחיפות (ראו [מגבלות](#מגבלות)).
453
 
454
  ![מהירות מול גודל](assets/speed_he.png)
455
 
 
457
 
458
  | איך מריצים | חציון על 50 שאלות בדיקה (בממוצע 135 טוקנים) | הודעה קצרה (38 טוקנים) | קלט ארוך (430 טוקנים) |
459
  |---|---|---|---|
460
+ | laya.exe, קובץ Q8, כרטיס מסך (Vulkan) | 48 ms | 37 ms | 127 ms |
461
+ | laya.exe, קובץ F16, כרטיס מסך (Vulkan) | 50 ms | 43 ms | 111 ms |
462
+ | פייתון (laya), כרטיס מסך (PyTorch XPU) | 60 ms | 44 ms | 139 ms |
463
+ | laya.exe, קובץ Q8, מעבד בלבד (16 תהליכונים) | 309 ms | 160 ms | 1137 ms |
464
+ | פייתון (laya), מעבד בלבד | 224 ms | 112 ms | 1022 ms |
465
 
466
  בעמודות של ההודעה הקצרה והקלט הארוך אותה שאלה רצה 30 פעמים, אחרי 5 הרצות חימום. הזמנים על המחשב הזה משתנים הרבה בין
467
  הפעלה להפעלה (מדידה קודמת של אותה הגדרה יצאה איטית פי כמה), אז אלה מספרים בקירוב.
 
470
 
471
  | קובץ, מכשיר | אותה תשובה מובילה, 50 שאלות | הפרש הסתברות מרבי | אותה תשובה מובילה, 596 שאלות ספאם | הפרש מרבי | הפרש ממוצע |
472
  |---|---|---|---|---|---|
473
+ | Q8, כרטיס מסך | 50/50 | 0.0375 | 596/596 | 0.0166 | 0.00130 |
474
+ | F16, כרטיס מסך | 50/50 | 0.0017 | 596/596 | 0.0015 | 0.00020 |
475
+ | Q8, מעבד | 50/50 | 0.0549 | 595/596 | 0.0207 | 0.00192 |
476
+ | F16, מעבד | 50/50 | 0.0032 | 596/596 | 0.0012 | 0.00018 |
477
 
478
  פער של 0.01 פירושו, למשל, 0.83 מול 0.84. השאלות המעטות שבהן התשובה המובילה משתנה הן כאלה שבהן שתי התשובות הטובות
479
+ היו כמעט שוות. בשתי שאלות הספאם יחד קובץ Q8 על כרטיס המסך צודק ב-84.7%, מול 84.7% למודל הפייתון. קובץ F16
480
+ קרוב יותר (84.7%).
481
 
482
  ### שימוש בעסק
483
 
 
489
  - **צהוב** (לא בטוח): לבדיקה של אדם, או של מודל שפה גדול אם אתם משתמשים בו.
490
  - **אדום** (בטוח שזו הונאה): חסימה או הסגר.
491
 
492
+ השרטוט הוא דוגמה על 298 הודעות המבחן. הסף העליון, 0.49, הוא הסף שנבחר לשאלת ההונאה על 1188 הודעות אימון
493
+ שהופרדו מראש (אף פעם לא על המבחן). הסף התחתון, 0.35, נבחר ביד. איתם 207 הודעות הולכות לירוק (13 מהן הן בעצם הונאה),
494
+ 9 לצהוב (2 הונאות), ו-82 לאדום (9 מהן בעצם תקינות). כלומר אדום צריך להיות "הסגר ובדיקה", לא "מחיקה".
495
+ בחרו ספים משלכם על מדגם של ההודעות שלכם, והחליטו כמה טעויות בירוק ובאדום אתם מוכנים לקבל. בסף 0.49 לבדו, תשובת
496
+ ההונאה נכונה ב-92.3% מהמקרים בסך הכול וב-83.1% על 65 המקרים הקשים (בסף 0.50: 92.0% ו-83.1%).
497
 
498
  המודל זול מספיק כדי להריץ אותו על כל הודעה. האדם (או מודל השפה) רואה רק את החלק הצהוב. **אל תשתמשו בו כקו הגנה
499
  יחיד בהחלטות שיכולות לפגוע במישהו.**
 
504
 
505
  1. **ללמוד לקרוא.** HalleluBERT-large אומן קודם למצוא את התשובה לשאלה בתוך קטע, על 27,085 שאלות אימון של HeQ
506
  (CC BY 4.0). קטעים שחפפו לקטע מבחן הוסרו.
507
+ 2. **ללמוד להחליט.** מעליו הונח ראש החלטות של laya, וכל המודל אומן על 116,854 פריטים (112,029 לאימון,
508
+ ו-4,825 הופרדו מראש כדי לבחור את הטוב מבין 2 מעברים ולכייל את הטמפרטורות).
509
 
510
  | מקור | פריטים | רישיון | איך נוצר |
511
  |---|---|---|---|
 
523
  | קריאה, 4 אפשרויות ("איזו מהבאות לא", תשובה במילים אחרות, כמה משפטים) | 6,825 | קטעים: FineWeb-2, ODC-By 1.0. שאלות: פלט מודלים, שייך לנו | נכתב בידי DeepSeek V4.1 Flash ונבדק בידי Gemma 4 31B עם הקטע ושוב בלי הקטע. נזרק אם אפשר היה לענות בלי הקטע |
524
  | נושא של קטע, 7 אפשרויות | 4,000 | קטעים: FineWeb-2, ODC-By 1.0. תוויות: פלט מודלים | סומן בידי DeepSeek V4.1 Flash ו-Gemma 4 31B |
525
  | האם השולח לקוח משלם קיים (כן/לא) | 575 | פלט מודלים, שייך לנו. ניסוחי השאלה נכתבו עם GPT | פניות תמיכה שנכתבו בידי DeepSeek V4.1 Flash ונבדקו בידי Gemma 4 31B |
526
+ | הונאה או לא: אזהרות מפני הונאה, הונאות קצרות בלי קישור, והודעות לגיטימיות שנראות כמוהן | 3,746 | פלט מודלים, שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0). ניסוחי השאלה נכתבו עם GPT | נכתב בידי DeepSeek V4.1 Flash. נשמר רק אם שני המסמנים הסכימו עם הכותב, חוץ מהונאה שרק DeepSeek זיהה, שנשמרה עם תווית רכה יותר (70%) |
527
+ | סוג ההודעה של אותן הודעות | 1,901 | פלט מודלים, שייך לנו | הסוג ששני המסמנים הסכימו עליו, רק כשהוא מתאים להחלטה אם זו הונאה |
528
+ | "אף אחת מהאפשרויות": פריטי אימון קיימים עם אפשרות אחת ששונתה | 4,484 | כמו הפריט המקורי (MASSIVE ו-HeQ: CC BY 4.0. קריאה: קטעים מ-FineWeb-2, ODC-By 1.0) | התשובה הנכונה הוסרה ונוספה אפשרות "אף אחת מהאפשרויות". בשליש מהם הוסרה במקום זה אפשרות שגויה, כך ש"אף אחת" שגויה שם |
529
 
530
  - **הודעות סינתטיות (35,000 פריטים).** DeepSeek V4.1 Flash (MIT) כתב הודעות SMS, וואטסאפ ומייל בסגנון ישראלי
531
  לפי תוכנית (התווית המתוכננת, נושא, משלב, מספרי טלפון וקישורים מזויפים ומגוונים). אחר כך DeepSeek ו-Gemma 4 31B
 
538
  הקטע נזרקה. 6,000 הודעות שקשה להבחין ביניהן (הונאות מוסוות, והודעות אמיתיות שנראות חשודות), שנכתבו בידי
539
  DeepSeek ונשמרו רק אם שני המסמנים הסכימו זה עם זה וגם עם הכותב. 4,000 קטעים שסומנו לפי נושא, 575 פניות
540
  תמיכה לשאלת הלקוח המשלם, ו-4,750 פריטי HeQ חוזרים, כדי שליכולות של HeQ יישאר משקל בתערובת.
541
+ - **דפוסי הונאה שנראו בהודעות אמיתיות.** הודעות אמיתיות בעברית הראו שתי נקודות חלשות: אזהרות אמיתיות מפני הונאה (מבנקים,
542
+ מהמשטרה, מחברות) שסומנו כהונאה, והונאות קצרות בלי קישור שפוספסו. לכן DeepSeek V4.1 Flash כתב עוד 3,746 הודעות:
543
+ 1,000 אזהרות לגיטימיות מפני הונאה, 1,476 הונאות קצרות בלי קישור (חוב קטן או אגרה שלא שולמו, "חבר" עם מספר חדש, מתנה,
544
+ אישור תשלום מזויף, אפליקציית תשלומים), 389 זוגות של הונאה ואזהרה על אותה הונאה, ו-492 הודעות לגיטימיות
545
+ שנראות כמו ההונאות האלה. כלל השמירה השתנה בהן: הודעה נשמרה אם שני המסמנים הסכימו עם הכותב, והונאה שרק DeepSeek זיהה
546
+ נשמרה גם כן, עם תווית רכה יותר (70% במקום קרוב ל-100%). 134 הודעות הן מהסוג הזה. הודעות שדמו לאחת מההודעות
547
+ האמיתיות שבדקנו נזרקו לפני הסימון, וההודעות האמיתיות עצמן לא שימשו לאימון. 134 שורות אימון קיימות
548
+ (67 הודעות: בקשה מ"מספר חדש" או על חוב קטן, בלי קישור) תויגו מחדש כהונאה אחרי שהמסמנים קבעו שהן הונאה
549
+ (עם תווית 70% כשרק DeepSeek קבע). 4,484 פריטי "אף אחת מהאפשרויות" נבנו מפריטי אימון קיימים של MASSIVE, HeQ
550
+ וקריאה: ב-2,990 מהם התשובה הנכונה הוסרה ונוספה אפשרות "אף אחת מהאפשרויות", וב-1,494 הוסרה במקום זה
551
+ אפשרות שגויה, כך ש"אף אחת" שגויה שם.
552
+ - **נכתב עם GPT.** חלק מנתוני האימון נכתב עם GPT של OpenAI, דרך מנוי ChatGPT שלנו: 376 ניסוחים חלופיים של השאלות ו-34 סטים חלופיים של אפשרויות תשובה, שמופיעים ב-35,189 פריטי אימון (כולל כל 575 הפריטים של שאלת הלקוח המשלם), ו-120 תרחישים קצרים שלפיהם DeepSeek V4.1 Flash כתב 6,000 הודעות. GPT לא כתב אף הודעה ואף תווית. בסך הכול 37,849 מתוך 116,854 פריטי האימון (32%) משתמשים בטקסט שנכתב עם GPT. הפלט של GPT כפוף לתנאי השימוש של OpenAI, לא לרישיון פתוח.
553
  - **נתונים פתוחים.** HeQ v1.1 (CC BY 4.0. בערך חצי מהשאלות שלו על כתבות של Geektime, ומחברי HeQ משתפים אותם באותו רישיון) ו-MASSIVE
554
  he-IL (CC BY 4.0), רק מפיצולי האימון. קטעים בעברית מ-FineWeb-2 (ODC-By 1.0) לשאלות הקריאה והנושא.
555
  - **בלי דליפה מהמבחנים.** כל פריט אימון הושווה לכל שאלת מבחן. כל מה שחולק רצף של 8 מילים עם טקסט מבחן נזרק, וגם
 
570
  | SIB-200, נושא ידיעה | 204 | CC BY-SA 4.0, רק לבדיקה |
571
  | Belebele, הבנת הנקרא | 900 | CC BY-SA 4.0, רק לבדיקה |
572
  | HeQ, אימות ו"אין תשובה" | 600 + 600 | פיצול המבחן של HeQ v1.1, CC BY 4.0 |
573
+ | "אף אחת מהאפשרויות" (מדווח בנפרד, ראו [מגבלות](#מגבלות)) | 400 (200 ש"אף אחת" נכונה בהן, 200 שהיא שגויה בהן) | שאלות מבחן של MASSIVE (CC BY 4.0) ו-Belebele (CC BY-SA 4.0) עם אפשרות "אף אחת מהאפשרויות" שנוספה, רק לבדיקה |
574
 
575
  הסתייגויות שמשנות כמה לסמוך על המספרים:
576
  - **מבחן הספאם והעסקים נכתב בידי מודל AI ונבדק בידי מודל AI, לא בידי אדם.** תיבות דואר אמיתיות ייראו אחרת. הוא ננעל
577
  אחרי הבדיקה הזו (4 תוויות שונו, 2 הודעות הוסרו).
578
+ - **שלוש ריצות אימון, עם שלושה זרעים אקראיים, וזו אחת מהן.** על פריטי האימון שהופרדו מראש שלוש הריצות כמעט שוות (95.0% כאן מול 95.2% ו-95.2%). הריצה הזו נבחרה כי היא הייתה הטובה ביותר על 188 ההודעות המוחזקות מנתוני ההונאה שנוספו (98.4% מול 97.3% ו-97.3%) ועל ההודעות האמיתיות שבדקנו, לא לפי המבחנים בכרטיס הזה. במבחנים הגדולים שלוש הריצות קרובות: הונאה או לא 92.0% כאן מול 92.0% ו-90.6%, נושא ידיעה 79.4% מול 82.8% ו-80.4%, הבנת הנקרא 63.2% מול 60.9% ו-63.1%. הריצה הזו וכל אחת מהשתיים האחרות נותנות אותה תשובה על 90.1% עד 90.6% מכל שאלות המבחן. במבחנים של 30 פניות ההבדלים גדולים יותר: דחיפות הפנייה 70.0% כאן מול 73.3% ו-76.7%.
579
  - **תוויות האימון באות משני מודלי AI.** איפה ששניהם טועים באותו אופן, ניצוץ למד את הטעות שלהם.
580
  - התשובות השגויות של HeQ במבחן נבחרו בקוד, ולא נבדקו בידי אדם.
581
 
582
  ### מגבלות
583
 
584
+ - **הבנת הנקרא של קטעים ארוכים מוגבלת.** Belebele: 63.2%, כשניחוש עיוור מקבל 25%, ועדיין פחות ממודל
585
+ laya האחר הטוב ביותר במבחן הזה (ראו [מבחנים](#מבחנים)). כשהתשובה הנכונה כתובה בקטע מילה במילה הוא מקבל 77% (252 שאלות).
586
+ כשהתשובה נאמרת במילים אחרות הוא מקבל 58% (648 שאלות). בשאלות "איזו מהבאות לא" 59% (158 שאלות).
587
  אל תשאלו אותו אם מסמך ארוך תומך בטענה.
588
+ - **בדיקה אם קטע תומך בתשובה (HeQ) עובדת.** 94.8% עם הקטע. כשמוחקים את הקטע זה יורד ל-51.0%, הטלת מטבע. כלומר בשאלות מהסוג הזה הוא באמת קורא את הקטע.
589
  - **מספרים, תאריכים, סכומים וכללים**: לא אומן עליהם ולא נמדד. חשבו אותם בקוד והעבירו את התוצאה.
590
  - **המקרים הקשים עדיין נקודת התורפה**: הונאות שנכתבו כדי להיראות לגיטימיות ("ספק" שמחליף פרטי חשבון בנק, "המנכ"ל"
591
  שמבקש העברה), והודעות אמיתיות שנראות כמו הונאה (התראה אמיתית מהבנק עם קישור, קוד אימות אמיתי). על 65 המקרים הקשים
592
+ שאלת ההונאה צודקת ב-83.1% מהמקרים (תשובה קבועה "לא הונאה" הייתה מקבלת שם 70.8%), עם AUC 0.84. בסף שנבחר הוא
593
+ עדיין מפספס 5 מתוך 19 הונאות שנכתבו כדי להיראות לגיטימיות, ומתריע על 6 מתוך 46 הודעות אמיתיות שנראות כמו הונאה.
594
  - **דחיפות היא עניין סובייקטיבי, והתשובה עליה תלויה בניסוח.** אפילו שני מודלי המורה הסכימו עם הדחיפות המתוכננת רק בכ-62
595
+ עד 64% מהמקרים. כששואלים את שאלת הדחיפות במילים אחרות, התשובה משתנה ב-63% מתוך 30 פניות המבחן.
596
  - **המבחנים הפנימיים וחלק גדול מנתוני האימון נכתבו בידי מודלי AI, לא בידי אנשים.** זה כולל את מבחן הספאם והעסקים, את
597
  מבחן פניות התמיכה ואת שאלות הקריאה על FineWeb-2. הודעות אמיתיות ייראו אחרת.
598
  - **מבחני פניות התמיכה קטנים**: 30 פניות לכל שאלה, כך שפנייה אחת מזיזה ציון ב-3.3 נקודות.
 
600
  - **512 טוקנים** (בערך 300 עד 400 מילים בעברית) לשאלה. הודעה ארוכה יותר נחתכת מהסוף בלי אזהרה.
601
  - **עברית בלבד.** לא אומן ולא נבדק על אנגלית או ערבית.
602
  - **זו לא מערכת הגנה לבד.** הוא טועה לשני הכיוונים. השאירו אדם בתהליך בכל דבר שיכול לפגוע במישהו.
603
+ - **כשמוסיפים אפשרות "אף אחת מהאפשרויות", הוא בוחר בה יותר מדי.** נמדד על מבחן קפוא נפרד של 400 שאלות, שאלות מבחן
604
+ של MASSIVE ו-Belebele שנבנו מחדש עם אפשרות "אף אחת מהאפשרויות". כש"אף אחת" היא התשובה הנכונה, הוא בוחר בה ב-76%
605
+ מהמקרים. כשהתשובה הנכונה נמצאת ברשימה, הוא עדיין בוחר "אף אחת" ב-31% מהשאלות, וצודק רק ב-61% מהן
606
+ (77% בפקודות קוליות, 45% בשאלות קריאה). כשמוחקים את הטקסט הוא בוחר "אף אחת" כמעט תמיד
607
+ (98%). אם אתם מציעים אפשרות כזו, בדקו אותה קודם על השאלות שלכם.
608
+ - **הודעות אמיתיות: עוד לא נמדד על סט בלתי תלוי.** נוספו נתוני אימון לאזהרות אמיתיות מפני הונאה ולהונאות קצרות בלי קישור,
609
+ שתי הנקודות החלשות שהודעות אמיתיות הראו. מדידה על סט בלתי תלוי של הודעות אמיתיות עוד לא נעשתה, ולכן בכרטיס הזה אין
610
+ מספר על הודעות אמיתיות.
611
 
612
  ### רישיון וקרדיטים
613
 
assets/banner_en.png CHANGED
assets/banner_he.png CHANGED
assets/bars_en.png CHANGED

Git LFS Details

  • SHA256: 2352043ce0194ed9d1b3bd2979a387bd989331aa3fd3ef4056dcc0feab2f59a2
  • Pointer size: 131 Bytes
  • Size of remote file: 112 kB

Git LFS Details

  • SHA256: ac2b479849940f05b4caefd9c517f97acb1c9045a9eeb3ee5f66aa76892a893a
  • Pointer size: 131 Bytes
  • Size of remote file: 112 kB
assets/bars_he.png CHANGED

Git LFS Details

  • SHA256: de7b8cea8285ac3bec810277e8357becbb4cb6431c92f22b6edbacebf118d6cf
  • Pointer size: 130 Bytes
  • Size of remote file: 84.7 kB

Git LFS Details

  • SHA256: f28b55b2771164514e5e2cfa51eafc99658f1648bc15d7e4cac335ad8d8a1fcd
  • Pointer size: 130 Bytes
  • Size of remote file: 85.4 kB
assets/calibration_en.png CHANGED

Git LFS Details

  • SHA256: 0f13b39b3b57c8e30e52d20b64a6909c6228a97bf94142cdc6967087991fa497
  • Pointer size: 131 Bytes
  • Size of remote file: 109 kB

Git LFS Details

  • SHA256: fdf90fac118adbaa5111e124e99d9b7c84058596b917b26d94105d88401c9ee8
  • Pointer size: 131 Bytes
  • Size of remote file: 107 kB
assets/calibration_he.png CHANGED

Git LFS Details

  • SHA256: bce20a494d210414b792a388bbd2ee0e23261097199805241d15c912a8358dd7
  • Pointer size: 130 Bytes
  • Size of remote file: 88.1 kB

Git LFS Details

  • SHA256: e3f84a6b305bef0ae6b1d7f1f7255c4b63693e8fc21e55e14c287aad49233270
  • Pointer size: 130 Bytes
  • Size of remote file: 87.1 kB
assets/radar_en.png CHANGED

Git LFS Details

  • SHA256: a2424988ab51e8e1029025d37ff2748e3e6295cd3bbe95a3f681ff2d2e141706
  • Pointer size: 131 Bytes
  • Size of remote file: 247 kB

Git LFS Details

  • SHA256: 390f345c20cf3e99f932cc0d874728fc5b60f5ac252d59c0fec691ad56aa9b9b
  • Pointer size: 131 Bytes
  • Size of remote file: 249 kB
assets/radar_he.png CHANGED

Git LFS Details

  • SHA256: a9816bd7bd293cad1a1cf353fddfd005510dcb5ab7fe7ca07286b6fa52fb9760
  • Pointer size: 131 Bytes
  • Size of remote file: 232 kB

Git LFS Details

  • SHA256: 62125389ba919e61cfd882f0a4f6c914bb87ae0ae720097a0483fbd821eb69c3
  • Pointer size: 131 Bytes
  • Size of remote file: 235 kB
assets/ramzor_en.svg CHANGED
assets/ramzor_he.svg CHANGED
assets/speed_en.png CHANGED

Git LFS Details

  • SHA256: 664333364d9612a637a9fae54208a0073d6eb717cb1176a3f6eefc67a8b343fb
  • Pointer size: 130 Bytes
  • Size of remote file: 89.3 kB

Git LFS Details

  • SHA256: d9e70a459be2a18d6c44af387b9f5033845f5068b9fdb1cbf716655fddfcc23d
  • Pointer size: 130 Bytes
  • Size of remote file: 88.9 kB
assets/speed_he.png CHANGED
assets/tiles_en.png CHANGED
assets/tiles_he.png CHANGED
calibration.json CHANGED
@@ -1,57 +1,61 @@
1
  {
2
  "temperatures_per_question_type": {
3
- "choice": 1.1032,
4
- "score": 1.1106,
5
- "noul": 1.196
6
  },
7
  "temperature_fit_items": {
8
- "choice": 2446,
9
  "score": 261,
10
- "noul": 1843
11
  },
12
  "note": "The temperatures are already applied by laya (rl_agent_config.json 'temperature') and baked into the GGUF files. They were fitted on held-out training items, never on the test sets.",
13
  "scam_threshold": {
14
  "question": "noul, 'is this message a scam'",
15
- "threshold": 0.5151,
16
- "chosen_on": "1000 held-out training messages, never the test",
17
- "held_out_acc_at_threshold": 0.98,
18
  "held_out_pools": {
 
 
 
 
19
  "held-out hard scam and legitimate messages": {
20
  "n": 300,
21
- "acc_at_threshold": 0.9967
22
  },
23
  "held-out spam messages": {
24
  "n": 700,
25
- "acc_at_threshold": 0.9729
26
  }
27
  },
28
- "test_acc_at_0.5": 0.9128,
29
- "test_acc_at_threshold": 0.9128,
30
- "test_hard_acc_at_0.5": 0.8,
31
- "test_hard_acc_at_threshold": 0.8,
32
- "test_auc": 0.9761
33
  },
34
  "ramzor_example_thresholds": {
35
  "green_below": 0.35,
36
- "red_from": 0.52,
37
  "zones_on_test": {
38
  "g": {
39
  "n": 207,
40
  "share": 0.6946308724832215,
41
- "scam": 11,
42
- "legit": 196
43
  },
44
  "y": {
45
- "n": 5,
46
- "share": 0.016778523489932886,
47
- "scam": 3,
48
- "legit": 2
49
  },
50
  "r": {
51
- "n": 86,
52
- "share": 0.28859060402684567,
53
- "scam": 74,
54
- "legit": 12
55
  }
56
  }
57
  }
 
1
  {
2
  "temperatures_per_question_type": {
3
+ "choice": 1.0971,
4
+ "score": 1.139,
5
+ "noul": 1.162
6
  },
7
  "temperature_fit_items": {
8
+ "choice": 2533,
9
  "score": 261,
10
+ "noul": 2031
11
  },
12
  "note": "The temperatures are already applied by laya (rl_agent_config.json 'temperature') and baked into the GGUF files. They were fitted on held-out training items, never on the test sets.",
13
  "scam_threshold": {
14
  "question": "noul, 'is this message a scam'",
15
+ "threshold": 0.4861,
16
+ "chosen_on": "1188 held-out training messages, never the test",
17
+ "held_out_acc_at_threshold": 0.979,
18
  "held_out_pools": {
19
+ "held-out warnings about scams, short scams without a link and look-alike messages": {
20
+ "n": 188,
21
+ "acc_at_threshold": 0.984
22
+ },
23
  "held-out hard scam and legitimate messages": {
24
  "n": 300,
25
+ "acc_at_threshold": 0.9933
26
  },
27
  "held-out spam messages": {
28
  "n": 700,
29
+ "acc_at_threshold": 0.9714
30
  }
31
  },
32
+ "test_acc_at_0.5": 0.9195,
33
+ "test_acc_at_threshold": 0.9228,
34
+ "test_hard_acc_at_0.5": 0.8308,
35
+ "test_hard_acc_at_threshold": 0.8308,
36
+ "test_auc": 0.969
37
  },
38
  "ramzor_example_thresholds": {
39
  "green_below": 0.35,
40
+ "red_from": 0.49,
41
  "zones_on_test": {
42
  "g": {
43
  "n": 207,
44
  "share": 0.6946308724832215,
45
+ "scam": 13,
46
+ "legit": 194
47
  },
48
  "y": {
49
+ "n": 9,
50
+ "share": 0.030201342281879196,
51
+ "scam": 2,
52
+ "legit": 7
53
  },
54
  "r": {
55
+ "n": 82,
56
+ "share": 0.2751677852348993,
57
+ "scam": 73,
58
+ "legit": 9
59
  }
60
  }
61
  }
eval_results.json CHANGED
@@ -17,17 +17,17 @@
17
  "n": 298,
18
  "chance": 0.5,
19
  "acc": {
20
- "nz": 0.9128,
21
  "roeig": 0.3221,
22
  "laya_ml": 0.5034
23
  },
24
  "brier": {
25
- "nz": 0.126,
26
  "roeig": 0.7604,
27
  "laya_ml": 0.8042
28
  },
29
  "ece": {
30
- "nz": 0.0548,
31
  "roeig": 0.3802,
32
  "laya_ml": 0.38
33
  },
@@ -39,13 +39,13 @@
39
  "vs_nitzotz": {
40
  "roeig": {
41
  "nz_only": 186,
42
- "other_only": 10,
43
- "p": 3.8398295923083015e-43
44
  },
45
  "laya_ml": {
46
- "nz_only": 136,
47
  "other_only": 14,
48
- "p": 2.7927572561866964e-26
49
  }
50
  }
51
  },
@@ -54,17 +54,17 @@
54
  "n": 65,
55
  "chance": 0.5,
56
  "acc": {
57
- "nz": 0.8,
58
  "roeig": 0.3231,
59
  "laya_ml": 0.3538
60
  },
61
  "brier": {
62
- "nz": 0.2474,
63
  "roeig": 0.7076,
64
  "laya_ml": 1.1057
65
  },
66
  "ece": {
67
- "nz": 0.0959,
68
  "roeig": 0.3391,
69
  "laya_ml": 0.5564
70
  },
@@ -76,13 +76,13 @@
76
  "vs_nitzotz": {
77
  "roeig": {
78
  "nz_only": 34,
79
- "other_only": 3,
80
- "p": 1.233129296451807e-07
81
  },
82
  "laya_ml": {
83
- "nz_only": 34,
84
- "other_only": 5,
85
- "p": 2.4299079086631536e-06
86
  }
87
  }
88
  },
@@ -91,17 +91,17 @@
91
  "n": 298,
92
  "chance": 0.1667,
93
  "acc": {
94
- "nz": 0.7617,
95
  "roeig": 0.5268,
96
  "laya_ml": 0.2081
97
  },
98
  "brier": {
99
- "nz": 0.3177,
100
  "roeig": 0.6489,
101
  "laya_ml": 1.0746
102
  },
103
  "ece": {
104
- "nz": 0.0872,
105
  "roeig": 0.1043,
106
  "laya_ml": 0.4131
107
  },
@@ -112,14 +112,14 @@
112
  },
113
  "vs_nitzotz": {
114
  "roeig": {
115
- "nz_only": 104,
116
- "other_only": 34,
117
- "p": 1.9152268319067766e-09
118
  },
119
  "laya_ml": {
120
- "nz_only": 174,
121
- "other_only": 9,
122
- "p": 8.929600498958945e-41
123
  }
124
  }
125
  },
@@ -128,17 +128,17 @@
128
  "n": 65,
129
  "chance": 0.1667,
130
  "acc": {
131
- "nz": 0.6308,
132
  "roeig": 0.4462,
133
  "laya_ml": 0.2462
134
  },
135
  "brier": {
136
- "nz": 0.4714,
137
  "roeig": 0.7315,
138
  "laya_ml": 1.1094
139
  },
140
  "ece": {
141
- "nz": 0.0653,
142
  "roeig": 0.1877,
143
  "laya_ml": 0.4109
144
  },
@@ -149,14 +149,14 @@
149
  },
150
  "vs_nitzotz": {
151
  "roeig": {
152
- "nz_only": 22,
153
- "other_only": 10,
154
- "p": 0.050102459732443094
155
  },
156
  "laya_ml": {
157
- "nz_only": 29,
158
- "other_only": 4,
159
- "p": 1.0928604751825333e-05
160
  }
161
  }
162
  },
@@ -165,17 +165,17 @@
165
  "n": 30,
166
  "chance": 0.2,
167
  "acc": {
168
- "nz": 0.8667,
169
  "roeig": 0.8,
170
  "laya_ml": 0.5667
171
  },
172
  "brier": {
173
- "nz": 0.2164,
174
  "roeig": 0.3191,
175
  "laya_ml": 0.6667
176
  },
177
  "ece": {
178
- "nz": 0.1364,
179
  "roeig": 0.1321,
180
  "laya_ml": 0.2772
181
  },
@@ -186,14 +186,14 @@
186
  },
187
  "vs_nitzotz": {
188
  "roeig": {
189
- "nz_only": 4,
190
  "other_only": 2,
191
- "p": 0.6875
192
  },
193
  "laya_ml": {
194
- "nz_only": 11,
195
  "other_only": 2,
196
- "p": 0.0224609375
197
  }
198
  }
199
  },
@@ -202,35 +202,35 @@
202
  "n": 30,
203
  "chance": 0.2,
204
  "acc": {
205
- "nz": 0.8,
206
  "roeig": 0.3667,
207
  "laya_ml": 0.3333
208
  },
209
  "brier": {
210
- "nz": 0.4576,
211
  "roeig": 0.708,
212
  "laya_ml": 0.7501
213
  },
214
  "ece": {
215
- "nz": 0.3313,
216
  "roeig": 0.0322,
217
  "laya_ml": 0.2167
218
  },
219
  "empty_acc": {
220
- "nz": 0.1,
221
  "roeig": 0.2333,
222
  "laya_ml": 0.0667
223
  },
224
  "vs_nitzotz": {
225
  "roeig": {
226
- "nz_only": 15,
227
- "other_only": 2,
228
- "p": 0.002349853515625
229
  },
230
  "laya_ml": {
231
- "nz_only": 17,
232
- "other_only": 3,
233
- "p": 0.0025768280029296875
234
  }
235
  }
236
  },
@@ -239,17 +239,17 @@
239
  "n": 30,
240
  "chance": 0.5,
241
  "acc": {
242
- "nz": 0.7667,
243
  "roeig": 0.5667,
244
  "laya_ml": 0.5
245
  },
246
  "brier": {
247
- "nz": 0.4053,
248
  "roeig": 0.4338,
249
  "laya_ml": 0.8183
250
  },
251
  "ece": {
252
- "nz": 0.2204,
253
  "roeig": 0.2362,
254
  "laya_ml": 0.4284
255
  },
@@ -260,14 +260,14 @@
260
  },
261
  "vs_nitzotz": {
262
  "roeig": {
263
- "nz_only": 6,
264
  "other_only": 0,
265
- "p": 0.03125
266
  },
267
  "laya_ml": {
268
  "nz_only": 14,
269
- "other_only": 6,
270
- "p": 0.11531829833984375
271
  }
272
  }
273
  },
@@ -281,30 +281,30 @@
281
  "laya_ml": 0.474
282
  },
283
  "brier": {
284
- "nz": 0.1701,
285
  "roeig": 0.3801,
286
  "laya_ml": 0.7338
287
  },
288
  "ece": {
289
- "nz": 0.0412,
290
  "roeig": 0.0395,
291
  "laya_ml": 0.1661
292
  },
293
  "empty_acc": {
294
- "nz": 0.058,
295
  "roeig": 0.052,
296
  "laya_ml": 0.044
297
  },
298
  "vs_nitzotz": {
299
  "roeig": {
300
- "nz_only": 100,
301
- "other_only": 13,
302
- "p": 8.471157872064744e-18
303
  },
304
  "laya_ml": {
305
- "nz_only": 220,
306
- "other_only": 7,
307
- "p": 5.374231848579288e-56
308
  }
309
  }
310
  },
@@ -313,35 +313,35 @@
313
  "n": 500,
314
  "chance": 0.25,
315
  "acc": {
316
- "nz": 0.97,
317
  "roeig": 0.908,
318
  "laya_ml": 0.694
319
  },
320
  "brier": {
321
- "nz": 0.0514,
322
  "roeig": 0.1369,
323
  "laya_ml": 0.3987
324
  },
325
  "ece": {
326
- "nz": 0.0443,
327
  "roeig": 0.0393,
328
  "laya_ml": 0.0468
329
  },
330
  "empty_acc": {
331
- "nz": 0.24,
332
  "roeig": 0.242,
333
  "laya_ml": 0.256
334
  },
335
  "vs_nitzotz": {
336
  "roeig": {
337
- "nz_only": 37,
338
- "other_only": 6,
339
- "p": 1.636124125070637e-06
340
  },
341
  "laya_ml": {
342
- "nz_only": 142,
343
- "other_only": 4,
344
- "p": 4.188799933293481e-37
345
  }
346
  }
347
  },
@@ -350,17 +350,17 @@
350
  "n": 204,
351
  "chance": 0.1429,
352
  "acc": {
353
- "nz": 0.8039,
354
  "roeig": 0.8235,
355
  "laya_ml": 0.6618
356
  },
357
  "brier": {
358
- "nz": 0.3002,
359
  "roeig": 0.285,
360
  "laya_ml": 0.4918
361
  },
362
  "ece": {
363
- "nz": 0.0857,
364
  "roeig": 0.0751,
365
  "laya_ml": 0.1784
366
  },
@@ -371,14 +371,14 @@
371
  },
372
  "vs_nitzotz": {
373
  "roeig": {
374
- "nz_only": 15,
375
- "other_only": 19,
376
- "p": 0.6075913612730801
377
  },
378
  "laya_ml": {
379
  "nz_only": 47,
380
- "other_only": 18,
381
- "p": 0.0004221303234752052
382
  }
383
  }
384
  },
@@ -392,30 +392,30 @@
392
  "laya_ml": 0.4883
393
  },
394
  "brier": {
395
- "nz": 0.0899,
396
  "roeig": 0.0926,
397
  "laya_ml": 0.8274
398
  },
399
  "ece": {
400
- "nz": 0.0484,
401
  "roeig": 0.0178,
402
  "laya_ml": 0.3946
403
  },
404
  "empty_acc": {
405
- "nz": 0.5033,
406
  "roeig": 0.8,
407
  "laya_ml": 0.59
408
  },
409
  "vs_nitzotz": {
410
  "roeig": {
411
- "nz_only": 33,
412
- "other_only": 29,
413
- "p": 0.7035366713552734
414
  },
415
  "laya_ml": {
416
- "nz_only": 299,
417
- "other_only": 23,
418
- "p": 2.1018222751120708e-62
419
  }
420
  }
421
  },
@@ -424,35 +424,35 @@
424
  "n": 600,
425
  "chance": 0.5,
426
  "acc": {
427
- "nz": 0.9067,
428
  "roeig": 0.53,
429
  "laya_ml": 0.5367
430
  },
431
  "brier": {
432
- "nz": 0.159,
433
  "roeig": 0.856,
434
  "laya_ml": 0.7348
435
  },
436
  "ece": {
437
- "nz": 0.0185,
438
  "roeig": 0.4336,
439
  "laya_ml": 0.3421
440
  },
441
  "empty_acc": {
442
- "nz": 0.505,
443
  "roeig": 0.5033,
444
  "laya_ml": 0.5167
445
  },
446
  "vs_nitzotz": {
447
  "roeig": {
448
- "nz_only": 249,
449
- "other_only": 23,
450
- "p": 4.261733883444211e-49
451
  },
452
  "laya_ml": {
453
- "nz_only": 244,
454
- "other_only": 22,
455
- "p": 1.5017318497531166e-48
456
  }
457
  }
458
  },
@@ -461,35 +461,35 @@
461
  "n": 900,
462
  "chance": 0.25,
463
  "acc": {
464
- "nz": 0.6267,
465
  "roeig": 0.7544,
466
  "laya_ml": 0.3144
467
  },
468
  "brier": {
469
- "nz": 0.5567,
470
  "roeig": 0.3442,
471
  "laya_ml": 0.8341
472
  },
473
  "ece": {
474
- "nz": 0.1917,
475
  "roeig": 0.0324,
476
  "laya_ml": 0.2307
477
  },
478
  "empty_acc": {
479
- "nz": 0.3133,
480
  "roeig": 0.3189,
481
  "laya_ml": 0.2844
482
  },
483
  "vs_nitzotz": {
484
  "roeig": {
485
- "nz_only": 66,
486
- "other_only": 181,
487
- "p": 1.4586191292115705e-13
488
  },
489
  "laya_ml": {
490
- "nz_only": 365,
491
- "other_only": 84,
492
- "p": 8.256351793141917e-43
493
  }
494
  }
495
  }
@@ -497,33 +497,33 @@
497
  "significance": "paired exact McNemar against Nitzotz on the same questions",
498
  "latency_ms": {
499
  "q8": {
500
- "median": 50.2,
501
- "p90": 76.5,
502
  "n": 50
503
  },
504
  "f16": {
505
- "median": 52.4,
506
- "p90": 79.5,
507
  "n": 50
508
  },
509
  "py": {
510
- "median": 70.2,
511
- "p90": 109.6,
512
  "n": 50
513
  },
514
  "q8_cpu": {
515
- "median": 339.5,
516
- "p90": 710.2,
517
  "n": 50
518
  },
519
  "f16_cpu": {
520
- "median": 491.5,
521
- "p90": 988.3,
522
  "n": 50
523
  },
524
  "py_cpu": {
525
- "median": 342.0,
526
- "p90": 666.5,
527
  "n": 50
528
  }
529
  },
@@ -532,32 +532,32 @@
532
  "short": {
533
  "item": "massive_he_4-0075",
534
  "input_tokens": 38,
535
- "median_ms": 38.1,
536
- "p90_ms": 39.6,
537
- "min_ms": 36.5
538
  },
539
  "long": {
540
  "item": "heq_unanswerable-0522",
541
  "input_tokens": 430,
542
- "median_ms": 138.4,
543
- "p90_ms": 141.0,
544
- "min_ms": 130.6
545
  }
546
  },
547
  "f16": {
548
  "short": {
549
  "item": "massive_he_4-0075",
550
  "input_tokens": 38,
551
- "median_ms": 43.5,
552
- "p90_ms": 44.8,
553
- "min_ms": 41.6
554
  },
555
  "long": {
556
  "item": "heq_unanswerable-0522",
557
  "input_tokens": 430,
558
- "median_ms": 137.2,
559
- "p90_ms": 148.4,
560
- "min_ms": 113.0
561
  }
562
  },
563
  "py": {
@@ -565,117 +565,117 @@
565
  "item": "massive_he_4-0075",
566
  "input_tokens": 38,
567
  "median_ms": 43.6,
568
- "p90_ms": 46.9,
569
- "min_ms": 41.8
570
  },
571
  "long": {
572
  "item": "heq_unanswerable-0522",
573
  "input_tokens": 430,
574
- "median_ms": 192.0,
575
- "p90_ms": 210.9,
576
- "min_ms": 149.4
577
  }
578
  },
579
  "q8_cpu": {
580
  "short": {
581
  "item": "massive_he_4-0075",
582
  "input_tokens": 38,
583
- "median_ms": 186.0,
584
- "p90_ms": 199.5,
585
- "min_ms": 165.2
586
  },
587
  "long": {
588
  "item": "heq_unanswerable-0522",
589
  "input_tokens": 430,
590
- "median_ms": 1270.3,
591
- "p90_ms": 1352.6,
592
- "min_ms": 1138.8
593
  }
594
  },
595
  "py_cpu": {
596
  "short": {
597
  "item": "massive_he_4-0075",
598
  "input_tokens": 38,
599
- "median_ms": 134.3,
600
- "p90_ms": 147.2,
601
- "min_ms": 126.7
602
  },
603
  "long": {
604
  "item": "heq_unanswerable-0522",
605
  "input_tokens": 430,
606
- "median_ms": 1236.2,
607
- "p90_ms": 1317.7,
608
- "min_ms": 1186.6
609
  }
610
  }
611
  },
612
  "latency_machine": "Intel Core Ultra 9 285H laptop, built-in Arc 140T GPU, Windows 11",
613
  "gguf_parity_50": {
614
  "q8": {
615
- "max_abs_diff": 0.0456,
616
- "mean_abs_diff": 0.001536,
617
  "argmax_agree": "50/50"
618
  },
619
  "f16": {
620
- "max_abs_diff": 0.0031,
621
- "mean_abs_diff": 0.00032,
622
  "argmax_agree": "50/50"
623
  },
624
  "q8_cpu": {
625
- "max_abs_diff": 0.0563,
626
- "mean_abs_diff": 0.00241,
627
  "argmax_agree": "50/50"
628
  },
629
  "f16_cpu": {
630
- "max_abs_diff": 0.0015,
631
- "mean_abs_diff": 0.000242,
632
  "argmax_agree": "50/50"
633
  }
634
  },
635
  "gguf_parity_spam_596": {
636
  "q8": {
637
- "max_abs_diff": 0.0189,
638
- "mean_abs_diff": 0.001162,
639
- "argmax_agree": "595/596",
640
- "accuracy_vs_gold": 0.8356
641
  },
642
  "f16": {
643
- "max_abs_diff": 0.0023,
644
- "mean_abs_diff": 0.000227,
645
  "argmax_agree": "596/596",
646
- "accuracy_vs_gold": 0.8372
647
  },
648
  "py": {
649
- "max_abs_diff": 0.0001,
650
  "mean_abs_diff": 0.0,
651
  "argmax_agree": "596/596",
652
- "accuracy_vs_gold": 0.8372
653
  },
654
  "q8_cpu": {
655
- "max_abs_diff": 0.0278,
656
- "mean_abs_diff": 0.001854,
657
  "argmax_agree": "595/596",
658
- "accuracy_vs_gold": 0.8356
659
  },
660
  "f16_cpu": {
661
- "max_abs_diff": 0.0017,
662
- "mean_abs_diff": 0.000203,
663
  "argmax_agree": "596/596",
664
- "accuracy_vs_gold": 0.8372
665
  },
666
  "py_cpu": {
667
  "max_abs_diff": 0.0,
668
  "mean_abs_diff": 0.0,
669
  "argmax_agree": "596/596",
670
- "accuracy_vs_gold": 0.8372
671
  }
672
  },
673
  "belebele_split": {
674
  "verbatim_n": 252,
675
  "reworded_n": 648,
676
  "nz": {
677
- "verbatim": 0.7896825396825397,
678
- "reworded": 0.5632716049382716
679
  },
680
  "roeig": {
681
  "verbatim": 0.8412698412698413,
@@ -688,74 +688,74 @@
688
  },
689
  "calibration_pooled": {
690
  "n": 3990,
691
- "ece_15bin": 0.0274
692
  },
693
  "wording_robustness": {
694
- "suite": "wording_he v1 (frozen): 1786 questions, each asked in 6 wordings with the same meaning",
695
  "flip_rate": "share of questions whose top answer is not the same in all 6 wordings (lower is better)",
696
  "models": {
697
  "BrainboxAI/nitzotz": {
698
  "all_item_weighted": {
699
- "flip_rate": 0.0442,
700
- "flip_vs_orig": 0.0142,
701
- "mean_acc": 0.9081,
702
- "worst_acc": 0.8998,
703
- "flip_rate_spam_type_without_form_swap": 0.0431
704
  },
705
  "parts": {
706
  "heq_verify": {
707
  "n_items": 600,
708
- "mean_acc": 0.945,
709
- "worst_acc": 0.9417,
710
- "flip_rate": 0.0233
711
  },
712
  "massive_he_4": {
713
  "n_items": 500,
714
- "mean_acc": 0.967,
715
- "worst_acc": 0.958,
716
- "flip_rate": 0.034
717
  },
718
  "spam_business_he:category": {
719
  "n_items": 298,
720
- "mean_acc": 0.7645,
721
- "worst_acc": 0.7584,
722
- "flip_rate": 0.0671
723
  },
724
  "spam_business_he:category:hard": {
725
  "n_items": 65,
726
- "mean_acc": 0.641,
727
- "worst_acc": 0.6154,
728
  "flip_rate": 0.1077
729
  },
730
  "spam_business_he:scam": {
731
  "n_items": 298,
732
- "mean_acc": 0.9144,
733
- "worst_acc": 0.9094,
734
- "flip_rate": 0.0235
735
  },
736
  "spam_business_he:scam:hard": {
737
  "n_items": 65,
738
- "mean_acc": 0.8103,
739
  "worst_acc": 0.8,
740
- "flip_rate": 0.0462
741
  },
742
  "triage30:category": {
743
  "n_items": 30,
744
- "mean_acc": 0.8723,
745
- "worst_acc": 0.8667,
746
  "flip_rate": 0.0333
747
  },
748
  "triage30:paying": {
749
  "n_items": 30,
750
- "mean_acc": 0.7722,
751
  "worst_acc": 0.7333,
752
- "flip_rate": 0.2333
753
  },
754
  "triage30:urgency": {
755
  "n_items": 30,
756
- "mean_acc": 0.7222,
757
  "worst_acc": 0.6,
758
- "flip_rate": 0.4333
759
  }
760
  },
761
  "single_answer_wordings": {}
@@ -764,51 +764,145 @@
764
  },
765
  "belebele_negation": {
766
  "n": 158,
767
- "acc": 0.519
768
  },
769
  "seeds": {
770
- "heldout_this": 0.958,
771
  "others": [
772
  {
773
  "name": "seed 2",
774
- "heldout": 0.9492,
775
  "acc": {
776
- "spam_scam": 0.9128,
777
- "spam_scam_hard": 0.8154,
778
- "spam_type": 0.7718,
779
- "spam_type_hard": 0.6308,
780
- "triage_type": 0.8667,
781
- "triage_urgency": 0.6667,
782
- "triage_paying": 0.7667,
783
- "massive_20": 0.902,
784
- "massive_4": 0.964,
785
- "sib200": 0.8088,
786
  "heq_verify": 0.9467,
787
- "heq_unans": 0.8933,
788
- "belebele": 0.6311
789
  }
790
  },
791
  {
792
  "name": "seed 3",
793
- "heldout": 0.953,
794
  "acc": {
795
- "spam_scam": 0.9094,
796
- "spam_scam_hard": 0.8308,
797
- "spam_type": 0.7886,
798
- "spam_type_hard": 0.6154,
799
- "triage_type": 0.9333,
800
- "triage_urgency": 0.6667,
801
- "triage_paying": 0.8,
802
  "massive_20": 0.902,
803
- "massive_4": 0.966,
804
- "sib200": 0.8235,
805
- "heq_verify": 0.945,
806
- "heq_unans": 0.9,
807
- "belebele": 0.6344
808
  }
809
  }
810
  ],
811
- "agreement_min": 0.9058,
812
- "agreement_max": 0.9105
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
813
  }
814
  }
 
17
  "n": 298,
18
  "chance": 0.5,
19
  "acc": {
20
+ "nz": 0.9195,
21
  "roeig": 0.3221,
22
  "laya_ml": 0.5034
23
  },
24
  "brier": {
25
+ "nz": 0.1321,
26
  "roeig": 0.7604,
27
  "laya_ml": 0.8042
28
  },
29
  "ece": {
30
+ "nz": 0.0714,
31
  "roeig": 0.3802,
32
  "laya_ml": 0.38
33
  },
 
39
  "vs_nitzotz": {
40
  "roeig": {
41
  "nz_only": 186,
42
+ "other_only": 8,
43
+ "p": 3.5766488810898616e-45
44
  },
45
  "laya_ml": {
46
+ "nz_only": 138,
47
  "other_only": 14,
48
+ "p": 8.462475337125398e-27
49
  }
50
  }
51
  },
 
54
  "n": 65,
55
  "chance": 0.5,
56
  "acc": {
57
+ "nz": 0.8308,
58
  "roeig": 0.3231,
59
  "laya_ml": 0.3538
60
  },
61
  "brier": {
62
+ "nz": 0.2842,
63
  "roeig": 0.7076,
64
  "laya_ml": 1.1057
65
  },
66
  "ece": {
67
+ "nz": 0.1132,
68
  "roeig": 0.3391,
69
  "laya_ml": 0.5564
70
  },
 
76
  "vs_nitzotz": {
77
  "roeig": {
78
  "nz_only": 34,
79
+ "other_only": 1,
80
+ "p": 2.0954757928848267e-09
81
  },
82
  "laya_ml": {
83
+ "nz_only": 35,
84
+ "other_only": 4,
85
+ "p": 3.353161446284503e-07
86
  }
87
  }
88
  },
 
91
  "n": 298,
92
  "chance": 0.1667,
93
  "acc": {
94
+ "nz": 0.7785,
95
  "roeig": 0.5268,
96
  "laya_ml": 0.2081
97
  },
98
  "brier": {
99
+ "nz": 0.3312,
100
  "roeig": 0.6489,
101
  "laya_ml": 1.0746
102
  },
103
  "ece": {
104
+ "nz": 0.0683,
105
  "roeig": 0.1043,
106
  "laya_ml": 0.4131
107
  },
 
112
  },
113
  "vs_nitzotz": {
114
  "roeig": {
115
+ "nz_only": 108,
116
+ "other_only": 33,
117
+ "p": 1.6862794099851315e-10
118
  },
119
  "laya_ml": {
120
+ "nz_only": 177,
121
+ "other_only": 7,
122
+ "p": 1.0714267651676255e-43
123
  }
124
  }
125
  },
 
128
  "n": 65,
129
  "chance": 0.1667,
130
  "acc": {
131
+ "nz": 0.6769,
132
  "roeig": 0.4462,
133
  "laya_ml": 0.2462
134
  },
135
  "brier": {
136
+ "nz": 0.4865,
137
  "roeig": 0.7315,
138
  "laya_ml": 1.1094
139
  },
140
  "ece": {
141
+ "nz": 0.1027,
142
  "roeig": 0.1877,
143
  "laya_ml": 0.4109
144
  },
 
149
  },
150
  "vs_nitzotz": {
151
  "roeig": {
152
+ "nz_only": 23,
153
+ "other_only": 8,
154
+ "p": 0.010673840530216694
155
  },
156
  "laya_ml": {
157
+ "nz_only": 31,
158
+ "other_only": 3,
159
+ "p": 7.660128176212311e-07
160
  }
161
  }
162
  },
 
165
  "n": 30,
166
  "chance": 0.2,
167
  "acc": {
168
+ "nz": 0.9,
169
  "roeig": 0.8,
170
  "laya_ml": 0.5667
171
  },
172
  "brier": {
173
+ "nz": 0.1838,
174
  "roeig": 0.3191,
175
  "laya_ml": 0.6667
176
  },
177
  "ece": {
178
+ "nz": 0.1292,
179
  "roeig": 0.1321,
180
  "laya_ml": 0.2772
181
  },
 
186
  },
187
  "vs_nitzotz": {
188
  "roeig": {
189
+ "nz_only": 5,
190
  "other_only": 2,
191
+ "p": 0.453125
192
  },
193
  "laya_ml": {
194
+ "nz_only": 12,
195
  "other_only": 2,
196
+ "p": 0.012939453125
197
  }
198
  }
199
  },
 
202
  "n": 30,
203
  "chance": 0.2,
204
  "acc": {
205
+ "nz": 0.7,
206
  "roeig": 0.3667,
207
  "laya_ml": 0.3333
208
  },
209
  "brier": {
210
+ "nz": 0.4868,
211
  "roeig": 0.708,
212
  "laya_ml": 0.7501
213
  },
214
  "ece": {
215
+ "nz": 0.2041,
216
  "roeig": 0.0322,
217
  "laya_ml": 0.2167
218
  },
219
  "empty_acc": {
220
+ "nz": 0.4,
221
  "roeig": 0.2333,
222
  "laya_ml": 0.0667
223
  },
224
  "vs_nitzotz": {
225
  "roeig": {
226
+ "nz_only": 14,
227
+ "other_only": 4,
228
+ "p": 0.0308837890625
229
  },
230
  "laya_ml": {
231
+ "nz_only": 15,
232
+ "other_only": 4,
233
+ "p": 0.0192108154296875
234
  }
235
  }
236
  },
 
239
  "n": 30,
240
  "chance": 0.5,
241
  "acc": {
242
+ "nz": 0.8,
243
  "roeig": 0.5667,
244
  "laya_ml": 0.5
245
  },
246
  "brier": {
247
+ "nz": 0.2841,
248
  "roeig": 0.4338,
249
  "laya_ml": 0.8183
250
  },
251
  "ece": {
252
+ "nz": 0.1466,
253
  "roeig": 0.2362,
254
  "laya_ml": 0.4284
255
  },
 
260
  },
261
  "vs_nitzotz": {
262
  "roeig": {
263
+ "nz_only": 7,
264
  "other_only": 0,
265
+ "p": 0.015625
266
  },
267
  "laya_ml": {
268
  "nz_only": 14,
269
+ "other_only": 5,
270
+ "p": 0.063568115234375
271
  }
272
  }
273
  },
 
281
  "laya_ml": 0.474
282
  },
283
  "brier": {
284
+ "nz": 0.1671,
285
  "roeig": 0.3801,
286
  "laya_ml": 0.7338
287
  },
288
  "ece": {
289
+ "nz": 0.0583,
290
  "roeig": 0.0395,
291
  "laya_ml": 0.1661
292
  },
293
  "empty_acc": {
294
+ "nz": 0.064,
295
  "roeig": 0.052,
296
  "laya_ml": 0.044
297
  },
298
  "vs_nitzotz": {
299
  "roeig": {
300
+ "nz_only": 98,
301
+ "other_only": 11,
302
+ "p": 1.3281749873159828e-18
303
  },
304
  "laya_ml": {
305
+ "nz_only": 222,
306
+ "other_only": 9,
307
+ "p": 2.661469004156119e-54
308
  }
309
  }
310
  },
 
313
  "n": 500,
314
  "chance": 0.25,
315
  "acc": {
316
+ "nz": 0.972,
317
  "roeig": 0.908,
318
  "laya_ml": 0.694
319
  },
320
  "brier": {
321
+ "nz": 0.0462,
322
  "roeig": 0.1369,
323
  "laya_ml": 0.3987
324
  },
325
  "ece": {
326
+ "nz": 0.0501,
327
  "roeig": 0.0393,
328
  "laya_ml": 0.0468
329
  },
330
  "empty_acc": {
331
+ "nz": 0.254,
332
  "roeig": 0.242,
333
  "laya_ml": 0.256
334
  },
335
  "vs_nitzotz": {
336
  "roeig": {
337
+ "nz_only": 35,
338
+ "other_only": 3,
339
+ "p": 6.677873898297548e-08
340
  },
341
  "laya_ml": {
342
+ "nz_only": 141,
343
+ "other_only": 2,
344
+ "p": 1.846933796755538e-39
345
  }
346
  }
347
  },
 
350
  "n": 204,
351
  "chance": 0.1429,
352
  "acc": {
353
+ "nz": 0.7941,
354
  "roeig": 0.8235,
355
  "laya_ml": 0.6618
356
  },
357
  "brier": {
358
+ "nz": 0.3027,
359
  "roeig": 0.285,
360
  "laya_ml": 0.4918
361
  },
362
  "ece": {
363
+ "nz": 0.0708,
364
  "roeig": 0.0751,
365
  "laya_ml": 0.1784
366
  },
 
371
  },
372
  "vs_nitzotz": {
373
  "roeig": {
374
+ "nz_only": 12,
375
+ "other_only": 18,
376
+ "p": 0.361594608053565
377
  },
378
  "laya_ml": {
379
  "nz_only": 47,
380
+ "other_only": 20,
381
+ "p": 0.0013071686857636005
382
  }
383
  }
384
  },
 
392
  "laya_ml": 0.4883
393
  },
394
  "brier": {
395
+ "nz": 0.0938,
396
  "roeig": 0.0926,
397
  "laya_ml": 0.8274
398
  },
399
  "ece": {
400
+ "nz": 0.0365,
401
  "roeig": 0.0178,
402
  "laya_ml": 0.3946
403
  },
404
  "empty_acc": {
405
+ "nz": 0.51,
406
  "roeig": 0.8,
407
  "laya_ml": 0.59
408
  },
409
  "vs_nitzotz": {
410
  "roeig": {
411
+ "nz_only": 30,
412
+ "other_only": 26,
413
+ "p": 0.6888797607233396
414
  },
415
  "laya_ml": {
416
+ "nz_only": 296,
417
+ "other_only": 20,
418
+ "p": 3.520918733970021e-64
419
  }
420
  }
421
  },
 
424
  "n": 600,
425
  "chance": 0.5,
426
  "acc": {
427
+ "nz": 0.8917,
428
  "roeig": 0.53,
429
  "laya_ml": 0.5367
430
  },
431
  "brier": {
432
+ "nz": 0.1739,
433
  "roeig": 0.856,
434
  "laya_ml": 0.7348
435
  },
436
  "ece": {
437
+ "nz": 0.0484,
438
  "roeig": 0.4336,
439
  "laya_ml": 0.3421
440
  },
441
  "empty_acc": {
442
+ "nz": 0.5033,
443
  "roeig": 0.5033,
444
  "laya_ml": 0.5167
445
  },
446
  "vs_nitzotz": {
447
  "roeig": {
448
+ "nz_only": 243,
449
+ "other_only": 26,
450
+ "p": 2.502982180208191e-45
451
  },
452
  "laya_ml": {
453
+ "nz_only": 240,
454
+ "other_only": 27,
455
+ "p": 7.334224519429334e-44
456
  }
457
  }
458
  },
 
461
  "n": 900,
462
  "chance": 0.25,
463
  "acc": {
464
+ "nz": 0.6322,
465
  "roeig": 0.7544,
466
  "laya_ml": 0.3144
467
  },
468
  "brier": {
469
+ "nz": 0.5672,
470
  "roeig": 0.3442,
471
  "laya_ml": 0.8341
472
  },
473
  "ece": {
474
+ "nz": 0.1919,
475
  "roeig": 0.0324,
476
  "laya_ml": 0.2307
477
  },
478
  "empty_acc": {
479
+ "nz": 0.3011,
480
  "roeig": 0.3189,
481
  "laya_ml": 0.2844
482
  },
483
  "vs_nitzotz": {
484
  "roeig": {
485
+ "nz_only": 72,
486
+ "other_only": 182,
487
+ "p": 3.66511528586527e-12
488
  },
489
  "laya_ml": {
490
+ "nz_only": 366,
491
+ "other_only": 80,
492
+ "p": 9.158467645861244e-45
493
  }
494
  }
495
  }
 
497
  "significance": "paired exact McNemar against Nitzotz on the same questions",
498
  "latency_ms": {
499
  "q8": {
500
+ "median": 47.7,
501
+ "p90": 75.6,
502
  "n": 50
503
  },
504
  "f16": {
505
+ "median": 50.5,
506
+ "p90": 77.0,
507
  "n": 50
508
  },
509
  "py": {
510
+ "median": 60.0,
511
+ "p90": 98.2,
512
  "n": 50
513
  },
514
  "q8_cpu": {
515
+ "median": 309.1,
516
+ "p90": 711.0,
517
  "n": 50
518
  },
519
  "f16_cpu": {
520
+ "median": 425.9,
521
+ "p90": 854.0,
522
  "n": 50
523
  },
524
  "py_cpu": {
525
+ "median": 223.9,
526
+ "p90": 452.7,
527
  "n": 50
528
  }
529
  },
 
532
  "short": {
533
  "item": "massive_he_4-0075",
534
  "input_tokens": 38,
535
+ "median_ms": 36.9,
536
+ "p90_ms": 38.4,
537
+ "min_ms": 35.6
538
  },
539
  "long": {
540
  "item": "heq_unanswerable-0522",
541
  "input_tokens": 430,
542
+ "median_ms": 127.4,
543
+ "p90_ms": 130.0,
544
+ "min_ms": 108.0
545
  }
546
  },
547
  "f16": {
548
  "short": {
549
  "item": "massive_he_4-0075",
550
  "input_tokens": 38,
551
+ "median_ms": 43.2,
552
+ "p90_ms": 44.3,
553
+ "min_ms": 41.2
554
  },
555
  "long": {
556
  "item": "heq_unanswerable-0522",
557
  "input_tokens": 430,
558
+ "median_ms": 111.2,
559
+ "p90_ms": 118.2,
560
+ "min_ms": 108.7
561
  }
562
  },
563
  "py": {
 
565
  "item": "massive_he_4-0075",
566
  "input_tokens": 38,
567
  "median_ms": 43.6,
568
+ "p90_ms": 44.7,
569
+ "min_ms": 42.4
570
  },
571
  "long": {
572
  "item": "heq_unanswerable-0522",
573
  "input_tokens": 430,
574
+ "median_ms": 139.3,
575
+ "p90_ms": 157.0,
576
+ "min_ms": 137.1
577
  }
578
  },
579
  "q8_cpu": {
580
  "short": {
581
  "item": "massive_he_4-0075",
582
  "input_tokens": 38,
583
+ "median_ms": 159.5,
584
+ "p90_ms": 185.0,
585
+ "min_ms": 117.5
586
  },
587
  "long": {
588
  "item": "heq_unanswerable-0522",
589
  "input_tokens": 430,
590
+ "median_ms": 1137.4,
591
+ "p90_ms": 1243.6,
592
+ "min_ms": 1072.4
593
  }
594
  },
595
  "py_cpu": {
596
  "short": {
597
  "item": "massive_he_4-0075",
598
  "input_tokens": 38,
599
+ "median_ms": 112.0,
600
+ "p90_ms": 131.3,
601
+ "min_ms": 97.1
602
  },
603
  "long": {
604
  "item": "heq_unanswerable-0522",
605
  "input_tokens": 430,
606
+ "median_ms": 1021.8,
607
+ "p90_ms": 1100.7,
608
+ "min_ms": 918.8
609
  }
610
  }
611
  },
612
  "latency_machine": "Intel Core Ultra 9 285H laptop, built-in Arc 140T GPU, Windows 11",
613
  "gguf_parity_50": {
614
  "q8": {
615
+ "max_abs_diff": 0.0375,
616
+ "mean_abs_diff": 0.001814,
617
  "argmax_agree": "50/50"
618
  },
619
  "f16": {
620
+ "max_abs_diff": 0.0017,
621
+ "mean_abs_diff": 0.000288,
622
  "argmax_agree": "50/50"
623
  },
624
  "q8_cpu": {
625
+ "max_abs_diff": 0.0549,
626
+ "mean_abs_diff": 0.003498,
627
  "argmax_agree": "50/50"
628
  },
629
  "f16_cpu": {
630
+ "max_abs_diff": 0.0032,
631
+ "mean_abs_diff": 0.000278,
632
  "argmax_agree": "50/50"
633
  }
634
  },
635
  "gguf_parity_spam_596": {
636
  "q8": {
637
+ "max_abs_diff": 0.0166,
638
+ "mean_abs_diff": 0.001297,
639
+ "argmax_agree": "596/596",
640
+ "accuracy_vs_gold": 0.8473
641
  },
642
  "f16": {
643
+ "max_abs_diff": 0.0015,
644
+ "mean_abs_diff": 0.000203,
645
  "argmax_agree": "596/596",
646
+ "accuracy_vs_gold": 0.8473
647
  },
648
  "py": {
649
+ "max_abs_diff": 0.0,
650
  "mean_abs_diff": 0.0,
651
  "argmax_agree": "596/596",
652
+ "accuracy_vs_gold": 0.8473
653
  },
654
  "q8_cpu": {
655
+ "max_abs_diff": 0.0207,
656
+ "mean_abs_diff": 0.001918,
657
  "argmax_agree": "595/596",
658
+ "accuracy_vs_gold": 0.8456
659
  },
660
  "f16_cpu": {
661
+ "max_abs_diff": 0.0012,
662
+ "mean_abs_diff": 0.000185,
663
  "argmax_agree": "596/596",
664
+ "accuracy_vs_gold": 0.8473
665
  },
666
  "py_cpu": {
667
  "max_abs_diff": 0.0,
668
  "mean_abs_diff": 0.0,
669
  "argmax_agree": "596/596",
670
+ "accuracy_vs_gold": 0.8473
671
  }
672
  },
673
  "belebele_split": {
674
  "verbatim_n": 252,
675
  "reworded_n": 648,
676
  "nz": {
677
+ "verbatim": 0.7658730158730159,
678
+ "reworded": 0.5802469135802469
679
  },
680
  "roeig": {
681
  "verbatim": 0.8412698412698413,
 
688
  },
689
  "calibration_pooled": {
690
  "n": 3990,
691
+ "ece_15bin": 0.0309
692
  },
693
  "wording_robustness": {
694
+ "suite": "wording_he (frozen): 1786 questions, each asked in 6 wordings with the same meaning",
695
  "flip_rate": "share of questions whose top answer is not the same in all 6 wordings (lower is better)",
696
  "models": {
697
  "BrainboxAI/nitzotz": {
698
  "all_item_weighted": {
699
+ "flip_rate": 0.0504,
700
+ "flip_vs_orig": 0.018,
701
+ "mean_acc": 0.9118,
702
+ "worst_acc": 0.9048,
703
+ "flip_rate_spam_type_without_form_swap": 0.0481
704
  },
705
  "parts": {
706
  "heq_verify": {
707
  "n_items": 600,
708
+ "mean_acc": 0.9478,
709
+ "worst_acc": 0.945,
710
+ "flip_rate": 0.0283
711
  },
712
  "massive_he_4": {
713
  "n_items": 500,
714
+ "mean_acc": 0.97,
715
+ "worst_acc": 0.962,
716
+ "flip_rate": 0.036
717
  },
718
  "spam_business_he:category": {
719
  "n_items": 298,
720
+ "mean_acc": 0.7752,
721
+ "worst_acc": 0.7685,
722
+ "flip_rate": 0.0805
723
  },
724
  "spam_business_he:category:hard": {
725
  "n_items": 65,
726
+ "mean_acc": 0.6692,
727
+ "worst_acc": 0.6462,
728
  "flip_rate": 0.1077
729
  },
730
  "spam_business_he:scam": {
731
  "n_items": 298,
732
+ "mean_acc": 0.915,
733
+ "worst_acc": 0.9128,
734
+ "flip_rate": 0.0201
735
  },
736
  "spam_business_he:scam:hard": {
737
  "n_items": 65,
738
+ "mean_acc": 0.8154,
739
  "worst_acc": 0.8,
740
+ "flip_rate": 0.0615
741
  },
742
  "triage30:category": {
743
  "n_items": 30,
744
+ "mean_acc": 0.9166,
745
+ "worst_acc": 0.9,
746
  "flip_rate": 0.0333
747
  },
748
  "triage30:paying": {
749
  "n_items": 30,
750
+ "mean_acc": 0.7944,
751
  "worst_acc": 0.7333,
752
+ "flip_rate": 0.1667
753
  },
754
  "triage30:urgency": {
755
  "n_items": 30,
756
+ "mean_acc": 0.6611,
757
  "worst_acc": 0.6,
758
+ "flip_rate": 0.6333
759
  }
760
  },
761
  "single_answer_wordings": {}
 
764
  },
765
  "belebele_negation": {
766
  "n": 158,
767
+ "acc": 0.5886
768
  },
769
  "seeds": {
770
+ "heldout_this": 0.9505,
771
  "others": [
772
  {
773
  "name": "seed 2",
774
+ "heldout": 0.9519,
775
  "acc": {
776
+ "spam_scam": 0.9195,
777
+ "spam_scam_hard": 0.8308,
778
+ "spam_type": 0.802,
779
+ "spam_type_hard": 0.7077,
780
+ "triage_type": 0.9,
781
+ "triage_urgency": 0.7333,
782
+ "triage_paying": 0.8,
783
+ "massive_20": 0.894,
784
+ "massive_4": 0.96,
785
+ "sib200": 0.8284,
786
  "heq_verify": 0.9467,
787
+ "heq_unans": 0.8983,
788
+ "belebele": 0.6089
789
  }
790
  },
791
  {
792
  "name": "seed 3",
793
+ "heldout": 0.9521,
794
  "acc": {
795
+ "spam_scam": 0.906,
796
+ "spam_scam_hard": 0.7846,
797
+ "spam_type": 0.7785,
798
+ "spam_type_hard": 0.6615,
799
+ "triage_type": 0.8333,
800
+ "triage_urgency": 0.7667,
801
+ "triage_paying": 0.8333,
802
  "massive_20": 0.902,
803
+ "massive_4": 0.968,
804
+ "sib200": 0.8039,
805
+ "heq_verify": 0.95,
806
+ "heq_unans": 0.8983,
807
+ "belebele": 0.6311
808
  }
809
  }
810
  ],
811
+ "agreement_min": 0.9008,
812
+ "agreement_max": 0.9063,
813
+ "picked_on": {
814
+ "this": {
815
+ "n": 188,
816
+ "acc": 0.984
817
+ },
818
+ "others": [
819
+ {
820
+ "n": 188,
821
+ "acc": 0.9734
822
+ },
823
+ {
824
+ "n": 188,
825
+ "acc": 0.9734
826
+ }
827
+ ]
828
+ }
829
+ },
830
+ "none_of_the_options": {
831
+ "suite": "notinlist_he (frozen): 400 MASSIVE and Belebele test questions with a 'none of the options' choice added; in half of them 'none' is right, in the other half the right answer is still in the list",
832
+ "model": "BrainboxAI/nitzotz",
833
+ "parts": {
834
+ "all": {
835
+ "n": 400,
836
+ "acc": 0.685,
837
+ "picks_none": 0.535
838
+ },
839
+ "belebele_he/control": {
840
+ "n": 100,
841
+ "acc": 0.45,
842
+ "picks_none": 0.43
843
+ },
844
+ "belebele_he/none": {
845
+ "n": 100,
846
+ "acc": 0.76,
847
+ "picks_none": 0.76
848
+ },
849
+ "control": {
850
+ "n": 200,
851
+ "acc": 0.61,
852
+ "picks_none": 0.31
853
+ },
854
+ "massive_he_20/control": {
855
+ "n": 100,
856
+ "acc": 0.77,
857
+ "picks_none": 0.19
858
+ },
859
+ "massive_he_20/none": {
860
+ "n": 100,
861
+ "acc": 0.76,
862
+ "picks_none": 0.76
863
+ },
864
+ "none": {
865
+ "n": 200,
866
+ "acc": 0.76,
867
+ "picks_none": 0.76
868
+ }
869
+ },
870
+ "text_removed": {
871
+ "all": {
872
+ "n": 400,
873
+ "acc": 0.49,
874
+ "picks_none": 0.9825
875
+ },
876
+ "belebele_he/control": {
877
+ "n": 100,
878
+ "acc": 0.01,
879
+ "picks_none": 0.99
880
+ },
881
+ "belebele_he/none": {
882
+ "n": 100,
883
+ "acc": 0.95,
884
+ "picks_none": 0.95
885
+ },
886
+ "control": {
887
+ "n": 200,
888
+ "acc": 0.005,
889
+ "picks_none": 0.99
890
+ },
891
+ "massive_he_20/control": {
892
+ "n": 100,
893
+ "acc": 0.0,
894
+ "picks_none": 0.99
895
+ },
896
+ "massive_he_20/none": {
897
+ "n": 100,
898
+ "acc": 1.0,
899
+ "picks_none": 1.0
900
+ },
901
+ "none": {
902
+ "n": 200,
903
+ "acc": 0.975,
904
+ "picks_none": 0.975
905
+ }
906
+ }
907
  }
908
  }
ggmlc/nitzotz_trunk.py CHANGED
@@ -1,7 +1,7 @@
1
  """Exportable Nitzotz trunk for ggmlc: HalleluBERT (RoBERTa) encoder + the laya DecisionModel head.
2
 
3
  Adapted from ggmlc examples/laya/laya_trunk.py (Apache-2.0, monatis/ggmlc v0.9.5), which does the same for the
4
- ModernBERT laya checkpoints. What changed: the encoder is a post-LN RoBERTa (learned absolute positions, one token
5
  type, q/k/v with bias, exact GELU), and the option axis has MAX_OPTS = 20 slots because MASSIVE-20 asks 20 options.
6
  The head, the marker gather, the scorer and the act head are the laya ones, unchanged.
7
 
 
1
  """Exportable Nitzotz trunk for ggmlc: HalleluBERT (RoBERTa) encoder + the laya DecisionModel head.
2
 
3
  Adapted from ggmlc examples/laya/laya_trunk.py (Apache-2.0, monatis/ggmlc v0.9.5), which does the same for the
4
+ ModernBERT laya checkpoints. The difference: the encoder is a post-LN RoBERTa (learned absolute positions, one token
5
  type, q/k/v with bias, exact GELU), and the option axis has MAX_OPTS = 20 slots because MASSIVE-20 asks 20 options.
6
  The head, the marker gather, the scorer and the act head are the laya ones, unchanged.
7
 
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:0abc05a7e052aca7960bfe31962d64d399e233558373de321f34aeceb9b5b530
3
  size 767366700
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4bcd5c6ca58d4c57c3a2421f20fdece0144a0fa98cce05d97594aca6632805df
3
  size 767366700
nitzotz-f16.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b06086caefbf8b626dbadcfb02fe6a6b43047e4624296f74bc5228c50c964313
3
  size 770204096
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:05154d4577843f514a15a8ed857e9f14098d506f66340b0ddefa999d83eaf0e1
3
  size 770204096
nitzotz-q8_0.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:ea5f6fbff83711dc3f8dfa2c26919898cbfc109a4710de6ec46ffe792ab3cf30
3
  size 412580576
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:27b65ac93d8fd360105eb521d65863da878a3f1cf6611e2dcf8d15cc5ece981e
3
  size 412580576
rl_agent_config.json CHANGED
@@ -9,72 +9,76 @@
9
  "amp_dtype": "bf16",
10
  "model_name": "nitzotz",
11
  "temperature": [
12
- 1.1031566858291626,
13
- 1.110580563545227,
14
- 1.195957064628601
15
  ],
16
  "training": {
17
  "encoder_init": "HalleluBERT-large after extractive QA on HeQ v1.1 train (test-overlapping passages removed)",
18
- "data": "106,723 items: synthetic Hebrew messages, HeQ v1.1 train, MASSIVE he-IL train, reading and topic items on FineWeb-2 Hebrew passages, reworded copies (see README)",
19
- "items_total": 106723,
20
- "train_items": 102173,
21
- "calib_items": 4550,
22
  "epochs_planned": 2,
23
  "epochs_run": 2,
24
  "kept_epoch": 2,
25
  "heldout_by_epoch": [
26
  {
27
- "acc": 0.9424,
28
- "soft_nll": 0.4824,
29
- "acc_base_families": 0.9407,
30
  "by_family": {
31
- "aug_heq": 0.9344,
32
- "aug_massive": 0.9286,
33
- "aug_scam": 0.9677,
34
- "aug_type": 0.9482,
35
- "aug_urgency": 0.814,
36
- "claim": 0.968,
37
- "heq_read": 0.972,
38
- "heq_unans": 0.9,
39
- "heq_verify": 0.9533,
40
- "massive_intent": 0.9229,
41
- "rel_scam": 0.9833,
42
- "rel_type": 0.9254,
43
- "routing": 0.9338,
44
- "spam_type": 0.9471,
45
- "urgency": 0.8914
 
 
46
  },
47
  "epoch": 1,
48
- "step": 3193,
49
- "train_seconds": 1619.1
50
  },
51
  {
52
- "acc": 0.958,
53
- "soft_nll": 0.4701,
54
- "acc_base_families": 0.9545,
55
  "by_family": {
56
- "aug_heq": 0.9672,
57
- "aug_massive": 0.9643,
58
  "aug_scam": 0.9731,
59
- "aug_type": 0.9512,
60
- "aug_urgency": 0.9186,
61
  "claim": 0.9658,
62
  "heq_read": 0.976,
63
- "heq_unans": 0.936,
64
- "heq_verify": 0.9567,
65
- "massive_intent": 0.9429,
66
- "rel_scam": 0.9867,
67
- "rel_type": 0.9552,
68
- "routing": 0.9543,
69
- "spam_type": 0.9586,
70
- "urgency": 0.9257
 
 
71
  },
72
  "epoch": 2,
73
- "step": 6386,
74
- "train_seconds": 1630.1
75
  }
76
  ],
77
- "steps": 6386,
78
  "batch": 32,
79
  "micro_batch": 16,
80
  "lr_encoder": 1e-05,
@@ -86,7 +90,7 @@
86
  "loss": "soft cross-entropy",
87
  "amp": "bf16",
88
  "device": "cuda",
89
- "train_seconds": 3265.6,
90
  "head_init": "fresh (no laya checkpoint weights)",
91
  "class_weights": "none",
92
  "option_shuffle_families": [
 
9
  "amp_dtype": "bf16",
10
  "model_name": "nitzotz",
11
  "temperature": [
12
+ 1.0970542430877686,
13
+ 1.1389989852905273,
14
+ 1.1619657278060913
15
  ],
16
  "training": {
17
  "encoder_init": "HalleluBERT-large after extractive QA on HeQ v1.1 train (test-overlapping passages removed)",
18
+ "data": "116,854 items: synthetic Hebrew messages, HeQ v1.1 train, MASSIVE he-IL train, reading and topic items on FineWeb-2 Hebrew passages, warnings about scams, short scams and look-alike messages, 'none of the options' items, reworded copies (see README)",
19
+ "items_total": 116854,
20
+ "train_items": 112029,
21
+ "calib_items": 4825,
22
  "epochs_planned": 2,
23
  "epochs_run": 2,
24
  "kept_epoch": 2,
25
  "heldout_by_epoch": [
26
  {
27
+ "acc": 0.9389,
28
+ "soft_nll": 0.4909,
29
+ "acc_base_families": 0.9359,
30
  "by_family": {
31
+ "aug_heq": 0.9399,
32
+ "aug_massive": 0.9464,
33
+ "aug_scam": 0.9624,
34
+ "aug_type": 0.9299,
35
+ "aug_urgency": 0.8488,
36
+ "claim": 0.9498,
37
+ "heq_read": 0.964,
38
+ "heq_unans": 0.892,
39
+ "heq_verify": 0.9733,
40
+ "massive_intent": 0.9286,
41
+ "rel6_scam": 0.9574,
42
+ "rel6_type": 0.931,
43
+ "rel_scam": 0.9767,
44
+ "rel_type": 0.9216,
45
+ "routing": 0.9224,
46
+ "spam_type": 0.9329,
47
+ "urgency": 0.92
48
  },
49
  "epoch": 1,
50
+ "step": 3501,
51
+ "train_seconds": 1746.6
52
  },
53
  {
54
+ "acc": 0.9505,
55
+ "soft_nll": 0.4741,
56
+ "acc_base_families": 0.9452,
57
  "by_family": {
58
+ "aug_heq": 0.9563,
59
+ "aug_massive": 0.9375,
60
  "aug_scam": 0.9731,
61
+ "aug_type": 0.9482,
62
+ "aug_urgency": 0.8721,
63
  "claim": 0.9658,
64
  "heq_read": 0.976,
65
+ "heq_unans": 0.896,
66
+ "heq_verify": 0.98,
67
+ "massive_intent": 0.9314,
68
+ "rel6_scam": 0.9787,
69
+ "rel6_type": 0.9425,
70
+ "rel_scam": 0.9833,
71
+ "rel_type": 0.9515,
72
+ "routing": 0.9361,
73
+ "spam_type": 0.9486,
74
+ "urgency": 0.8971
75
  },
76
  "epoch": 2,
77
+ "step": 7002,
78
+ "train_seconds": 1757.1
79
  }
80
  ],
81
+ "steps": 7002,
82
  "batch": 32,
83
  "micro_batch": 16,
84
  "lr_encoder": 1e-05,
 
90
  "loss": "soft cross-entropy",
91
  "amp": "bf16",
92
  "device": "cuda",
93
+ "train_seconds": 3520.6,
94
  "head_init": "fresh (no laya checkpoint weights)",
95
  "class_weights": "none",
96
  "option_shuffle_families": [