BrainboxAI commited on
Commit
8ae5104
·
verified ·
1 Parent(s): ad9ac22

Rewrite the model card in English in the BrainboxAI house style; Hebrew kept only in example prompts; repository id unchanged

Browse files
Files changed (1) hide show
  1. README.md +110 -101
README.md CHANGED
@@ -32,62 +32,60 @@ model-index:
32
 
33
  # bx-code-nogah
34
 
35
- ### מזהה המאגר: `BrainboxAI/code-il-E4B`
36
 
37
- **עוזר כתיבת קוד ל-Python ול-TypeScript שרץ כולו על המחשב שלך. שורת קוד לא יוצאת החוצה.**
38
 
39
  [![HF Model](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Model-yellow)](https://huggingface.co/BrainboxAI/code-il-E4B)
40
  [![Dataset](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Dataset-blue)](https://huggingface.co/datasets/BrainboxAI/code-training-il)
41
  [![Safetensors](https://img.shields.io/badge/Format-Safetensors-green)](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors)
42
  [![License](https://img.shields.io/badge/License-Apache_2.0-lightgrey)](https://www.apache.org/licenses/LICENSE-2.0)
43
 
44
- > **על השם.** `bx-code-nogah` הוא שם המודל במוסכמת השמות של BrainboxAI: `bx` לחברה, `code` לתחום, ו-`nogah` (נוגה) לדרגת הגודל האמצעית. **מזהה המאגר נשאר `BrainboxAI/code-il-E4B` ולא ישתנה** — כל קישור וסקריפט קיים ממשיך לעבוד.
45
 
46
- > **על יציבות הגרסה.** אימון חדש על אותה משימה נדחף לאותו מאגר ומעדכן את המשקולות במקום. כלומר מי שיוריד היום ושוב בעוד חודשיים עלול לקבל משקולות שונות תחת אותו שם. מי שצריך יציבות מוחלטת — יצמיד את עצמו ל-commit מסוים ולא לענף הראשי.
47
-
48
- > **English.** `code-il-E4B` (brand name `bx-code-nogah`) is an on-device Python and TypeScript coding assistant, fine-tuned from `unsloth/gemma-4-E4B-it` on a 40,330-example test-filtered corpus. No formal benchmark (HumanEval, MBPP) has been run on it. This card is in Hebrew; identifiers, code and the recommended system prompt are in English.
49
 
50
  ---
51
 
52
- ## מה זה, בקצרה
53
 
54
- מודל שכותב ובודק קוד ב-Python וב-TypeScript, ורץ אצלך על המחשב. הוא בנוי על [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it) של גוגל, ואומן מעליו על 40,330 דוגמאות שסוננו לפי קריטריון אחד פשוט: **האם הקוד בדוגמה באמת עבר את הבדיקות שלו.**
55
 
56
- הכל נכנס בקובץ אחד של כ-5.3 ג'יגה-בייט. הוא רץ על:
57
 
58
- - מעבד של מחשב נייד מודרני — איטי, אבל עובד.
59
- - כל כרטיס מסך ביתי עם 6 ג'יגה זיכרון ומעלה.
60
- - מחשבי Apple עם שבב M, דרך llama.cpp.
61
 
62
- בלי חיבור לאינטרנט, בלי דיווח, בלי ששורת קוד אחת יוצאת מהמכונה.
63
 
64
- ## למה הוא קיים
65
 
66
- כל תו שנשלח לעוזר קוד בענן הוא דליפה פוטנציאלית. לחברה שבונה מערכת קניינית — ובמיוחד בפיננסים, ברפואה ובביטחון — זה פשוט לא עובר.
67
 
68
- המודל הזה הוא החלופה הפרטית: קטן מספיק לרוץ מקומית, מכוון לשתי השפות שרוב החברות באמת כותבות בהן.
69
 
70
- **הוא לא מתחרה ב-Claude או ב-GPT ביכולת גולמית, והוא לא מנסה.** הוא מציע משהו אחר: עזרה סבירה, בלי רשת, בלי שאף אחד אחר רואה את הקוד.
71
 
72
- ## למה הוא מיועד
73
 
74
- - השלמה וסקירה של קוד בסביבה מפוקחת שאסור לה לצאת החוצה.
75
- - התקנה בתוך חברה שיש לה כללי מקום-אחסון נוקשים.
76
- - עבודה בזוג עם מפתח כשהאינטרנט לא אמין או לא קיים.
77
- - הטמעה בתוך כלי פיתוח פנימי שאסור לו לקרוא ל-API חיצוני.
78
- - מפתחים דוברי עברית — המודל עונה בעברית כשפונים אליו בעברית, והקוד עצמו נשאר באנגלית.
79
 
80
- ## מה הוא **לא**, ומה אסור לעשות איתו
81
 
82
- - **הוא לא תחליף למודל גדול** בשאלות ארכיטקטורה, בקוד שפרוס על הרבה קבצים, או בכל דבר שדורש להחזיק הקשר ארוך בראש.
83
- - **אסור לשלוח את הקוד שלו לייצור בלי שאדם קרא אותו.** הוא מייצר קוד שנראה נכון ולא רץ. זו לא תקלה נדירה.
84
- - **הוא ממציא ממשקים של ספריות.** חתימות של פונקציות שלא קיימות, פרמטרים שלא קיימים, גרסאות שלא קיימות. תמיד לבדוק מול התיעוד.
85
- - **הוא מכיר רק Python ו-TypeScript.** בכל שפה אחרת הכיסוי מינימלי, והתחביר שיצא לא יהיה בהכרח נכון או מקובל.
86
- - **יש לו תאריך ידע.** ספריות וכלים שיצאו אחרי שהדאטה נאסף, בתחילת 2026, פשוט לא קיימים בשבילו.
87
- - **אין לו כלים.** הוא לא מריץ פקודות, לא קורא קבצים ולא בודק את עצמו. בשביל התנהגות של סוכן צריך לבנות מסביבו תשתית.
88
- - **אין לו ציון על מבחן מוכר.** ראה את פרק ההערכה — הבדיקות שכן נעשו קטנות מאוד ואינן מבחן.
89
 
90
- ## איך מריצים
91
 
92
  ### Ollama
93
 
@@ -98,7 +96,7 @@ ollama run hf.co/BrainboxAI/code-il-E4B:Q4_K_M
98
 
99
  ### llama.cpp
100
 
101
- שם הקובץ בתוך המאגר הוא `gemma-4-e4b-it.Q4_K_M.gguf`. השם נשמר משלב הבנייה — זה הקובץ של המודל **המאומן**, ולא של מודל הבסיס.
102
 
103
  ```bash
104
  ./llama-cli -m gemma-4-e4b-it.Q4_K_M.gguf \
@@ -106,7 +104,18 @@ ollama run hf.co/BrainboxAI/code-il-E4B:Q4_K_M
106
  --temp 0.2 --top-p 0.95 -n 1024
107
  ```
108
 
109
- ### פייתון, דרך מאגר ה-safetensors
 
 
 
 
 
 
 
 
 
 
 
110
 
111
  ```python
112
  from transformers import AutoTokenizer, AutoModelForCausalLM
@@ -126,24 +135,24 @@ outputs = model.generate(inputs, max_new_tokens=1024, temperature=0.2, top_p=0.9
126
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
127
  ```
128
 
129
- ### פרמטרי הרצה מומלצים
130
 
131
- | פרמטר | ערך | למה |
132
  |---|---|---|
133
- | `temperature` | 0.2 | יצירתיות נמוכה. בקוד רוצים תשובה צפויה, לא מקורית |
134
- | `top_p` | 0.95 | מעט גבוה מהמודל המשפטי, כדי לאפשר גיוון בסגנון כתיבה |
135
- | `max_new_tokens` | 1024 | מספיק לרוב הפונקציות |
136
- | `repetition_penalty` | 1.0 | קנס על חזרות פוגע בקוד — הזחה ושמות משתנים חוזרים על עצמם בכוונה |
137
 
138
- ## פרומפט המערכת המומלץ — זה מה שמשנה הכי הרבה
139
 
140
- מודל בגודל הזה כותב קוד **הרבה** יותר טוב כשמכריחים אותו לחשוב בחמישה שלבים לפני שהוא כותב שורה. בלי זה הוא קופץ ישר לקוד — קוד שמתקמפל ונופל על מקרה קצה, בלי בדיקות ובלי אזהרה.
141
 
142
- חמשת השלבים: להבין את הבעיה, למנות את מקרי הקצה, לכתוב את הקוד, לכתוב בדיקות, ולומר בכנות מה הקוד לא מכסה.
143
 
144
- **וזאת התרשמות, לא מדידה.** לא רצה השוואה מספרית בין הרצה עם הפרומפט להרצה בלעדיו.
145
 
146
- ### הפרומפט עצמו — העתק כמו שהוא
147
 
148
  ```text
149
  DEFINITIONS:
@@ -222,7 +231,7 @@ VERIFICATION:
222
  - regression check: No "production-ready" claims unless edge cases match limitations.
223
  ```
224
 
225
- ### דוגמת שימוש עם הפרומפט
226
 
227
  ```python
228
  from transformers import AutoTokenizer, AutoModelForCausalLM
@@ -234,7 +243,7 @@ model = AutoModelForCausalLM.from_pretrained(
234
  device_map="auto",
235
  )
236
 
237
- # הדבק כאן את הפרומפט המלא מהבלוק שלמעלה
238
  SYSTEM_PROMPT = """[paste the full prompt from the code block above]"""
239
 
240
  messages = [
@@ -247,86 +256,86 @@ outputs = model.generate(inputs, max_new_tokens=1500, temperature=0.2, top_p=0.9
247
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
248
  ```
249
 
250
- ### התאמות אפשריות
251
 
252
- - רוצה רק קוד בלי הסברים? החלף את `OUTPUT_FORMAT` ב-"Code blocks only".
253
- - בונה כלי לסקירת קוד? הוסף ל-`REQUIREMENTS` דרישה שהפלט יהיה בפורמט diff.
254
- - צריך רק TypeScript? הוסף ל-`REQUIREMENTS` שכל תשובה תהיה ב-TypeScript עם טיפוסים.
255
- - עובד על קוד רגיש אבטחתית? הוסף סעיף ל-`OUTPUT_FORMAT` בשם "Security Review".
256
 
257
- ## פרטי האימון
258
 
259
- | מאפיין | ערך |
260
  |---|---|
261
- | **מודל בסיס** | [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it) |
262
- | **שיטה** | QLoRA — מודל הבסיס נטען ב-4 ביט בזמן האימון |
263
- | **מסגרת** | Unsloth |
264
- | **חומרה** | NVIDIA RTX 5090 |
265
- | **שורות אימון** | 38,314 |
266
- | **שורות מבחן שהוחזקו בצד** | 2,016 |
267
- | **חלוקה** | 95% / 5%, seed 3407 |
268
- | **היפר-פרמטרים, זמן ועלות** | לא נכתבים כאן — ראה ההערה מתחת לטבלה |
269
 
270
- > **למה חסרים כאן מספרים.** רשומות האימון של המודל הזה שרדו בשתי גרסאות שסותרות זו את זו בדיוק על דרגת ה-LoRA ועל שאר ההיפר-פרמטרים. אי אפשר לדעת מהמקורות שקיימים איזו מהן מתארת את המשקולות שפורסמו כאן, ולכן השורות האלה הוסרו במקום להישאר ולהיראות כמו עובדה. מה שכן נשאר — מודל הבסיס, החומרה, וספירת השורות — מופיע באופן זהה בשני המקורות, וספירת השורות אף נקראה מקובץ סטטיסטיקה שנוצר על ידי המכונה עצמה.
271
 
272
- ### ממה מורכב הדאטה
273
 
274
- | מקור | כמות | תוכן |
275
  |---|---|---|
276
- | [`nvidia/OpenCodeInstruct`](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) | 20,000 | Python — רק דוגמאות שהקוד בהן עבר לפחות 50% מהבדיקות שלו |
277
  | [`bleugreen/typescript-instruct`](https://huggingface.co/datasets/bleugreen/typescript-instruct) | 20,000 | TypeScript |
278
- | סט זהות שנכתב ביד | 330 | 165 זוגות שאלה-תשובה, כל אחד נכלל פעמיים. עברית ואנגלית |
279
- | **סך הכל** | **40,330** | |
280
 
281
- **הסינון הוא העיקר כאן.** מקור ה-Python הוא קורפוס ענק. עברנו עליו וחתכנו לפי מבחן אחד: האם הקוד בדוגמה עבר את הבדיקות שנכתבו עבורו. דוגמאות בלי תוצאות בדיקה נזרקו, דוגמאות שעברו פחות מחצי נזרקו, וכפילויות לפי שאלה זהה נזרקו. גם אורך הטקסט נחתך ב-6,000 תווים.
282
 
283
- זו הייתה ההחלטה שהשפיעה הכי הרבה על התוצאה: אימון על הקורפוס המלא, בלי סינון, ייצר מודל רועש יותר.
284
 
285
- הכל מפורסם בכרטיס הדאטה של [`code-training-il`](https://huggingface.co/datasets/BrainboxAI/code-training-il).
286
 
287
- ## הערכה
288
 
289
- **לא רץ מבחן מוכר על המודל הזה. אין ציון HumanEval, אין MBPP, ואין שום מספר שאפשר להשוות מולו למודל אחר.**
290
 
291
- מה שכן נעשה — שתי בדיקות ידניות, קטנות:
292
 
293
- | מה נבדק | כמה מקרים | תוצאה |
294
  |---|---|---|
295
- | FizzBuzz, דרך לולאת סוכן | 5 | 5 מתוך 5, ב-6 צעדים, בלי סבב תיקון |
296
- | חיפוש בינארי עם 11 מקרי קצה | 11 | 11 מתוך 11, כולל טיפול בכפילות השמאלית |
297
 
298
- **איך לקרוא את זה, בכנות.** אלה 16 מקרים בסך הכל, ששני בני אדם הריצו ביד. אין קובץ תוצאות, אין קוד בדיקה שפורסם, ואי אפשר לשחזר את זה מבחוץ. זה מספיק כדי לומר "המודל עובד ולא קורס". זה **לא** מבחן, ו��סור להשוות אותו למספרים של מודלים אחרים.
299
 
300
- מבחן אמיתי הוא עבודה פתוחה. אם וכאשר הוא ירוץ, התוצאה תופיע כאן.
301
 
302
- ## מגבלות
303
 
304
- - **מודל קטן.** בגודל הזה יש טעויות בשאלות ארכיטקטורה ובהקשר ארוך. זו ודאות, לא אפשרות.
305
- - **שתי שפות.** חזק ב-Python וב-TypeScript, חלש בכל השאר.
306
- - **אין שימוש בכלים מהקופסה.** הוא מדבר, הוא לא מריץ. סוכן דורש עבודת אינטגרציה.
307
- - **תאריך ידע.** מה שיצא אחרי תחילת 2026 לא קיים בשבילו.
308
- - **הוא ממציא קוד שנראה נכון.** תמיד להריץ ולבדוק.
309
- - **אין מבחן.** ראה את פרק ההערכה.
310
- - **זהו אימון מעל `unsloth/gemma-4-E4B-it`.** כל מגבלה של מודל הבסיס נמצאת גם כאן.
311
 
312
- ## הקבצים והמאגרים
313
 
314
- | מאגר | מה יש בפנים | למי זה |
315
  |---|---|---|
316
- | [`BrainboxAI/code-il-E4B`](https://huggingface.co/BrainboxAI/code-il-E4B) | `gemma-4-e4b-it.Q4_K_M.gguf` (5.3 ג'יגה) והכרטיס הזה | Ollama, llama.cpp, LM Studio |
317
- | [`BrainboxAI/code-il-E4B-safetensors`](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors) | משקולות מלאות ב-16 ביט (16.0 ג'יגה) | transformers, והמשך אימון |
318
 
319
- במאגר יושב גם `gemma-4-e4b-it.BF16-mmproj.gguf` (0.99 ג'יגה). זהו רכיב הראייה של Gemma-4, שנחוץ רק אם רוצים להזין תמונות. לעבודת קוד אין בו צורך.
320
 
321
- ## רישיון
322
 
323
- Apache 2.0. מותר להשתמש, לשנות, להפיץ ולמכור נגזרות, עם ייחוס.
324
 
325
- זהו אימון מעל [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it), ולכן התנאים של מודל הבסיס חלים גם על המודל הזה. מודל הבסיס מפורסם תחת Apache 2.0 ומפנה גם אל [תנאי השימוש של Gemma](https://ai.google.dev/gemma/docs/gemma_4_license). כדאי לקרוא אותם לפני שימוש מסחרי.
326
 
327
- לחומר האימון יש רישיונות משלו, של המקורות שממנו הוא נבנה. ראה את כרטיס הדאטה.
328
 
329
- ## ציטוט
330
 
331
  ```bibtex
332
  @misc{elyasi2026codeil,
@@ -339,10 +348,10 @@ Apache 2.0. מותר להשתמש, לשנות, להפיץ ולמכור נגזר
339
  }
340
  ```
341
 
342
- ## מי בנה את זה
343
 
344
- נבנה על ידי [**נתנאל אליאסי**](https://huggingface.co/BrainboxAI), מייסד [BrainboxAI](https://brainboxai.io) — סטודיו ישראלי לבינה מלאכותית יישומית, שבונה מודלים קטנים, פרטיים ומתמחים.
345
 
346
- לכוונון מודל קוד על בסיס הקוד הפרטי של החברה שלך: [netanele@brainboxai.io](mailto:netanele@brainboxai.io).
347
 
348
- *חלק ממשפחת המודלים של BrainboxAI שרצים על החומרה שלך. ראה גם [`law-il-E2B`](https://huggingface.co/BrainboxAI/law-il-E2B) (משפט) ו-[`cyber-analyst-4B`](https://huggingface.co/BrainboxAI/cyber-analyst-4B) (סייבר).*
 
32
 
33
  # bx-code-nogah
34
 
35
+ ### Repository id: `BrainboxAI/code-il-E4B`
36
 
37
+ **A Python and TypeScript coding assistant that runs entirely on your own machine. Not one line of your code leaves it.**
38
 
39
  [![HF Model](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Model-yellow)](https://huggingface.co/BrainboxAI/code-il-E4B)
40
  [![Dataset](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Dataset-blue)](https://huggingface.co/datasets/BrainboxAI/code-training-il)
41
  [![Safetensors](https://img.shields.io/badge/Format-Safetensors-green)](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors)
42
  [![License](https://img.shields.io/badge/License-Apache_2.0-lightgrey)](https://www.apache.org/licenses/LICENSE-2.0)
43
 
44
+ > **About the name.** `bx-code-nogah` is this model's name under the BrainboxAI naming convention: `bx` for the lab, `code` for the domain, and `nogah` (Hebrew for the planet Venus, the morning star) for the middle size tier. **The repository id stays `BrainboxAI/code-il-E4B` and will not change.** Every existing link and script keeps working.
45
 
46
+ > **About version stability.** Retraining on the same task is pushed to the same repository and updates the weights in place. Someone who downloads today and again in two months may get different weights under the same name. If you need absolute stability, pin yourself to a specific commit rather than to the main branch.
 
 
47
 
48
  ---
49
 
50
+ ## What it is
51
 
52
+ A model that writes and reviews Python and TypeScript, running on your own hardware. It is built on Google's [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it) and fine-tuned on 40,330 examples filtered by one simple test: **did the code in the example actually pass its own tests?**
53
 
54
+ The whole thing fits in one file of about 5.3 GB. It runs on:
55
 
56
+ - A modern laptop CPU. Slow, but it works.
57
+ - Any consumer GPU with 6 GB of VRAM or more.
58
+ - Apple Silicon, through llama.cpp.
59
 
60
+ No network, no telemetry, and no line of code leaving the machine.
61
 
62
+ ## Why it exists
63
 
64
+ Every keystroke sent to a cloud coding assistant is a potential leak. For a company building a proprietary system, and especially in finance, healthcare or defence, that simply does not pass review.
65
 
66
+ This model is the private alternative: small enough to run locally, tuned for the two languages most companies actually write in.
67
 
68
+ **It does not compete with Claude or GPT on raw capability, and it is not trying to.** It offers something different: useful help, with no network, and nobody else reading your code.
69
 
70
+ ## What it is for
71
 
72
+ - Code completion and review inside a regulated environment that cannot reach the internet.
73
+ - On-premise deployment for companies with strict data-residency rules.
74
+ - Pair programming when the connection is unreliable or absent.
75
+ - Embedding into an internal developer tool that is not allowed to call an external API.
76
+ - Hebrew-speaking developers. The model answers in Hebrew when addressed in Hebrew, and the code itself stays in English.
77
 
78
+ ## What it is not, and what you must not do with it
79
 
80
+ - **It is not a replacement for a frontier model** on architecture questions, on code spread across many files, or on anything that needs a long context held in mind.
81
+ - **Do not ship its output to production without a person reading it.** It produces code that looks right and does not run. That is not a rare failure.
82
+ - **It invents library APIs.** Function signatures that do not exist, parameters that do not exist, versions that do not exist. Always check against the documentation.
83
+ - **It knows Python and TypeScript only.** Coverage of any other language is minimal, and the syntax it produces will not reliably be correct or idiomatic.
84
+ - **It has a knowledge cutoff.** Libraries and tools released after the data was collected in early 2026 simply do not exist for it.
85
+ - **It has no tool use out of the box.** It talks; it does not run commands, read files or check itself. Agent behaviour requires integration work around it.
86
+ - **It has no score on a recognised benchmark.** See the Evaluation section. The checks that were done are very small and are not a benchmark.
87
 
88
+ ## How to run it
89
 
90
  ### Ollama
91
 
 
96
 
97
  ### llama.cpp
98
 
99
+ The file inside the repository is named `gemma-4-e4b-it.Q4_K_M.gguf`. The name is left over from the build step. It is the **fine-tuned** model, not the base model.
100
 
101
  ```bash
102
  ./llama-cli -m gemma-4-e4b-it.Q4_K_M.gguf \
 
104
  --temp 0.2 --top-p 0.95 -n 1024
105
  ```
106
 
107
+ The model also takes Hebrew. Same request, asked in Hebrew:
108
+
109
+ ```bash
110
+ # Prompt: "Write me a Python function that parses ISO-8601 dates with timezones."
111
+ ./llama-cli -m gemma-4-e4b-it.Q4_K_M.gguf \
112
+ -p "תכתוב לי פונקציה בפייתון שמפרסרת תאריכים בפורמט ISO-8601 עם אזורי זמן." \
113
+ --temp 0.2 --top-p 0.95 -n 1024
114
+ ```
115
+
116
+ The explanation comes back in Hebrew. The code, the identifiers and the library names stay in English.
117
+
118
+ ### Python, through the safetensors repository
119
 
120
  ```python
121
  from transformers import AutoTokenizer, AutoModelForCausalLM
 
135
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
136
  ```
137
 
138
+ ### Recommended generation parameters
139
 
140
+ | Parameter | Value | Why |
141
  |---|---|---|
142
+ | `temperature` | 0.2 | Low creativity. Code wants the predictable answer, not the original one |
143
+ | `top_p` | 0.95 | Slightly higher than the legal model, to allow some idiom variety |
144
+ | `max_new_tokens` | 1024 | Enough for most function-level work |
145
+ | `repetition_penalty` | 1.0 | Penalising repetition hurts code. Indentation and variable names repeat on purpose |
146
 
147
+ ## The recommended system prompt, which matters more than anything else here
148
 
149
+ A model this size writes **much** better code when it is forced through five explicit steps before it writes a line. Without that it jumps straight to code, and the code compiles and then falls over on an edge case, with no tests and no warning.
150
 
151
+ The five steps: understand the problem, enumerate the edge cases, write the code, write tests, and state honestly what the code does not cover.
152
 
153
+ **And this is an impression, not a measurement.** No numerical comparison was run between the model with this prompt and without it.
154
 
155
+ ### The system prompt (copy as-is)
156
 
157
  ```text
158
  DEFINITIONS:
 
231
  - regression check: No "production-ready" claims unless edge cases match limitations.
232
  ```
233
 
234
+ ### Usage example with the system prompt
235
 
236
  ```python
237
  from transformers import AutoTokenizer, AutoModelForCausalLM
 
243
  device_map="auto",
244
  )
245
 
246
+ # Paste the full prompt from the code block above.
247
  SYSTEM_PROMPT = """[paste the full prompt from the code block above]"""
248
 
249
  messages = [
 
256
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
257
  ```
258
 
259
+ ### Customisation
260
 
261
+ - Want code only, with no prose? Replace `OUTPUT_FORMAT` with "Code blocks only".
262
+ - Building a code review tool? Add a requirement that output comes back as a diff.
263
+ - Want TypeScript only? Add a requirement that every answer is TypeScript with type annotations.
264
+ - Working on a security-sensitive codebase? Add a "Security Review" section to `OUTPUT_FORMAT`.
265
 
266
+ ## Training details
267
 
268
+ | Attribute | Value |
269
  |---|---|
270
+ | **Base model** | [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it) |
271
+ | **Method** | QLoRA. The base model is loaded in 4 bits during training |
272
+ | **Framework** | Unsloth |
273
+ | **Hardware** | NVIDIA RTX 5090 |
274
+ | **Training rows** | 38,314 |
275
+ | **Held-out rows** | 2,016 |
276
+ | **Split** | 95% / 5%, seed 3407 |
277
+ | **Hyperparameters, wall time and cost** | Not stated here. See the note below |
278
 
279
+ > **Why numbers are missing.** The training records for this model survived in two versions that contradict each other precisely on the LoRA rank and the rest of the hyperparameters. Nothing in the surviving sources says which version describes the weights published here, so those rows were removed rather than left on the card looking like fact. What did survive (the base model, the hardware and the row counts) appears identically in both sources, and the row counts were read from a statistics file written by the machine itself.
280
 
281
+ ### Dataset composition
282
 
283
+ | Source | Count | Content |
284
  |---|---|---|
285
+ | [`nvidia/OpenCodeInstruct`](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) | 20,000 | Python. Only examples whose code passed at least 50% of its own tests |
286
  | [`bleugreen/typescript-instruct`](https://huggingface.co/datasets/bleugreen/typescript-instruct) | 20,000 | TypeScript |
287
+ | Hand-written identity set | 330 | 165 question-and-answer pairs, each included twice. Hebrew and English |
288
+ | **Total** | **40,330** | |
289
 
290
+ **The filtering is the point here.** The Python source is an enormous corpus. It was cut down on one test: did the code in the example pass the tests written for it. Examples with no test results were dropped, examples that passed less than half were dropped, and duplicates by prompt hash were dropped. Text length was also capped at 6,000 characters.
291
 
292
+ That was the decision that moved the result most. Training on the full unfiltered corpus produced a noisier model.
293
 
294
+ The full account is on the [`code-training-il`](https://huggingface.co/datasets/BrainboxAI/code-training-il) dataset card.
295
 
296
+ ## Evaluation
297
 
298
+ **No recognised benchmark was run on this model. There is no HumanEval score, no MBPP score, and no number you can compare against another model.**
299
 
300
+ What was done instead: two small checks, run by hand.
301
 
302
+ | What was tested | Cases | Result |
303
  |---|---|---|
304
+ | FizzBuzz, through an agent loop | 5 | 5 of 5, in 6 steps, with no correction rounds |
305
+ | Binary search with 11 edge cases | 11 | 11 of 11, including leftmost-duplicate handling |
306
 
307
+ **How to read that, honestly.** Sixteen cases in total, run by hand. There is no results file, no published test code, and no way to reproduce it from outside. It is enough to say the model works and does not fall over. It is **not** a benchmark, and it must not be compared with other models' numbers.
308
 
309
+ A real benchmark is open work. If and when one is run, the result will appear here.
310
 
311
+ ## Limitations
312
 
313
+ - **It is a small model.** At this size there will be mistakes on architecture questions and long-context reasoning. That is a certainty, not a possibility.
314
+ - **Two languages.** Strong on Python and TypeScript, weak on everything else.
315
+ - **No tool use out of the box.** It talks, it does not run. An agent needs integration work.
316
+ - **Knowledge cutoff.** Anything released after early 2026 does not exist for it.
317
+ - **It produces code that looks right.** Always run it and test it.
318
+ - **No benchmark.** See the Evaluation section.
319
+ - **It is a fine-tune of `unsloth/gemma-4-E4B-it`.** Every limit of that model is still here.
320
 
321
+ ## Files and repositories
322
 
323
+ | Repository | What is inside | Who wants it |
324
  |---|---|---|
325
+ | [`BrainboxAI/code-il-E4B`](https://huggingface.co/BrainboxAI/code-il-E4B) | `gemma-4-e4b-it.Q4_K_M.gguf` (5.3 GB) and this card | Ollama, llama.cpp, LM Studio |
326
+ | [`BrainboxAI/code-il-E4B-safetensors`](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors) | Merged 16-bit weights (16.0 GB) | `transformers`, and continued training |
327
 
328
+ The repository also holds `gemma-4-e4b-it.BF16-mmproj.gguf` (0.99 GB). That is Gemma-4's vision component, needed only if you want to feed it images. Code work does not need it.
329
 
330
+ ## License
331
 
332
+ Apache 2.0. You may use, modify, distribute and sell derivatives, with attribution.
333
 
334
+ This is a fine-tune of [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it), so the terms of that model apply to this one as well. The base model is published under Apache 2.0 and also points to the [Gemma 4 licence terms](https://ai.google.dev/gemma/docs/gemma_4_license). Read those before relying on this line commercially.
335
 
336
+ The training material carries the licences of the sources it was built from. See the dataset card.
337
 
338
+ ## Citation
339
 
340
  ```bibtex
341
  @misc{elyasi2026codeil,
 
348
  }
349
  ```
350
 
351
+ ## Author
352
 
353
+ Built by [**Netanel Elyasi**](https://huggingface.co/BrainboxAI), founder of [BrainboxAI](https://brainboxai.io), an Israeli applied-AI studio building small, private, domain-specialised models.
354
 
355
+ For tuning a coding model on your company's own codebase: [netanele@brainboxai.io](mailto:netanele@brainboxai.io).
356
 
357
+ *Part of the BrainboxAI family of on-device models. See also [`law-il-E2B`](https://huggingface.co/BrainboxAI/law-il-E2B) (law) and [`cyber-analyst-4B`](https://huggingface.co/BrainboxAI/cyber-analyst-4B) (security).*