Text Generation
Transformers
GGUF
English
Hebrew
gemma4
image-text-to-text
code
python
typescript
coding-assistant
llama.cpp
ollama
unsloth
qlora
on-device
private-first
conversational
Instructions to use BrainboxAI/code-il-E4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BrainboxAI/code-il-E4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="BrainboxAI/code-il-E4B")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BrainboxAI/code-il-E4B") model = AutoModelForMultimodalLM.from_pretrained("BrainboxAI/code-il-E4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BrainboxAI/code-il-E4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BrainboxAI/code-il-E4B:BF16 # Run inference directly in the terminal: llama cli -hf BrainboxAI/code-il-E4B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BrainboxAI/code-il-E4B:BF16 # Run inference directly in the terminal: llama cli -hf BrainboxAI/code-il-E4B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BrainboxAI/code-il-E4B:BF16 # Run inference directly in the terminal: ./llama-cli -hf BrainboxAI/code-il-E4B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BrainboxAI/code-il-E4B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf BrainboxAI/code-il-E4B:BF16
Use Docker
docker model run hf.co/BrainboxAI/code-il-E4B:BF16
- LM Studio
- Jan
- vLLM
How to use BrainboxAI/code-il-E4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BrainboxAI/code-il-E4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BrainboxAI/code-il-E4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BrainboxAI/code-il-E4B:BF16
- SGLang
How to use BrainboxAI/code-il-E4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BrainboxAI/code-il-E4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BrainboxAI/code-il-E4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BrainboxAI/code-il-E4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BrainboxAI/code-il-E4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use BrainboxAI/code-il-E4B with Ollama:
ollama run hf.co/BrainboxAI/code-il-E4B:BF16
- Unsloth Desktop
- Pi
How to use BrainboxAI/code-il-E4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BrainboxAI/code-il-E4B:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "BrainboxAI/code-il-E4B:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use BrainboxAI/code-il-E4B with Docker Model Runner:
docker model run hf.co/BrainboxAI/code-il-E4B:BF16
- Lemonade
How to use BrainboxAI/code-il-E4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BrainboxAI/code-il-E4B:BF16
Run and chat with the model
lemonade run user.code-il-E4B-BF16
List all available models
lemonade list
- Hermes Agent
How to use BrainboxAI/code-il-E4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BrainboxAI/code-il-E4B:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default BrainboxAI/code-il-E4B:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use BrainboxAI/code-il-E4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BrainboxAI/code-il-E4B:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "BrainboxAI/code-il-E4B:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Rewrite the model card in English in the BrainboxAI house style; Hebrew kept only in example prompts; repository id unchanged
Browse files
README.md
CHANGED
|
@@ -32,62 +32,60 @@ model-index:
|
|
| 32 |
|
| 33 |
# bx-code-nogah
|
| 34 |
|
| 35 |
-
###
|
| 36 |
|
| 37 |
-
**
|
| 38 |
|
| 39 |
[](https://huggingface.co/BrainboxAI/code-il-E4B)
|
| 40 |
[](https://huggingface.co/datasets/BrainboxAI/code-training-il)
|
| 41 |
[](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors)
|
| 42 |
[](https://www.apache.org/licenses/LICENSE-2.0)
|
| 43 |
|
| 44 |
-
> **
|
| 45 |
|
| 46 |
-
> **
|
| 47 |
-
|
| 48 |
-
> **English.** `code-il-E4B` (brand name `bx-code-nogah`) is an on-device Python and TypeScript coding assistant, fine-tuned from `unsloth/gemma-4-E4B-it` on a 40,330-example test-filtered corpus. No formal benchmark (HumanEval, MBPP) has been run on it. This card is in Hebrew; identifiers, code and the recommended system prompt are in English.
|
| 49 |
|
| 50 |
---
|
| 51 |
|
| 52 |
-
##
|
| 53 |
|
| 54 |
-
|
| 55 |
|
| 56 |
-
|
| 57 |
|
| 58 |
-
-
|
| 59 |
-
-
|
| 60 |
-
-
|
| 61 |
|
| 62 |
-
|
| 63 |
|
| 64 |
-
##
|
| 65 |
|
| 66 |
-
|
| 67 |
|
| 68 |
-
|
| 69 |
|
| 70 |
-
**
|
| 71 |
|
| 72 |
-
##
|
| 73 |
|
| 74 |
-
-
|
| 75 |
-
-
|
| 76 |
-
-
|
| 77 |
-
-
|
| 78 |
-
-
|
| 79 |
|
| 80 |
-
##
|
| 81 |
|
| 82 |
-
- **
|
| 83 |
-
- **
|
| 84 |
-
- **
|
| 85 |
-
- **
|
| 86 |
-
- **
|
| 87 |
-
- **
|
| 88 |
-
- **
|
| 89 |
|
| 90 |
-
##
|
| 91 |
|
| 92 |
### Ollama
|
| 93 |
|
|
@@ -98,7 +96,7 @@ ollama run hf.co/BrainboxAI/code-il-E4B:Q4_K_M
|
|
| 98 |
|
| 99 |
### llama.cpp
|
| 100 |
|
| 101 |
-
|
| 102 |
|
| 103 |
```bash
|
| 104 |
./llama-cli -m gemma-4-e4b-it.Q4_K_M.gguf \
|
|
@@ -106,7 +104,18 @@ ollama run hf.co/BrainboxAI/code-il-E4B:Q4_K_M
|
|
| 106 |
--temp 0.2 --top-p 0.95 -n 1024
|
| 107 |
```
|
| 108 |
|
| 109 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 110 |
|
| 111 |
```python
|
| 112 |
from transformers import AutoTokenizer, AutoModelForCausalLM
|
|
@@ -126,24 +135,24 @@ outputs = model.generate(inputs, max_new_tokens=1024, temperature=0.2, top_p=0.9
|
|
| 126 |
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 127 |
```
|
| 128 |
|
| 129 |
-
###
|
| 130 |
|
| 131 |
-
|
|
| 132 |
|---|---|---|
|
| 133 |
-
| `temperature` | 0.2 |
|
| 134 |
-
| `top_p` | 0.95 |
|
| 135 |
-
| `max_new_tokens` | 1024 |
|
| 136 |
-
| `repetition_penalty` | 1.0 |
|
| 137 |
|
| 138 |
-
##
|
| 139 |
|
| 140 |
-
|
| 141 |
|
| 142 |
-
|
| 143 |
|
| 144 |
-
**
|
| 145 |
|
| 146 |
-
###
|
| 147 |
|
| 148 |
```text
|
| 149 |
DEFINITIONS:
|
|
@@ -222,7 +231,7 @@ VERIFICATION:
|
|
| 222 |
- regression check: No "production-ready" claims unless edge cases match limitations.
|
| 223 |
```
|
| 224 |
|
| 225 |
-
###
|
| 226 |
|
| 227 |
```python
|
| 228 |
from transformers import AutoTokenizer, AutoModelForCausalLM
|
|
@@ -234,7 +243,7 @@ model = AutoModelForCausalLM.from_pretrained(
|
|
| 234 |
device_map="auto",
|
| 235 |
)
|
| 236 |
|
| 237 |
-
#
|
| 238 |
SYSTEM_PROMPT = """[paste the full prompt from the code block above]"""
|
| 239 |
|
| 240 |
messages = [
|
|
@@ -247,86 +256,86 @@ outputs = model.generate(inputs, max_new_tokens=1500, temperature=0.2, top_p=0.9
|
|
| 247 |
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 248 |
```
|
| 249 |
|
| 250 |
-
###
|
| 251 |
|
| 252 |
-
-
|
| 253 |
-
-
|
| 254 |
-
-
|
| 255 |
-
-
|
| 256 |
|
| 257 |
-
##
|
| 258 |
|
| 259 |
-
|
|
| 260 |
|---|---|
|
| 261 |
-
| **
|
| 262 |
-
| **
|
| 263 |
-
| **
|
| 264 |
-
| **
|
| 265 |
-
| **
|
| 266 |
-
| **
|
| 267 |
-
| **
|
| 268 |
-
| **
|
| 269 |
|
| 270 |
-
> **
|
| 271 |
|
| 272 |
-
###
|
| 273 |
|
| 274 |
-
|
|
| 275 |
|---|---|---|
|
| 276 |
-
| [`nvidia/OpenCodeInstruct`](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) | 20,000 | Python
|
| 277 |
| [`bleugreen/typescript-instruct`](https://huggingface.co/datasets/bleugreen/typescript-instruct) | 20,000 | TypeScript |
|
| 278 |
-
|
|
| 279 |
-
| **
|
| 280 |
|
| 281 |
-
**
|
| 282 |
|
| 283 |
-
|
| 284 |
|
| 285 |
-
|
| 286 |
|
| 287 |
-
##
|
| 288 |
|
| 289 |
-
**
|
| 290 |
|
| 291 |
-
|
| 292 |
|
| 293 |
-
|
|
| 294 |
|---|---|---|
|
| 295 |
-
| FizzBuzz,
|
| 296 |
-
|
|
| 297 |
|
| 298 |
-
**
|
| 299 |
|
| 300 |
-
|
| 301 |
|
| 302 |
-
##
|
| 303 |
|
| 304 |
-
- **
|
| 305 |
-
- **
|
| 306 |
-
- **
|
| 307 |
-
- **
|
| 308 |
-
- **
|
| 309 |
-
- **
|
| 310 |
-
- **
|
| 311 |
|
| 312 |
-
##
|
| 313 |
|
| 314 |
-
|
|
| 315 |
|---|---|---|
|
| 316 |
-
| [`BrainboxAI/code-il-E4B`](https://huggingface.co/BrainboxAI/code-il-E4B) | `gemma-4-e4b-it.Q4_K_M.gguf` (5.3
|
| 317 |
-
| [`BrainboxAI/code-il-E4B-safetensors`](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors) |
|
| 318 |
|
| 319 |
-
|
| 320 |
|
| 321 |
-
##
|
| 322 |
|
| 323 |
-
Apache 2.0.
|
| 324 |
|
| 325 |
-
|
| 326 |
|
| 327 |
-
|
| 328 |
|
| 329 |
-
##
|
| 330 |
|
| 331 |
```bibtex
|
| 332 |
@misc{elyasi2026codeil,
|
|
@@ -339,10 +348,10 @@ Apache 2.0. מותר להשתמש, לשנות, להפיץ ולמכור נגזר
|
|
| 339 |
}
|
| 340 |
```
|
| 341 |
|
| 342 |
-
##
|
| 343 |
|
| 344 |
-
|
| 345 |
|
| 346 |
-
|
| 347 |
|
| 348 |
-
*
|
|
|
|
| 32 |
|
| 33 |
# bx-code-nogah
|
| 34 |
|
| 35 |
+
### Repository id: `BrainboxAI/code-il-E4B`
|
| 36 |
|
| 37 |
+
**A Python and TypeScript coding assistant that runs entirely on your own machine. Not one line of your code leaves it.**
|
| 38 |
|
| 39 |
[](https://huggingface.co/BrainboxAI/code-il-E4B)
|
| 40 |
[](https://huggingface.co/datasets/BrainboxAI/code-training-il)
|
| 41 |
[](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors)
|
| 42 |
[](https://www.apache.org/licenses/LICENSE-2.0)
|
| 43 |
|
| 44 |
+
> **About the name.** `bx-code-nogah` is this model's name under the BrainboxAI naming convention: `bx` for the lab, `code` for the domain, and `nogah` (Hebrew for the planet Venus, the morning star) for the middle size tier. **The repository id stays `BrainboxAI/code-il-E4B` and will not change.** Every existing link and script keeps working.
|
| 45 |
|
| 46 |
+
> **About version stability.** Retraining on the same task is pushed to the same repository and updates the weights in place. Someone who downloads today and again in two months may get different weights under the same name. If you need absolute stability, pin yourself to a specific commit rather than to the main branch.
|
|
|
|
|
|
|
| 47 |
|
| 48 |
---
|
| 49 |
|
| 50 |
+
## What it is
|
| 51 |
|
| 52 |
+
A model that writes and reviews Python and TypeScript, running on your own hardware. It is built on Google's [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it) and fine-tuned on 40,330 examples filtered by one simple test: **did the code in the example actually pass its own tests?**
|
| 53 |
|
| 54 |
+
The whole thing fits in one file of about 5.3 GB. It runs on:
|
| 55 |
|
| 56 |
+
- A modern laptop CPU. Slow, but it works.
|
| 57 |
+
- Any consumer GPU with 6 GB of VRAM or more.
|
| 58 |
+
- Apple Silicon, through llama.cpp.
|
| 59 |
|
| 60 |
+
No network, no telemetry, and no line of code leaving the machine.
|
| 61 |
|
| 62 |
+
## Why it exists
|
| 63 |
|
| 64 |
+
Every keystroke sent to a cloud coding assistant is a potential leak. For a company building a proprietary system, and especially in finance, healthcare or defence, that simply does not pass review.
|
| 65 |
|
| 66 |
+
This model is the private alternative: small enough to run locally, tuned for the two languages most companies actually write in.
|
| 67 |
|
| 68 |
+
**It does not compete with Claude or GPT on raw capability, and it is not trying to.** It offers something different: useful help, with no network, and nobody else reading your code.
|
| 69 |
|
| 70 |
+
## What it is for
|
| 71 |
|
| 72 |
+
- Code completion and review inside a regulated environment that cannot reach the internet.
|
| 73 |
+
- On-premise deployment for companies with strict data-residency rules.
|
| 74 |
+
- Pair programming when the connection is unreliable or absent.
|
| 75 |
+
- Embedding into an internal developer tool that is not allowed to call an external API.
|
| 76 |
+
- Hebrew-speaking developers. The model answers in Hebrew when addressed in Hebrew, and the code itself stays in English.
|
| 77 |
|
| 78 |
+
## What it is not, and what you must not do with it
|
| 79 |
|
| 80 |
+
- **It is not a replacement for a frontier model** on architecture questions, on code spread across many files, or on anything that needs a long context held in mind.
|
| 81 |
+
- **Do not ship its output to production without a person reading it.** It produces code that looks right and does not run. That is not a rare failure.
|
| 82 |
+
- **It invents library APIs.** Function signatures that do not exist, parameters that do not exist, versions that do not exist. Always check against the documentation.
|
| 83 |
+
- **It knows Python and TypeScript only.** Coverage of any other language is minimal, and the syntax it produces will not reliably be correct or idiomatic.
|
| 84 |
+
- **It has a knowledge cutoff.** Libraries and tools released after the data was collected in early 2026 simply do not exist for it.
|
| 85 |
+
- **It has no tool use out of the box.** It talks; it does not run commands, read files or check itself. Agent behaviour requires integration work around it.
|
| 86 |
+
- **It has no score on a recognised benchmark.** See the Evaluation section. The checks that were done are very small and are not a benchmark.
|
| 87 |
|
| 88 |
+
## How to run it
|
| 89 |
|
| 90 |
### Ollama
|
| 91 |
|
|
|
|
| 96 |
|
| 97 |
### llama.cpp
|
| 98 |
|
| 99 |
+
The file inside the repository is named `gemma-4-e4b-it.Q4_K_M.gguf`. The name is left over from the build step. It is the **fine-tuned** model, not the base model.
|
| 100 |
|
| 101 |
```bash
|
| 102 |
./llama-cli -m gemma-4-e4b-it.Q4_K_M.gguf \
|
|
|
|
| 104 |
--temp 0.2 --top-p 0.95 -n 1024
|
| 105 |
```
|
| 106 |
|
| 107 |
+
The model also takes Hebrew. Same request, asked in Hebrew:
|
| 108 |
+
|
| 109 |
+
```bash
|
| 110 |
+
# Prompt: "Write me a Python function that parses ISO-8601 dates with timezones."
|
| 111 |
+
./llama-cli -m gemma-4-e4b-it.Q4_K_M.gguf \
|
| 112 |
+
-p "תכתוב לי פונקציה בפייתון שמפרסרת תאריכים בפורמט ISO-8601 עם אזורי זמן." \
|
| 113 |
+
--temp 0.2 --top-p 0.95 -n 1024
|
| 114 |
+
```
|
| 115 |
+
|
| 116 |
+
The explanation comes back in Hebrew. The code, the identifiers and the library names stay in English.
|
| 117 |
+
|
| 118 |
+
### Python, through the safetensors repository
|
| 119 |
|
| 120 |
```python
|
| 121 |
from transformers import AutoTokenizer, AutoModelForCausalLM
|
|
|
|
| 135 |
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 136 |
```
|
| 137 |
|
| 138 |
+
### Recommended generation parameters
|
| 139 |
|
| 140 |
+
| Parameter | Value | Why |
|
| 141 |
|---|---|---|
|
| 142 |
+
| `temperature` | 0.2 | Low creativity. Code wants the predictable answer, not the original one |
|
| 143 |
+
| `top_p` | 0.95 | Slightly higher than the legal model, to allow some idiom variety |
|
| 144 |
+
| `max_new_tokens` | 1024 | Enough for most function-level work |
|
| 145 |
+
| `repetition_penalty` | 1.0 | Penalising repetition hurts code. Indentation and variable names repeat on purpose |
|
| 146 |
|
| 147 |
+
## The recommended system prompt, which matters more than anything else here
|
| 148 |
|
| 149 |
+
A model this size writes **much** better code when it is forced through five explicit steps before it writes a line. Without that it jumps straight to code, and the code compiles and then falls over on an edge case, with no tests and no warning.
|
| 150 |
|
| 151 |
+
The five steps: understand the problem, enumerate the edge cases, write the code, write tests, and state honestly what the code does not cover.
|
| 152 |
|
| 153 |
+
**And this is an impression, not a measurement.** No numerical comparison was run between the model with this prompt and without it.
|
| 154 |
|
| 155 |
+
### The system prompt (copy as-is)
|
| 156 |
|
| 157 |
```text
|
| 158 |
DEFINITIONS:
|
|
|
|
| 231 |
- regression check: No "production-ready" claims unless edge cases match limitations.
|
| 232 |
```
|
| 233 |
|
| 234 |
+
### Usage example with the system prompt
|
| 235 |
|
| 236 |
```python
|
| 237 |
from transformers import AutoTokenizer, AutoModelForCausalLM
|
|
|
|
| 243 |
device_map="auto",
|
| 244 |
)
|
| 245 |
|
| 246 |
+
# Paste the full prompt from the code block above.
|
| 247 |
SYSTEM_PROMPT = """[paste the full prompt from the code block above]"""
|
| 248 |
|
| 249 |
messages = [
|
|
|
|
| 256 |
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 257 |
```
|
| 258 |
|
| 259 |
+
### Customisation
|
| 260 |
|
| 261 |
+
- Want code only, with no prose? Replace `OUTPUT_FORMAT` with "Code blocks only".
|
| 262 |
+
- Building a code review tool? Add a requirement that output comes back as a diff.
|
| 263 |
+
- Want TypeScript only? Add a requirement that every answer is TypeScript with type annotations.
|
| 264 |
+
- Working on a security-sensitive codebase? Add a "Security Review" section to `OUTPUT_FORMAT`.
|
| 265 |
|
| 266 |
+
## Training details
|
| 267 |
|
| 268 |
+
| Attribute | Value |
|
| 269 |
|---|---|
|
| 270 |
+
| **Base model** | [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it) |
|
| 271 |
+
| **Method** | QLoRA. The base model is loaded in 4 bits during training |
|
| 272 |
+
| **Framework** | Unsloth |
|
| 273 |
+
| **Hardware** | NVIDIA RTX 5090 |
|
| 274 |
+
| **Training rows** | 38,314 |
|
| 275 |
+
| **Held-out rows** | 2,016 |
|
| 276 |
+
| **Split** | 95% / 5%, seed 3407 |
|
| 277 |
+
| **Hyperparameters, wall time and cost** | Not stated here. See the note below |
|
| 278 |
|
| 279 |
+
> **Why numbers are missing.** The training records for this model survived in two versions that contradict each other precisely on the LoRA rank and the rest of the hyperparameters. Nothing in the surviving sources says which version describes the weights published here, so those rows were removed rather than left on the card looking like fact. What did survive (the base model, the hardware and the row counts) appears identically in both sources, and the row counts were read from a statistics file written by the machine itself.
|
| 280 |
|
| 281 |
+
### Dataset composition
|
| 282 |
|
| 283 |
+
| Source | Count | Content |
|
| 284 |
|---|---|---|
|
| 285 |
+
| [`nvidia/OpenCodeInstruct`](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) | 20,000 | Python. Only examples whose code passed at least 50% of its own tests |
|
| 286 |
| [`bleugreen/typescript-instruct`](https://huggingface.co/datasets/bleugreen/typescript-instruct) | 20,000 | TypeScript |
|
| 287 |
+
| Hand-written identity set | 330 | 165 question-and-answer pairs, each included twice. Hebrew and English |
|
| 288 |
+
| **Total** | **40,330** | |
|
| 289 |
|
| 290 |
+
**The filtering is the point here.** The Python source is an enormous corpus. It was cut down on one test: did the code in the example pass the tests written for it. Examples with no test results were dropped, examples that passed less than half were dropped, and duplicates by prompt hash were dropped. Text length was also capped at 6,000 characters.
|
| 291 |
|
| 292 |
+
That was the decision that moved the result most. Training on the full unfiltered corpus produced a noisier model.
|
| 293 |
|
| 294 |
+
The full account is on the [`code-training-il`](https://huggingface.co/datasets/BrainboxAI/code-training-il) dataset card.
|
| 295 |
|
| 296 |
+
## Evaluation
|
| 297 |
|
| 298 |
+
**No recognised benchmark was run on this model. There is no HumanEval score, no MBPP score, and no number you can compare against another model.**
|
| 299 |
|
| 300 |
+
What was done instead: two small checks, run by hand.
|
| 301 |
|
| 302 |
+
| What was tested | Cases | Result |
|
| 303 |
|---|---|---|
|
| 304 |
+
| FizzBuzz, through an agent loop | 5 | 5 of 5, in 6 steps, with no correction rounds |
|
| 305 |
+
| Binary search with 11 edge cases | 11 | 11 of 11, including leftmost-duplicate handling |
|
| 306 |
|
| 307 |
+
**How to read that, honestly.** Sixteen cases in total, run by hand. There is no results file, no published test code, and no way to reproduce it from outside. It is enough to say the model works and does not fall over. It is **not** a benchmark, and it must not be compared with other models' numbers.
|
| 308 |
|
| 309 |
+
A real benchmark is open work. If and when one is run, the result will appear here.
|
| 310 |
|
| 311 |
+
## Limitations
|
| 312 |
|
| 313 |
+
- **It is a small model.** At this size there will be mistakes on architecture questions and long-context reasoning. That is a certainty, not a possibility.
|
| 314 |
+
- **Two languages.** Strong on Python and TypeScript, weak on everything else.
|
| 315 |
+
- **No tool use out of the box.** It talks, it does not run. An agent needs integration work.
|
| 316 |
+
- **Knowledge cutoff.** Anything released after early 2026 does not exist for it.
|
| 317 |
+
- **It produces code that looks right.** Always run it and test it.
|
| 318 |
+
- **No benchmark.** See the Evaluation section.
|
| 319 |
+
- **It is a fine-tune of `unsloth/gemma-4-E4B-it`.** Every limit of that model is still here.
|
| 320 |
|
| 321 |
+
## Files and repositories
|
| 322 |
|
| 323 |
+
| Repository | What is inside | Who wants it |
|
| 324 |
|---|---|---|
|
| 325 |
+
| [`BrainboxAI/code-il-E4B`](https://huggingface.co/BrainboxAI/code-il-E4B) | `gemma-4-e4b-it.Q4_K_M.gguf` (5.3 GB) and this card | Ollama, llama.cpp, LM Studio |
|
| 326 |
+
| [`BrainboxAI/code-il-E4B-safetensors`](https://huggingface.co/BrainboxAI/code-il-E4B-safetensors) | Merged 16-bit weights (16.0 GB) | `transformers`, and continued training |
|
| 327 |
|
| 328 |
+
The repository also holds `gemma-4-e4b-it.BF16-mmproj.gguf` (0.99 GB). That is Gemma-4's vision component, needed only if you want to feed it images. Code work does not need it.
|
| 329 |
|
| 330 |
+
## License
|
| 331 |
|
| 332 |
+
Apache 2.0. You may use, modify, distribute and sell derivatives, with attribution.
|
| 333 |
|
| 334 |
+
This is a fine-tune of [`unsloth/gemma-4-E4B-it`](https://huggingface.co/unsloth/gemma-4-E4B-it), so the terms of that model apply to this one as well. The base model is published under Apache 2.0 and also points to the [Gemma 4 licence terms](https://ai.google.dev/gemma/docs/gemma_4_license). Read those before relying on this line commercially.
|
| 335 |
|
| 336 |
+
The training material carries the licences of the sources it was built from. See the dataset card.
|
| 337 |
|
| 338 |
+
## Citation
|
| 339 |
|
| 340 |
```bibtex
|
| 341 |
@misc{elyasi2026codeil,
|
|
|
|
| 348 |
}
|
| 349 |
```
|
| 350 |
|
| 351 |
+
## Author
|
| 352 |
|
| 353 |
+
Built by [**Netanel Elyasi**](https://huggingface.co/BrainboxAI), founder of [BrainboxAI](https://brainboxai.io), an Israeli applied-AI studio building small, private, domain-specialised models.
|
| 354 |
|
| 355 |
+
For tuning a coding model on your company's own codebase: [netanele@brainboxai.io](mailto:netanele@brainboxai.io).
|
| 356 |
|
| 357 |
+
*Part of the BrainboxAI family of on-device models. See also [`law-il-E2B`](https://huggingface.co/BrainboxAI/law-il-E2B) (law) and [`cyber-analyst-4B`](https://huggingface.co/BrainboxAI/cyber-analyst-4B) (security).*
|