Text Generation
Transformers
Safetensors
English
qwen2
arduino
esp32
raspberry-pi
micropython
circuitpython
embedded-systems
code-generation
lora
continued-pretraining
conversational
text-generation-inference
Instructions to use EzioDevio/ArduinoLLM-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use EzioDevio/ArduinoLLM-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="EzioDevio/ArduinoLLM-7B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("EzioDevio/ArduinoLLM-7B") model = AutoModelForCausalLM.from_pretrained("EzioDevio/ArduinoLLM-7B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use EzioDevio/ArduinoLLM-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EzioDevio/ArduinoLLM-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EzioDevio/ArduinoLLM-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/EzioDevio/ArduinoLLM-7B
- SGLang
How to use EzioDevio/ArduinoLLM-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "EzioDevio/ArduinoLLM-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EzioDevio/ArduinoLLM-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "EzioDevio/ArduinoLLM-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EzioDevio/ArduinoLLM-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use EzioDevio/ArduinoLLM-7B with Docker Model Runner:
docker model run hf.co/EzioDevio/ArduinoLLM-7B
Upload folder using huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,56 +1,42 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
-
base_model:
|
| 4 |
tags:
|
| 5 |
- arduino
|
| 6 |
- esp32
|
| 7 |
- raspberry-pi
|
| 8 |
- micropython
|
| 9 |
- circuitpython
|
| 10 |
-
- embedded
|
|
|
|
| 11 |
- lora
|
| 12 |
-
-
|
| 13 |
language:
|
| 14 |
- en
|
| 15 |
-
|
| 16 |
---
|
| 17 |
|
| 18 |
-
#
|
| 19 |
|
| 20 |
-
A
|
| 21 |
|
| 22 |
-
Given a
|
| 23 |
|
| 24 |
-
|
| 25 |
|
| 26 |
-
|
| 27 |
|
| 28 |
-
|
| 29 |
-
|---|---|---|
|
| 30 |
-
| Format compliance | **83%** | 58% |
|
| 31 |
-
| Runtime contamination-free | **100%** | 83% |
|
| 32 |
-
| Code syntax valid | **~100%*** | 92% |
|
| 33 |
-
| Speed (RTX 4070 8GB laptop) | 12.0 tok/s | 13.7 tok/s |
|
| 34 |
-
| Speed (data-center GPU) | 73.2 tok/s | — |
|
| 35 |
-
|
| 36 |
-
\* Two apparent syntax failures were confirmed to be 900-token generation-cap truncation, not real errors.
|
| 37 |
-
|
| 38 |
-
"Runtime contamination" = wrong-framework API leaking into an answer (e.g. MicroPython's `machine.Pin` appearing in Arduino code, or Raspberry Pi's physical-pin numbering convention appearing in an ESP32 answer). Eliminating this was the primary goal of moving from the 1.5B to the 7B model plus a larger, validated dataset.
|
| 39 |
-
|
| 40 |
-
We also tested a LoRA rank-32 variant: no measurable quality improvement over rank-16 on any metric, so this release ships the more efficient rank-16 adapter.
|
| 41 |
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
- Internal inconsistencies within one answer (a wiring diagram wire contradicting its own table)
|
| 50 |
-
- A fabricated, non-functional low-level procedure for an I2C address conflict (invented register names/values presented with confidence)
|
| 51 |
-
- An invented sensor capability not physically possible on the requested hardware (a "CO2 estimate" from a sensor that cannot measure CO2)
|
| 52 |
|
| 53 |
-
|
| 54 |
|
| 55 |
## Usage
|
| 56 |
|
|
@@ -58,34 +44,31 @@ This is a v0.1 release trained on 373 examples — small by any standard. Manual
|
|
| 58 |
from unsloth import FastLanguageModel
|
| 59 |
|
| 60 |
model, tokenizer = FastLanguageModel.from_pretrained(
|
| 61 |
-
"
|
| 62 |
max_seq_length=2048,
|
| 63 |
load_in_4bit=True,
|
| 64 |
)
|
| 65 |
-
model.load_adapter("YOUR-USERNAME/arduino-embedded-qwen2.5-coder-7b")
|
| 66 |
FastLanguageModel.for_inference(model)
|
| 67 |
|
| 68 |
-
|
| 69 |
-
inputs = tokenizer.
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
outputs = model.generate(input_ids=inputs, max_new_tokens=1300, do_sample=False)
|
| 73 |
-
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
|
| 74 |
```
|
| 75 |
|
| 76 |
-
|
| 77 |
|
| 78 |
-
|
|
|
|
| 79 |
|
| 80 |
-
|
| 81 |
-
- **Method:** LoRA, rank 16, alpha 16, all attention + MLP projections
|
| 82 |
-
- **Dataset:** 373 examples, synthetically generated and validated against an automated contamination checker before merging into training data
|
| 83 |
-
- **Epochs:** 3
|
| 84 |
|
| 85 |
-
##
|
| 86 |
|
| 87 |
-
|
|
|
|
|
|
|
| 88 |
|
| 89 |
-
##
|
| 90 |
|
| 91 |
-
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
|
| 4 |
tags:
|
| 5 |
- arduino
|
| 6 |
- esp32
|
| 7 |
- raspberry-pi
|
| 8 |
- micropython
|
| 9 |
- circuitpython
|
| 10 |
+
- embedded-systems
|
| 11 |
+
- code-generation
|
| 12 |
- lora
|
| 13 |
+
- continued-pretraining
|
| 14 |
language:
|
| 15 |
- en
|
| 16 |
+
library_name: transformers
|
| 17 |
---
|
| 18 |
|
| 19 |
+
# ArduinoLLM-7B (v2)
|
| 20 |
|
| 21 |
+
A Qwen2.5-Coder-7B specialist for embedded systems wiring and code generation, covering Arduino C++, Raspberry Pi Python, MicroPython, and CircuitPython across 9 boards.
|
| 22 |
|
| 23 |
+
Given a component and a board, it generates a wiring table, an ASCII wiring diagram, working code, and a short explanation of the key design decision — in one consistent format.
|
| 24 |
|
| 25 |
+
## What's new in v2
|
| 26 |
|
| 27 |
+
v1 was trained entirely on LLM-generated synthetic examples. v2 adds a genuine second training stage: **continued pretraining on real library source code** (~500M tokens from 240,634 real files across the Arduino core libraries, `micropython-lib`, and the Adafruit CircuitPython Bundle), followed by re-running instruction fine-tuning on top.
|
| 28 |
|
| 29 |
+
**Result**: format compliance improved from 83% to 100% on a 12-question held-out benchmark, and three specific, previously-documented failure modes (an internally-inconsistent wiring diagram, a fabricated I2C conflict-resolution procedure, and an invented sensor capability) no longer reproduce.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
|
| 31 |
+
| Metric | v1 | v2 |
|
| 32 |
+
|---|---|---|
|
| 33 |
+
| Format compliance | 83% | **100%** |
|
| 34 |
+
| Contamination-free | 100% | 100% |
|
| 35 |
+
| Correct runtime API | 92% | 92% |
|
| 36 |
+
| Syntax valid | ~100% | 100% |
|
| 37 |
+
| Speed (RTX 5090) | 73.2 tok/s | 37.5 tok/s |
|
|
|
|
|
|
|
|
|
|
| 38 |
|
| 39 |
+
v2 currently runs slower than v1 — a known tradeoff from the local re-merge process, not a quality issue. A faster-quantized export may follow.
|
| 40 |
|
| 41 |
## Usage
|
| 42 |
|
|
|
|
| 44 |
from unsloth import FastLanguageModel
|
| 45 |
|
| 46 |
model, tokenizer = FastLanguageModel.from_pretrained(
|
| 47 |
+
model_name="EzioDevio/ArduinoLLM-7B",
|
| 48 |
max_seq_length=2048,
|
| 49 |
load_in_4bit=True,
|
| 50 |
)
|
|
|
|
| 51 |
FastLanguageModel.for_inference(model)
|
| 52 |
|
| 53 |
+
prompt = "How do I wire a BME280 sensor to an ESP32 over I2C, with Arduino code?"
|
| 54 |
+
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
|
| 55 |
+
outputs = model.generate(**inputs, max_new_tokens=1300)
|
| 56 |
+
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
|
|
|
|
|
|
| 57 |
```
|
| 58 |
|
| 59 |
+
## Training data
|
| 60 |
|
| 61 |
+
- **Synthetic**: 1,607 unique wiring/code examples, generated via the Anthropic API, Google's Gemini API, and a locally-hosted Qwen2.5-Coder-32B, validated for format compliance and runtime contamination.
|
| 62 |
+
- **Real corpus**: ~500M tokens (of a 2.59B-token total) from real Arduino/MicroPython/CircuitPython library source code.
|
| 63 |
|
| 64 |
+
OpenAI and DeepSeek were not used, since both explicitly prohibit using their API output to train competing models.
|
|
|
|
|
|
|
|
|
|
| 65 |
|
| 66 |
+
## Limitations
|
| 67 |
|
| 68 |
+
- Most reliable on well-represented single-component combinations.
|
| 69 |
+
- v2 used only ~19% of the available real-code corpus.
|
| 70 |
+
- Always verify wiring against the component's actual datasheet before connecting power.
|
| 71 |
|
| 72 |
+
## Full pipeline and dataset
|
| 73 |
|
| 74 |
+
See [github.com/EzioDEVio/ArduinoLLM-7B](https://github.com/EzioDEVio/ArduinoLLM-7B) for the complete training pipeline, dataset, and evaluation scripts.
|