EzioDevio commited on
Commit
a67baf3
·
verified ·
1 Parent(s): 86fd8f6

Upload folder using huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +34 -51
README.md CHANGED
@@ -1,56 +1,42 @@
1
  ---
2
  license: apache-2.0
3
- base_model: unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit
4
  tags:
5
  - arduino
6
  - esp32
7
  - raspberry-pi
8
  - micropython
9
  - circuitpython
10
- - embedded
 
11
  - lora
12
- - unsloth
13
  language:
14
  - en
15
- pipeline_tag: text-generation
16
  ---
17
 
18
- # Arduino/Embedded Wiring & Code Assistant (Qwen2.5-Coder-7B LoRA)
19
 
20
- A LoRA fine-tune of [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct), specialized for generating wiring instructions and working code for embedded/maker projects.
21
 
22
- Given a plain-language request like *"wire a BME280 to an ESP32 over I2C"*, it produces a wiring table, an ASCII wiring diagram, complete working code, and a short explanation — across Arduino C++, Raspberry Pi Python, MicroPython, and CircuitPython, on 7 boards (ESP32, Raspberry Pi, Raspberry Pi Pico, Arduino Uno, ESP8266, Teensy 4.0, M5Stack Core2).
23
 
24
- **Full training pipeline, dataset, and evaluation harness:** [GitHub repo](https://github.com/YOUR-USERNAME/arduino-embedded-llm)
25
 
26
- ## Benchmark (12-question held-out eval, deterministic decoding)
27
 
28
- | Metric | This model (7B, 373 examples) | Earlier 1.5B baseline (65 examples) |
29
- |---|---|---|
30
- | Format compliance | **83%** | 58% |
31
- | Runtime contamination-free | **100%** | 83% |
32
- | Code syntax valid | **~100%*** | 92% |
33
- | Speed (RTX 4070 8GB laptop) | 12.0 tok/s | 13.7 tok/s |
34
- | Speed (data-center GPU) | 73.2 tok/s | — |
35
-
36
- \* Two apparent syntax failures were confirmed to be 900-token generation-cap truncation, not real errors.
37
-
38
- "Runtime contamination" = wrong-framework API leaking into an answer (e.g. MicroPython's `machine.Pin` appearing in Arduino code, or Raspberry Pi's physical-pin numbering convention appearing in an ESP32 answer). Eliminating this was the primary goal of moving from the 1.5B to the 7B model plus a larger, validated dataset.
39
-
40
- We also tested a LoRA rank-32 variant: no measurable quality improvement over rank-16 on any metric, so this release ships the more efficient rank-16 adapter.
41
 
42
- ## ⚠️ Known limitations
43
-
44
- This is a v0.1 release trained on 373 examples — small by any standard. Manual testing (6 hand-picked questions, not from the eval set) found a clear, honest capability boundary:
45
-
46
- **Reliable (3/3 tested):** straightforward single-component wiring closely resembling the training distribution — e.g. "wire a BME280 to an ESP32 over I2C," "read a DHT22 on a Raspberry Pi," "wire an SSD1306 OLED to an ESP32 in MicroPython." All three produced correct pins, correct libraries, and internally consistent answers.
47
-
48
- **Unreliable (0/3 tested):** novel multi-component combinations and protocol-level troubleshooting outside the training distribution. Three distinct hallucination patterns were found:
49
- - Internal inconsistencies within one answer (a wiring diagram wire contradicting its own table)
50
- - A fabricated, non-functional low-level procedure for an I2C address conflict (invented register names/values presented with confidence)
51
- - An invented sensor capability not physically possible on the requested hardware (a "CO2 estimate" from a sensor that cannot measure CO2)
52
 
53
- **Always verify wiring against the actual component datasheet before connecting real hardware.** Incorrect voltage or pin assignments can damage components. Treat outputs as a first draft, not a final instruction set.
54
 
55
  ## Usage
56
 
@@ -58,34 +44,31 @@ This is a v0.1 release trained on 373 examples — small by any standard. Manual
58
  from unsloth import FastLanguageModel
59
 
60
  model, tokenizer = FastLanguageModel.from_pretrained(
61
- "unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit",
62
  max_seq_length=2048,
63
  load_in_4bit=True,
64
  )
65
- model.load_adapter("YOUR-USERNAME/arduino-embedded-qwen2.5-coder-7b")
66
  FastLanguageModel.for_inference(model)
67
 
68
- messages = [{"role": "user", "content": "How do I wire a BME280 to an ESP32 over I2C?"}]
69
- inputs = tokenizer.apply_chat_template(
70
- messages, tokenize=True, add_generation_prompt=True, return_tensors="pt"
71
- ).to("cuda")
72
- outputs = model.generate(input_ids=inputs, max_new_tokens=1300, do_sample=False)
73
- print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
74
  ```
75
 
76
- Minimum hardware: ~6GB VRAM for inference (tested at 8GB on an RTX 4070 laptop). A 0.5B distilled variant for CPU/Raspberry Pi Zero 2W deployment is planned.
77
 
78
- ## Training details
 
79
 
80
- - **Base model:** Qwen2.5-Coder-7B-Instruct (4-bit, via Unsloth)
81
- - **Method:** LoRA, rank 16, alpha 16, all attention + MLP projections
82
- - **Dataset:** 373 examples, synthetically generated and validated against an automated contamination checker before merging into training data
83
- - **Epochs:** 3
84
 
85
- ## License
86
 
87
- Apache 2.0, inherited from the base model. Free for commercial use, modification, and redistribution.
 
 
88
 
89
- ## Citation / Acknowledgments
90
 
91
- Built on [Qwen2.5-Coder](https://github.com/QwenLM/Qwen2.5-Coder) (Alibaba Qwen team) using [Unsloth](https://github.com/unslothai/unsloth).
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-Coder-7B-Instruct
4
  tags:
5
  - arduino
6
  - esp32
7
  - raspberry-pi
8
  - micropython
9
  - circuitpython
10
+ - embedded-systems
11
+ - code-generation
12
  - lora
13
+ - continued-pretraining
14
  language:
15
  - en
16
+ library_name: transformers
17
  ---
18
 
19
+ # ArduinoLLM-7B (v2)
20
 
21
+ A Qwen2.5-Coder-7B specialist for embedded systems wiring and code generation, covering Arduino C++, Raspberry Pi Python, MicroPython, and CircuitPython across 9 boards.
22
 
23
+ Given a component and a board, it generates a wiring table, an ASCII wiring diagram, working code, and a short explanation of the key design decision — in one consistent format.
24
 
25
+ ## What's new in v2
26
 
27
+ v1 was trained entirely on LLM-generated synthetic examples. v2 adds a genuine second training stage: **continued pretraining on real library source code** (~500M tokens from 240,634 real files across the Arduino core libraries, `micropython-lib`, and the Adafruit CircuitPython Bundle), followed by re-running instruction fine-tuning on top.
28
 
29
+ **Result**: format compliance improved from 83% to 100% on a 12-question held-out benchmark, and three specific, previously-documented failure modes (an internally-inconsistent wiring diagram, a fabricated I2C conflict-resolution procedure, and an invented sensor capability) no longer reproduce.
 
 
 
 
 
 
 
 
 
 
 
 
30
 
31
+ | Metric | v1 | v2 |
32
+ |---|---|---|
33
+ | Format compliance | 83% | **100%** |
34
+ | Contamination-free | 100% | 100% |
35
+ | Correct runtime API | 92% | 92% |
36
+ | Syntax valid | ~100% | 100% |
37
+ | Speed (RTX 5090) | 73.2 tok/s | 37.5 tok/s |
 
 
 
38
 
39
+ v2 currently runs slower than v1 — a known tradeoff from the local re-merge process, not a quality issue. A faster-quantized export may follow.
40
 
41
  ## Usage
42
 
 
44
  from unsloth import FastLanguageModel
45
 
46
  model, tokenizer = FastLanguageModel.from_pretrained(
47
+ model_name="EzioDevio/ArduinoLLM-7B",
48
  max_seq_length=2048,
49
  load_in_4bit=True,
50
  )
 
51
  FastLanguageModel.for_inference(model)
52
 
53
+ prompt = "How do I wire a BME280 sensor to an ESP32 over I2C, with Arduino code?"
54
+ inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
55
+ outputs = model.generate(**inputs, max_new_tokens=1300)
56
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
 
 
57
  ```
58
 
59
+ ## Training data
60
 
61
+ - **Synthetic**: 1,607 unique wiring/code examples, generated via the Anthropic API, Google's Gemini API, and a locally-hosted Qwen2.5-Coder-32B, validated for format compliance and runtime contamination.
62
+ - **Real corpus**: ~500M tokens (of a 2.59B-token total) from real Arduino/MicroPython/CircuitPython library source code.
63
 
64
+ OpenAI and DeepSeek were not used, since both explicitly prohibit using their API output to train competing models.
 
 
 
65
 
66
+ ## Limitations
67
 
68
+ - Most reliable on well-represented single-component combinations.
69
+ - v2 used only ~19% of the available real-code corpus.
70
+ - Always verify wiring against the component's actual datasheet before connecting power.
71
 
72
+ ## Full pipeline and dataset
73
 
74
+ See [github.com/EzioDEVio/ArduinoLLM-7B](https://github.com/EzioDEVio/ArduinoLLM-7B) for the complete training pipeline, dataset, and evaluation scripts.