EzioDevio commited on
Commit
4802bdc
Β·
verified Β·
1 Parent(s): f18b3b1

Upload folder using huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +60 -39
README.md CHANGED
@@ -1,54 +1,54 @@
1
- # ArduinoLLM-7B
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
 
3
- An open-source, LoRA-tuned Qwen2.5-Coder-7B specialist for embedded systems wiring and code generation β€” Arduino C++, Raspberry Pi Python, MicroPython, and CircuitPython, across 9 boards (ESP32, ESP32-S3, ESP32-C3, Raspberry Pi, Raspberry Pi Pico, Arduino Uno, ESP8266, Teensy 4.0, M5Stack Core2).
4
 
5
- Given a component and a board, it produces a wiring table, an ASCII wiring diagram, working code, and a short explanation β€” in one consistent format.
6
 
7
- ## v2: Corpus-Enhanced Training
8
 
9
- v1 was trained entirely on synthetic examples generated by LLMs (Claude, Gemini, and a locally-hosted Qwen2.5-Coder-32B). That approach is fast and controllable, but it has a real ceiling: a model trained only on another model's *description* of correct wiring can't be more accurate than the model that generated its training data.
10
 
11
- v2 adds a genuine second training stage: **continued pretraining directly on real library source code** β€” 240,634 real files (~2.59 billion tokens) from the actual Arduino core libraries, `micropython-lib`, and the Adafruit CircuitPython Bundle. This release trained on the first ~500 million tokens (~19% of the corpus) of that real code, then re-ran instruction fine-tuning on top to restore clean output formatting.
12
 
13
- ### Why this matters
14
 
15
- Manual testing of v1 found three concrete, reproducible failure modes:
16
- - A wiring diagram that internally contradicted its own wiring table
17
- - A fabricated, non-functional I2C address-conflict "fix" using invented register names
18
- - A weather-station answer that invented a CO2 reading from a sensor (BME280) that cannot measure CO2
19
-
20
- All three were re-tested after v2 training and no longer reproduce.
21
-
22
- ### Measured results (12-question held-out benchmark)
23
-
24
- | Metric | v1 (synthetic only) | v2 (corpus + synthetic) |
25
  |---|---|---|
26
  | Format compliance | 83% | **100%** |
27
- | Runtime-contamination-free | 100% | 100% |
28
- | Correct runtime API present | 92% | 92% |
29
- | Code syntax valid | ~100% | 100% |
30
- | Generation speed (RTX 5090) | 73.2 tok/s | 37.5 tok/s |
31
 
32
- **Known tradeoff**: v2 is currently slower than v1 due to being loaded from a locally re-merged checkpoint rather than a pre-optimized quantized repo. A faster-quantized export may follow.
33
 
34
- ## How it was built
35
 
36
- 1. **Dataset generation** β€” synthetic wiring/code examples generated via the Anthropic API, Google's Gemini API, and a locally-hosted Qwen2.5-Coder-32B (via vLLM), validated for format compliance and runtime contamination. 1,607 unique examples in `full_dataset.jsonl`.
37
- 2. **Continued pretraining** β€” `prepare_pretraining_corpus.py` chunks real library source into 2048-token blocks; `train_continued_pretrain.py` runs LoRA rank-32 training with `embed_tokens`/`lm_head` unlocked (needed to absorb new vocabulary/patterns from raw code, not just new behavior).
38
- 3. **Merge** β€” `merge_pretrain_checkpoint.py` combines the pretraining adapter into a standalone base model.
39
- 4. **Instruction fine-tuning** β€” `train_lora.py` re-runs standard LoRA fine-tuning (rank 16, attention/MLP only) on top of the corpus-enhanced base, using the same synthetic dataset as v1.
40
- 5. **Final merge** β€” `merge_final_release.py` combines both stages into the single published model.
41
- 6. **Evaluation** β€” `evaluate_model.py` runs the 12-question automated benchmark; manual testing checks specific documented failure modes via `ask_model.py`.
42
 
43
- ## Licensing note
 
 
44
 
45
- Built on Qwen2.5-Coder-7B-Instruct (Apache 2.0). Synthetic training data was generated using Anthropic's and Google's APIs, and a locally-run open-weight model β€” OpenAI and DeepSeek were deliberately not used for data generation, as both explicitly prohibit using their outputs to train competing models in their terms of service. The real-code corpus (Arduino core libraries, `micropython-lib`, CircuitPython Bundle) is used under each project's own open-source license.
46
-
47
- ## Limitations
48
-
49
- - Reliable on single-component, well-represented combinations; less reliable on genuinely novel multi-sensor combinations or troubleshooting scenarios not well-represented in the training data.
50
- - v2 trained on only ~19% of the available real-code corpus; the remaining ~81% has not yet been used.
51
- - Not a substitute for checking your own wiring against a component's actual datasheet, especially for voltage-sensitive components.
52
 
53
  ## Usage
54
 
@@ -61,7 +61,28 @@ model, tokenizer = FastLanguageModel.from_pretrained(
61
  load_in_4bit=True,
62
  )
63
  FastLanguageModel.for_inference(model)
 
 
 
 
 
64
  ```
65
 
66
- See `ask_model.py` for a full interactive example.
67
- - Character LCD displays (HD44780-style, in any interface variant β€” parallel, I2C backpack, or claimed SPI) are currently unreliable; wiring diagrams have shown pin fabrication and internal inconsistency with the accompanying code across multiple tests. Recommend manual verification against your specific module's datasheet for these displays specifically.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-Coder-7B-Instruct
4
+ tags:
5
+ - arduino
6
+ - esp32
7
+ - raspberry-pi
8
+ - micropython
9
+ - circuitpython
10
+ - embedded-systems
11
+ - code-generation
12
+ - lora
13
+ - continued-pretraining
14
+ language:
15
+ - en
16
+ library_name: transformers
17
+ ---
18
 
19
+ # ArduinoLLM-7B (v2)
20
 
21
+ A Qwen2.5-Coder-7B specialist for embedded systems wiring and code generation, covering Arduino C++, Raspberry Pi Python, MicroPython, and CircuitPython across 9 boards.
22
 
23
+ Given a component and a board, it generates a wiring table, an ASCII wiring diagram, working code, and a short explanation of the key design decision β€” in one consistent format.
24
 
25
+ ## What's new in v2
26
 
27
+ v1 was trained entirely on LLM-generated synthetic examples. v2 adds a genuine second training stage: **continued pretraining on real library source code** (~500M tokens from 240,634 real files across the Arduino core libraries, `micropython-lib`, and the Adafruit CircuitPython Bundle), followed by re-running instruction fine-tuning on top.
28
 
29
+ **Result**: format compliance improved from 83% to 100% on a 12-question held-out benchmark, and three specific, previously-documented failure modes (an internally-inconsistent wiring diagram, a fabricated I2C conflict-resolution procedure, and an invented sensor capability) no longer reproduce.
30
 
31
+ | Metric | v1 | v2 |
 
 
 
 
 
 
 
 
 
32
  |---|---|---|
33
  | Format compliance | 83% | **100%** |
34
+ | Contamination-free | 100% | 100% |
35
+ | Correct runtime API | 92% | 92% |
36
+ | Syntax valid | ~100% | 100% |
37
+ | Speed (RTX 5090) | 73.2 tok/s | 37.5 tok/s |
38
 
39
+ v2 currently runs slower than v1 β€” a known tradeoff from the local re-merge process, not a quality issue. A faster-quantized export may follow.
40
 
41
+ ## Post-release findings (ongoing)
42
 
43
+ After initial v2 publication, targeted fixes were applied for two mechanically-verified issues, now enforced by automated checks in `validate_dataset.py` rather than prompt instructions alone:
44
+ - ESP32/ESP8266-specific macros (e.g. `IRAM_ATTR`) appearing in AVR (Arduino Uno) code
45
+ - Resistive sensors (LDR, FSR, thermistor) wired without a required voltage-divider resistor
 
 
 
46
 
47
+ Further spot-testing has surfaced two additional, not-yet-fixed patterns worth knowing about before you rely on generated wiring:
48
+ - Occasional swapped SPI pin assignments on ESP8266, where hardware SPI pins are fixed in silicon but the generated wiring table sometimes assigns them incorrectly
49
+ - Character LCD displays (HD44780-style, in any interface variant β€” parallel, I2C backpack, or claimed SPI) are currently unreliable: wiring diagrams have shown fabricated pins and internal inconsistency with the accompanying code across multiple tests
50
 
51
+ These are documented here rather than left for users to discover. As with all outputs, verify pin assignments against your specific hardware's actual datasheet β€” especially for character LCDs until this is resolved.
 
 
 
 
 
 
52
 
53
  ## Usage
54
 
 
61
  load_in_4bit=True,
62
  )
63
  FastLanguageModel.for_inference(model)
64
+
65
+ prompt = "How do I wire a BME280 sensor to an ESP32 over I2C, with Arduino code?"
66
+ inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
67
+ outputs = model.generate(**inputs, max_new_tokens=1300)
68
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
69
  ```
70
 
71
+ ## Training data
72
+
73
+ - **Synthetic**: 1,634 unique wiring/code examples, generated via the Anthropic API, Google's Gemini API, and a locally-hosted Qwen2.5-Coder-32B, validated for format compliance, runtime contamination, AVR/ESP macro correctness, and resistive-sensor voltage-divider presence.
74
+ - **Real corpus**: ~500M tokens (of a 2.59B-token total) from real Arduino/MicroPython/CircuitPython library source code.
75
+
76
+ OpenAI and DeepSeek were not used, since both explicitly prohibit using their API output to train competing models.
77
+
78
+ ## Limitations
79
+
80
+ - Most reliable on well-represented single-component combinations.
81
+ - v2 used only ~19% of the available real-code corpus.
82
+ - Character LCD displays are currently unreliable (see Post-release findings above).
83
+ - Occasional SPI pin-assignment errors on ESP8266 have been observed.
84
+ - Always verify wiring against the component's actual datasheet before connecting power.
85
+
86
+ ## Full pipeline and dataset
87
+
88
+ See [github.com/EzioDEVio/ArduinoLLM-7B](https://github.com/EzioDEVio/ArduinoLLM-7B) for the complete training pipeline, dataset, and evaluation scripts.