CodeSprout 30M โ experimental code completion
An independently trained small decoder-only Transformer for HTML, JavaScript, Python and TypeScript. This is an early research checkpoint, not a general coding assistant or a chat model.
Actual training run
- Unique model parameters: 29,676,544.
- Architecture: 8 layers, width 512, 8 heads, 8,192-token byte-level BPE vocabulary, tied output/token weights.
- Context window: 512 tokens; training sequence length: 256.
- Hardware: NVIDIA GeForce RTX 3050 Laptop GPU.
- Selected inference checkpoint: step 4,324, after 17,711,104 sampled tokens.
- Final optimizer run reached step 6,079; later, more overfitted weights were not selected for inference.
- Elapsed training time, including checkpoints and evaluation: 0.47 hours.
- Peak allocated GPU memory: 788.3 MB.
Data and limitations
The starter corpus contains 812 exact-deduplicated files from pinned snapshots of MDN learning-area (CC0), Flask (BSD-3-Clause), Express (MIT), and TypeScript (Apache-2.0). It contains 732 training files and 80 held-out files. Files were split within the same repositories; validation is not repository-disjoint, and no claim of performance on unseen projects or coding benchmarks is made. Languages are sampled equally in training.
Full corpus sizes and pinned source revisions appear in data_report.json; original license texts are in the SOURCE_LICENSE_*.txt files. This release does not redistribute raw third-party training files. The small corpus can cause overfitting and memorized output.
After source-code pretraining, the model was aligned on original deterministic templates for small functions and HTML pages, plus about 10% source-token rehearsal. completion_data_report.json records their counts. These examples are not generated by a pretrained teacher. Their train/validation strings are disjoint, but their template families overlap. This makes the model most suitable for tutorial-sized code completion, not unfamiliar complex projects.
Validation
The following final next-token cross-entropy values are on completion-template variations plus source rehearsal, estimated from three fixed sampled batches per language. They are not losses on the original source-only validation split and must not be compared directly with that earlier stage:
| Language | Validation loss |
|---|---|
| html | 0.3679 |
| javascript | 1.2808 |
| python | 1.4414 |
| typescript | 0.8401 |
These measurements are language-modeling losses, not functional coding accuracy. samples.json contains unedited generated continuations through the supported completion wrapper. evaluation.json includes a tiny sanity check on trained arithmetic task families and balanced HTML tags, including the failures. This is not an unseen-task benchmark. Generated output may be syntactically invalid, incorrect, insecure or incomplete; inspect it before using it. This model was not chat/instruction-tuned and does not execute code.
The completion wrapper normalizes trailing prefix whitespace before tokenization and restores the matching whitespace in the continuation. This avoids a byte-level BPE boundary issue around indentation; use the supplied wrapper rather than blindly tokenizing a prefix ending in a newline.
Load and use
Install the supplied requirements and place the inference files together. generate.py loads tensors with torch.load(..., weights_only=True). The custom architecture is defined in model.py; this is not currently an AutoModel-compatible Transformers checkpoint.
from pathlib import Path
from generate import load, complete
model, tokenizer = load(Path("."), "cpu")
print(complete(model, tokenizer, "def add(a, b):\n", "python"))
The Gradio demo is app.py. No external API, paid inference provider or laptop connection is required once deployed. Generation is limited to 120 new tokens per request.
License
Model and project code are released under Apache-2.0. This does not replace the licenses or attribution requirements of the upstream training projects. Their license texts are included separately.
- Downloads last month
- 13