sprout

523m param llama-style model i trained from scratch on a free kaggle tpu, then finetuned into a chat model. its the bigger version of lilbase / lilchat.

held-out loss while pretraining

it answers in full sentences, stops when its done, writes ok python and gets simple facts right a lot more than lilchat did (capital of australia is canberra now, not melbourne). still a small model tho. maths is a coin flip (17 + 25 was 42 once and 31 the next time), haikus are bad, and it will confidently make stuff up.

run it

ollama run navthings/sprout

or try it in your browser: https://navthings.github.io/playground/?model=sprout

files

file size notes
model.safetensors + config + tokenizer 2.09gb transformers LlamaForCausalLM (fp32), chat template included
sprout-q8_0.gguf 556mb
sprout-q4_k_m.gguf 345mb smallest, what the playground uses

prompt format

llama-2 style, with </s> (id 2) as the start token and after every assistant reply:

</s>[INST] hi [/INST] Hello! How can I help you today?</s>[INST] next message [/INST]

system prompts go in <<SYS>>\n...\n<</SYS>>\n\n before the first [INST]. temperature 0.4 works well.

model

28 layers, d_model 1280, 20 query heads / 4 kv heads (gqa), head_dim 64, swiglu ffn 3456, rmsnorm, rope (theta 10000), tied embeddings, no biases. 1024 context. llama tokenizer (32k vocab).

pretraining

12b tokens (22,888 steps of 524k tokens) on a kaggle tpu v5e-8, data parallel over all 8 chips with the adam state sharded across them. ~130k tok/s, about 26 hours over 4 kaggle sessions.

data by tokens: fineweb-edu 55%, dclm 27%, cosmopedia v2 9%, project gutenberg 9%.

adamw (0.9/0.95, wd 0.1, clip 1.0), lr 3e-4 with 1,500 warmup steps, flat, then linear decay to 10% over the last 15%.

held-out loss at the end: fineweb-edu 2.387 (ppl 10.9), wikitext-103 2.723 (ppl 15.2). lilbase got 2.608 on the same fineweb-edu split.

finetune

full sft on smol-smoltalk, 2 epochs (4,702 steps of 131k tokens, 616m tokens), lr 1e-4, loss only on assistant turns. about 90 min on the same tpu. held-out loss went from 1.611 (plain sprout) to 0.955.

code

https://github.com/navthings/sprout

Downloads last month
16
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train navthings/sprout