Instructions to use navthings/sprout with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use navthings/sprout with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf navthings/sprout:Q4_K_M # Run inference directly in the terminal: llama cli -hf navthings/sprout:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf navthings/sprout:Q4_K_M # Run inference directly in the terminal: llama cli -hf navthings/sprout:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf navthings/sprout:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf navthings/sprout:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf navthings/sprout:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf navthings/sprout:Q4_K_M
Use Docker
docker model run hf.co/navthings/sprout:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use navthings/sprout with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "navthings/sprout" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "navthings/sprout", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/navthings/sprout:Q4_K_M
- Ollama
How to use navthings/sprout with Ollama:
ollama run hf.co/navthings/sprout:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use navthings/sprout with Docker Model Runner:
docker model run hf.co/navthings/sprout:Q4_K_M
- Lemonade
How to use navthings/sprout with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull navthings/sprout:Q4_K_M
Run and chat with the model
lemonade run user.sprout-Q4_K_M
List all available models
lemonade list
- Atomic Chat
sprout
523m param llama-style model i trained from scratch on a free kaggle tpu, then finetuned into a chat model. its the bigger version of lilbase / lilchat.
it answers in full sentences, stops when its done, writes ok python and gets simple facts right a lot more than lilchat did (capital of australia is canberra now, not melbourne). still a small model tho. maths is a coin flip (17 + 25 was 42 once and 31 the next time), haikus are bad, and it will confidently make stuff up.
run it
ollama run navthings/sprout
or try it in your browser: https://navthings.github.io/playground/?model=sprout
files
| file | size | notes |
|---|---|---|
model.safetensors + config + tokenizer |
2.09gb | transformers LlamaForCausalLM (fp32), chat template included |
sprout-q8_0.gguf |
556mb | |
sprout-q4_k_m.gguf |
345mb | smallest, what the playground uses |
prompt format
llama-2 style, with </s> (id 2) as the start token and after every assistant reply:
</s>[INST] hi [/INST] Hello! How can I help you today?</s>[INST] next message [/INST]
system prompts go in <<SYS>>\n...\n<</SYS>>\n\n before the first [INST]. temperature 0.4 works well.
model
28 layers, d_model 1280, 20 query heads / 4 kv heads (gqa), head_dim 64, swiglu ffn 3456, rmsnorm, rope (theta 10000), tied embeddings, no biases. 1024 context. llama tokenizer (32k vocab).
pretraining
12b tokens (22,888 steps of 524k tokens) on a kaggle tpu v5e-8, data parallel over all 8 chips with the adam state sharded across them. ~130k tok/s, about 26 hours over 4 kaggle sessions.
data by tokens: fineweb-edu 55%, dclm 27%, cosmopedia v2 9%, project gutenberg 9%.
adamw (0.9/0.95, wd 0.1, clip 1.0), lr 3e-4 with 1,500 warmup steps, flat, then linear decay to 10% over the last 15%.
held-out loss at the end: fineweb-edu 2.387 (ppl 10.9), wikitext-103 2.723 (ppl 15.2). lilbase got 2.608 on the same fineweb-edu split.
finetune
full sft on smol-smoltalk, 2 epochs (4,702 steps of 131k tokens, 616m tokens), lr 1e-4, loss only on assistant turns. about 90 min on the same tpu. held-out loss went from 1.611 (plain sprout) to 0.955.
code
- Downloads last month
- 16
