README / README.md
nkthebass's picture
Add the scaling study
3ee8b4c verified
|
Raw History Blame Contribute Delete
2.47 kB
---
title: README
emoji: 🧠
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
---
# TinyBrainBot
Small language models trained from scratch on limited hardware, released with the details of how they were made: the data, the training runs, and the things that went wrong along the way.
**Start here:** [TinyBrainBot-350M-v4-Thinking](https://huggingface.co/nkthebass/tinybrainbot-350m-v4-thinking) thinks through a question before it answers, and replies to small talk directly. A GGUF for llama.cpp, LM Studio and Ollama is included.
**Try it now:** [TinyBrainBot 100M Thinking demo](https://huggingface.co/spaces/nkthebass/tinybrainbot-100m-thinking) runs the 100M thinking model right in your browser, no install needed.
## Models
| Series | Models | Notes |
|---|---|---|
| [v4](https://huggingface.co/collections/nkthebass/tinybrainbot-v4-6ac14ac6a9a6a6806ba4a707) | [350M Thinking](https://huggingface.co/nkthebass/tinybrainbot-350m-v4-thinking), [100M Thinking](https://huggingface.co/nkthebass/tinybrainbot-100m-v4-thinking), [25M Base](https://huggingface.co/nkthebass/tinybrainbot-25m-v4-base) | Newest. Thinking models that reason inside `<think>` before answering, and a 25M base that beats Pythia-70M on 10 of 13 benchmarks |
| [v3](https://huggingface.co/collections/nkthebass/tinybrainbot-v3-6ac14ac4b73470ecf713992f) | 350M and 100M base, instruct and math | The 100M beats Supra2-100M-Instruct on 6 of 7 benchmarks |
| [v2](https://huggingface.co/collections/nkthebass/tinybrainbot-v2-6ac14ac3659e8a86ec3923cd) | 320M and 303M base, instruct and math | The 320M math model beats GPT-3 175B on 4 to 5 digit arithmetic |
| [v1](https://huggingface.co/collections/nkthebass/tinybrainbot-v1-6ac14ac38f73e742da1998dd) | 303M base and instruct, 216M demo | The first models |
## Research
**[Do data-mix rankings transfer across scale?](https://huggingface.co/nkthebass/tinybrainbot-pilot-100m-datamix)** A controlled study at 100M parameters: three models that are identical except for their training data, measured on held-out text, seven domain slices and fifteen benchmarks.
**[One recipe, five sizes](https://huggingface.co/nkthebass/tinybrainbot-v4-scaling-500k-25m)** Five models from 500K to 25M parameters trained on exactly the same 11B tokens, so the only difference is size. The 10M averages above Pythia-70M.
All models are currently published from [nkthebass](https://huggingface.co/nkthebass).