File size: 2,471 Bytes
bd8054a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3ee8b4c
 
bd8054a
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
---

title: README
emoji: 🧠
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
---


# TinyBrainBot

Small language models trained from scratch on limited hardware, released with the details of how they were made: the data, the training runs, and the things that went wrong along the way.

**Start here:** [TinyBrainBot-350M-v4-Thinking](https://huggingface.co/nkthebass/tinybrainbot-350m-v4-thinking) thinks through a question before it answers, and replies to small talk directly. A GGUF for llama.cpp, LM Studio and Ollama is included.

**Try it now:** [TinyBrainBot 100M Thinking demo](https://huggingface.co/spaces/nkthebass/tinybrainbot-100m-thinking) runs the 100M thinking model right in your browser, no install needed.

## Models

| Series | Models | Notes |
|---|---|---|
| [v4](https://huggingface.co/collections/nkthebass/tinybrainbot-v4-6ac14ac6a9a6a6806ba4a707) | [350M Thinking](https://huggingface.co/nkthebass/tinybrainbot-350m-v4-thinking), [100M Thinking](https://huggingface.co/nkthebass/tinybrainbot-100m-v4-thinking), [25M Base](https://huggingface.co/nkthebass/tinybrainbot-25m-v4-base) | Newest. Thinking models that reason inside `<think>` before answering, and a 25M base that beats Pythia-70M on 10 of 13 benchmarks |
| [v3](https://huggingface.co/collections/nkthebass/tinybrainbot-v3-6ac14ac4b73470ecf713992f) | 350M and 100M base, instruct and math | The 100M beats Supra2-100M-Instruct on 6 of 7 benchmarks |
| [v2](https://huggingface.co/collections/nkthebass/tinybrainbot-v2-6ac14ac3659e8a86ec3923cd) | 320M and 303M base, instruct and math | The 320M math model beats GPT-3 175B on 4 to 5 digit arithmetic |
| [v1](https://huggingface.co/collections/nkthebass/tinybrainbot-v1-6ac14ac38f73e742da1998dd) | 303M base and instruct, 216M demo | The first models |

## Research

**[Do data-mix rankings transfer across scale?](https://huggingface.co/nkthebass/tinybrainbot-pilot-100m-datamix)** A controlled study at 100M parameters: three models that are identical except for their training data, measured on held-out text, seven domain slices and fifteen benchmarks.

**[One recipe, five sizes](https://huggingface.co/nkthebass/tinybrainbot-v4-scaling-500k-25m)** Five models from 500K to 25M parameters trained on exactly the same 11B tokens, so the only difference is size. The 10M averages above Pythia-70M.

All models are currently published from [nkthebass](https://huggingface.co/nkthebass).