Instructions to use gszauer/Gab100M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use gszauer/Gab100M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="gszauer/Gab100M", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("gszauer/Gab100M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use gszauer/Gab100M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "gszauer/Gab100M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gszauer/Gab100M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/gszauer/Gab100M
- SGLang
How to use gszauer/Gab100M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "gszauer/Gab100M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gszauer/Gab100M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "gszauer/Gab100M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gszauer/Gab100M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use gszauer/Gab100M with Docker Model Runner:
docker model run hf.co/gszauer/Gab100M
Gab 100M
Try it in your browser: giftofgab.chat
Gab is a ~100 million parameter language model trained from scratch on a Mac Studio. It is a decoder-only transformer with pre-normalization (RMSNorm), rotary position embeddings, exact-GeLU feed-forward blocks with a hidden bias, causal self-attention, and tied input/output embeddings. The canonical model definition is reference.js (included in this repo) — a dependency-free JavaScript implementation the whole training suite is matched against.
It speaks English, holds short conversations, and supports an optional <think> reasoning trace inside assistant turns.
| Setting | Value |
|---|---|
| Parameters | 99,753,216 |
| Transformer blocks | 13 |
| Feature dimension | 768 |
| Attention heads | 12 |
| Head dimension | 64 |
| MLP hidden dimension | 3072 (4x, with a per-neuron bias) |
| Context length | 2048 |
| Vocabulary size | 10,000 |
| RoPE theta | 10,000 (interleaved pairs) |
| RMSNorm epsilon | 1e-5 |
| Precision | float32 |
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "gszauer/Gab100M"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)
messages = [{"role": "user", "content": "Hello, how are you?"}]
input_ids = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
# enable_thinking=True, # seed the reply with a <think> trace
return_tensors="pt",
)
output = model.generate(input_ids, max_new_tokens=128)
print(tokenizer.decode(output[0][input_ids.shape[1]:], skip_special_tokens=True))
The Transformers port is a simple, faithful implementation that recomputes the full sequence each generation step (no KV cache) — fine for a 100M model. The browser runtime and the training suite's inference.py both implement KV-cached generation.
Files
| File | What it is |
|---|---|
model.safetensors |
float32 weights with HF-style tensor names (per-head attention matrices fused, see below) |
gab100m.weights |
the native format: a headerless flat float32 buffer in reference.js parameters() order, loadable via deserializeFromArrayBuffer |
vocab.json |
the tokenizer: {reserved, merges} |
reference.js |
the canonical implementation — tensor library, autograd, tokenizer, model, KV cache, and AdamW trainer in dependency-free JavaScript |
modeling_gift_of_gab.py etc. |
the Transformers port (trust_remote_code) |
Architecture notes
Attention is trained per head. Each of the 12 heads owns its own Wq, Wk, Wv of shape [768, 64] and its own output projection Wo of shape [64, 768]. Heads are never concatenated — each head projects its result back up to 768 on its own and the per-head results are summed:
Attention(x) = sum_i( softmax(causal(q_i @ k_i.T / sqrt(64))) @ v_i @ Wo_i )
Summing per-head outputs is algebraically identical to a concat plus one big [768, 768] output projection, so model.safetensors ships fused [768, 768] tensors and the port runs the standard formulation — float32 results match the reference bit for bit. The unfused per-head matrices are preserved in gab100m.weights.
RoPE uses interleaved pairs. Adjacent channels (2i, 2i+1) rotate together with freq_i = 1/10000^(2i/64) — unlike Llama-style split-half RoPE. There is no position embedding table; position enters only through RoPE.
The MLP has a hidden bias and no gate. MLP(x) = GeLU(x @ Wup + b) @ Wdown with exact (erf) GeLU.
The unembedding is tied. logits = RMSNorm(h) @ E.T — no separate output head.
The tokenizer is byte-level BPE with sequential merge application. Ids 0–255 are raw bytes; every other id is a merge of two earlier ids, learned in order. Encoding applies every merge rule in sequence over the UTF-8 bytes. Reserved tokens (<|user|>, <|assistant|>, <|end|>, <|endoftext|>, <think>, </think>) are chained merges taught before training, so they encode to a single id and are atomic.
Chat format
No system role. Turns are concatenated directly:
<|user|>QUESTION<|end|><|assistant|>ANSWER<|end|>
Thinking is optional and lives inside the assistant turn:
<|user|>QUESTION<|end|><|assistant|><think>PRIVATE_REASONING</think>ANSWER<|end|>
The chat template supports enable_thinking=True to seed the reply with <think>.
Training
Pre-trained on 33,882,142,110 tokens (Gab100MPretrain plus Cosmopedia v2), then fine-tuned into a chat model on 1,016,529,229 tokens (Gab100MFinetune plus SmolTalk), learning only from assistant replies.
AdamW with linear warmup and cosine decay: peak learning rate 3e-4 (pre-train) / 3e-5 (fine-tune), betas 0.9/0.999, weight decay 0.01 (decoupled), global-norm gradient clip 1.0. Trained with MLX on a Mac Studio.
Limitations
A compact experimental model trained on a desktop machine. It can chat in English and follow simple instructions, but it is dumb by LLM standards — expect factual errors and quirks.
Links
- Try it out: giftofgab.chat
- Blog/site: gabormakesgames.com
- Downloads last month
- 22