Gab 100M

Try it in your browser: giftofgab.chat

Gab is a ~100 million parameter language model trained from scratch on a Mac Studio. It is a decoder-only transformer with pre-normalization (RMSNorm), rotary position embeddings, exact-GeLU feed-forward blocks with a hidden bias, causal self-attention, and tied input/output embeddings. The canonical model definition is reference.js (included in this repo) — a dependency-free JavaScript implementation the whole training suite is matched against.

It speaks English, holds short conversations, and supports an optional <think> reasoning trace inside assistant turns.

Setting Value
Parameters 99,753,216
Transformer blocks 13
Feature dimension 768
Attention heads 12
Head dimension 64
MLP hidden dimension 3072 (4x, with a per-neuron bias)
Context length 2048
Vocabulary size 10,000
RoPE theta 10,000 (interleaved pairs)
RMSNorm epsilon 1e-5
Precision float32

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "gszauer/Gab100M"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)

messages = [{"role": "user", "content": "Hello, how are you?"}]
input_ids = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    # enable_thinking=True,  # seed the reply with a <think> trace
    return_tensors="pt",
)
output = model.generate(input_ids, max_new_tokens=128)
print(tokenizer.decode(output[0][input_ids.shape[1]:], skip_special_tokens=True))

The Transformers port is a simple, faithful implementation that recomputes the full sequence each generation step (no KV cache) — fine for a 100M model. The browser runtime and the training suite's inference.py both implement KV-cached generation.

Files

File What it is
model.safetensors float32 weights with HF-style tensor names (per-head attention matrices fused, see below)
gab100m.weights the native format: a headerless flat float32 buffer in reference.js parameters() order, loadable via deserializeFromArrayBuffer
vocab.json the tokenizer: {reserved, merges}
reference.js the canonical implementation — tensor library, autograd, tokenizer, model, KV cache, and AdamW trainer in dependency-free JavaScript
modeling_gift_of_gab.py etc. the Transformers port (trust_remote_code)

Architecture notes

Attention is trained per head. Each of the 12 heads owns its own Wq, Wk, Wv of shape [768, 64] and its own output projection Wo of shape [64, 768]. Heads are never concatenated — each head projects its result back up to 768 on its own and the per-head results are summed:

Attention(x) = sum_i( softmax(causal(q_i @ k_i.T / sqrt(64))) @ v_i @ Wo_i )

Summing per-head outputs is algebraically identical to a concat plus one big [768, 768] output projection, so model.safetensors ships fused [768, 768] tensors and the port runs the standard formulation — float32 results match the reference bit for bit. The unfused per-head matrices are preserved in gab100m.weights.

RoPE uses interleaved pairs. Adjacent channels (2i, 2i+1) rotate together with freq_i = 1/10000^(2i/64) — unlike Llama-style split-half RoPE. There is no position embedding table; position enters only through RoPE.

The MLP has a hidden bias and no gate. MLP(x) = GeLU(x @ Wup + b) @ Wdown with exact (erf) GeLU.

The unembedding is tied. logits = RMSNorm(h) @ E.T — no separate output head.

The tokenizer is byte-level BPE with sequential merge application. Ids 0–255 are raw bytes; every other id is a merge of two earlier ids, learned in order. Encoding applies every merge rule in sequence over the UTF-8 bytes. Reserved tokens (<|user|>, <|assistant|>, <|end|>, <|endoftext|>, <think>, </think>) are chained merges taught before training, so they encode to a single id and are atomic.

Chat format

No system role. Turns are concatenated directly:

<|user|>QUESTION<|end|><|assistant|>ANSWER<|end|>

Thinking is optional and lives inside the assistant turn:

<|user|>QUESTION<|end|><|assistant|><think>PRIVATE_REASONING</think>ANSWER<|end|>

The chat template supports enable_thinking=True to seed the reply with <think>.

Training

Pre-trained on 33,882,142,110 tokens (Gab100MPretrain plus Cosmopedia v2), then fine-tuned into a chat model on 1,016,529,229 tokens (Gab100MFinetune plus SmolTalk), learning only from assistant replies.

AdamW with linear warmup and cosine decay: peak learning rate 3e-4 (pre-train) / 3e-5 (fine-tune), betas 0.9/0.999, weight decay 0.01 (decoupled), global-norm gradient clip 1.0. Trained with MLX on a Mac Studio.

Limitations

A compact experimental model trained on a desktop machine. It can chat in English and follow simple instructions, but it is dumb by LLM standards — expect factual errors and quirks.

Links

Downloads last month
22
Safetensors
Model size
99.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train gszauer/Gab100M