NanoChat-31M-V2
A ~31.4M parameter GPT-style decoder-only transformer, trained fully from scratch, that generates simulated live-stream chat comments -- both regular fast-paced "hype chat" and rarer, longer "superchat"-style paid messages/questions. Built to run fast on CPU, as a lightweight audience-simulation layer for AI VTuber / livestream projects.
v2 fixes v1's main complaint (garbled spelling) via a larger tokenizer vocabulary and new generation-time repetition controls -- see "What changed from v1" below.
Model details
- Architecture: custom GPT-style decoder-only transformer (not a registered
HF architecture -- see
model.pyin this repo for the class definition) - Parameters: 31,368,960 (~31.4M) -- up from v1's 27.8M due to a larger vocabulary
- Layers: 10 | Hidden size: 448 | Attention heads: 8 | FFN size: 1792
- Context length: 128 tokens
- Tokenizer: byte-level BPE, vocab size 16,000 (v1 was 8,000 -- see below)
- Weight tying: input/output embeddings are tied
- Inference: KV-cached autoregressive generation; fp32 recommended over int8 quantization at this model scale (quantization overhead outweighs its benefit once KV-caching is in use -- verified by direct benchmarking, not assumed)
What changed from v1
v1's biggest issue was garbled spelling -- individual words falling apart into nonsense fragments, even though grammar/structure was fine. Two changes targeted this:
- Vocabulary doubled (8,000 → 16,000 tokens). With a small vocabulary, common words get chopped into several sub-word pieces, and the model has to predict every piece correctly in sequence to spell a word right -- more pieces means more chances to go wrong. A larger vocabulary means more whole words get a single clean token instead of being fragmented, directly reducing the failure mode v1 had.
- Generation-time fixes in
generate():min_new_tokens: when a prompt is seeded with a topic keyword (see "relevance" below), the model would sometimes just echo the keyword back and immediately stop, since short 2-3 word messages are also completely normal, valid chat on their own. This forces a minimum number of new tokens before the model is allowed to end the message, when a keyword seed is used.repetition_penalty(default 1.2) andno_repeat_ngram_size(default 3): standard techniques that discourage the model from getting stuck in repetition loops (e.g. "the first game...the first game...the first"). Meaningfully reduces but does not 100% eliminate this failure mode -- occasional near-duplicate (not exact-duplicate) repeats can still slip through.
Training was also extended to 200,000 steps (from v1's 60,000), on a freshly-retrained tokenizer, so this is a from-scratch v2 model, not a fine-tune of v1.
Honest note on evaluation: validation loss on held-out real chat data is not directly comparable between v1 (3.90, vocab 8,000) and v2 (4.10, vocab 16,000) -- a larger vocabulary is an inherently harder per-token prediction task (more possible next tokens to choose from), which raises the loss floor independent of model quality. Normalizing against each model's own random-guess baseline, v2 actually edges out v1 slightly (beats its random baseline by ~5.58 nats vs. v1's ~5.09). But the real signal is qualitative: side-by-side generation testing showed v2 producing meaningfully fewer garbled words and more coherent full sentences than v1, which is the actual goal here.
Two generation modes
<chat>-- short, high-volume, low-effort hype comments ("LMAOOO", "W stream", emote spam), matching typical live chat cadence.<superchat>-- rarer, longer, more coherent messages: questions, personal questions, playful roasts, jokes, and support messages, matching the style of paid superchat/membership messages that streamers are expected to acknowledge.
Usage
import torch
from tokenizers import Tokenizer
from safetensors.torch import load_model
from model import ChatGPTMini, ModelConfig # model.py included in this repo
tokenizer = Tokenizer.from_file("tokenizer.json")
cfg = ModelConfig(vocab_size=16000, pad_token_id=tokenizer.token_to_id("<pad>"))
model = ChatGPTMini(cfg)
load_model(model, "model.safetensors")
model.eval()
bos_id = tokenizer.token_to_id("<bos>")
chat_id = tokenizer.token_to_id("<chat>")
eos_id = tokenizer.token_to_id("<eos>")
prompt = torch.tensor([[bos_id, chat_id]])
out = model.generate(prompt, max_new_tokens=30, temperature=0.9, top_k=40, top_p=0.9, eos_token_id=eos_id)
print(tokenizer.decode(out[0].tolist(), skip_special_tokens=True))
See example.py in this repo for a complete runnable script, including seeded
generation with min_new_tokens and batched generation (recommended for real usage).
Training data
- Regular chat: primarily real, public Twitch chat logs from the
lparkourer10/twitch_chatdataset (CC-BY-SA-4.0, usernames stripped), supplemented with a small amount of synthetic hype-phrase data. - Superchat: hand-written and template-generated (questions, personal questions, playful roasts, jokes), since no public dataset of real superchat/membership messages exists. This class is intentionally imbalanced relative to chat volume (~700 unique examples vs. millions of chat rows) and was oversampled during training via weighted sampling, not duplication.
Relevance / topic conditioning
The model itself has no context input -- it only knows "chat mode" vs. "superchat
mode." Topical relevance to a live stream is handled externally: a keyword is extracted
from the streamer's recent transcript (simple frequency-based extraction, no ML) and
occasionally used to seed the generation prompt. See keyword_utils.py in the training
repo (not included in this weights-only release) for the reference implementation.
Intended use
Generating background "audience chat" text for livestream/VTuber simulation projects, demos, or similar creative/entertainment use cases where large volumes of low-effort, stylistically plausible chat text are needed cheaply and quickly. Not intended as a factual, instruction-following, or general-purpose assistant model.
Limitations
- Small model, narrow training objective -- will still occasionally produce garbled text or repetition loops, though both are reduced from v1. This is generally visually consistent with how noisy real chat already looks, but shouldn't be mistaken for a general-purpose language model.
- Superchat generation may lean close to its (relatively small) training examples rather than fully generalizing.
- No built-in context/topic conditioning -- see "Relevance" above.
- The repetition fixes catch exact repeats reliably but can miss near-duplicate phrasing with minor word substitutions.
License
Released under CC-BY-SA-4.0, matching the license of the primary training dataset.
- Downloads last month
- 255