TinyJLLM โ€” 100M-parameter small language model built from scratch

A decoder-only Transformer (~102.5M parameters) pretrained from random initialization on ~5 GB of FineWeb (sample-10BT), 3 epochs / 108,000 optimizer steps. Every component โ€” tokenizer, data pipeline, model, training loop, evaluation, export โ€” is implemented from scratch in the TinyLLM repository.

Final metrics: validation loss 3.50 (perplexity 33.1); the best checkpoint reached 3.48 / 32.6.

Model family

Model Stage Link
TinyJLLM Base (this model) jaweed123/TinyJLLM
TinyJLLM-Instruct Supervised fine-tuning jaweed123/TinyJLLM-Instruct
TinyJLLM-Instruct-DPO DPO preference tuning jaweed123/TinyJLLM-Instruct-DPO

Model details

Property Value
Parameters 102,450,432 (~102.5M)
Architecture Llama-style decoder-only: RMSNorm, RoPE (half-split), SwiGLU, tied embeddings, no biases
Layers / heads / head_dim 11 / 12 / 64
Context length 512
Vocabulary 32,000 (custom byte-level BPE, <pad> <unk> <bos> <eos> = 0โ€“3)
Pretraining data FineWeb sample-10BT, ~5.37 GB raw, 1.75M documents
Tokens seen 3.54B (3 epochs)
Hardware RTX 4060 8 GB, ~30K tok/s (torch.compile)
Precision BF16 mixed precision, FP32 master weights

Intended use

  • Educational reference: inspect a small, complete, honest pretraining run.
  • Qualitative experimentation: prompt it (a base model โ€” it continues text, it does not follow instructions).
  • A starting point for post-training (see the Instruct and DPO variants).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("jaweed123/TinyJLLM")
tokenizer = AutoTokenizer.from_pretrained("jaweed123/TinyJLLM")

prompt = "The future of AI is"
inputs = tokenizer(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=50, temperature=0.8, top_k=50, top_p=0.95)
print(tokenizer.decode(out[0]))

llama.cpp / GGUF

This repo also ships GGUF files (F16, Q8_0, Q4_K_M) โ€” load directly with llama.cpp or llama-cpp-python.

Training details

  • Custom 32K byte-level BPE tokenizer (trained on a 512 MB FineWeb sample).
  • Tokens stored once as uint16 shards (591 train + 6 validation).
  • AdamW (lr 3e-4, wd 0.1, decay/no-decay groups), 1,000-step warmup + cosine decay to 1e-5, effective batch 32,768 tokens/step, gradient clipping 1.0.
  • Full run: ~35 h on a single RTX 4060.

Limitations

Small scale โ‡’ repetition in long generations, weak instruction following, limited and unreliable world knowledge. Not for factual or consequential use.

Citation

@misc{tinyjllm,
  title  = {TinyJLLM: A 100M-Parameter Small Language Model Built From Scratch},
  author = {Jaweed, Abdul},
  year   = {2026},
  url    = {https://github.com/Abdul-Jaweed/TinyLLM}
}

Acknowledgments

FineWeb (HuggingFaceFW), Hugging Face tokenizers / datasets, PyTorch, and llama.cpp. Built and measured with the TinyLLM pipeline.

Downloads last month
806
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jaweed123/TinyJLLM

Finetunes
1 model

Dataset used to train jaweed123/TinyJLLM