Forge-20M
A 20.75M parameter, GPT-2-style decoder-only transformer, architected and trained completely from scratch — no pretrained weights, no fine-tuning of an existing model. Trained end-to-end on a single NVIDIA T4 GPU (Kaggle, free tier).
Model Details
Model Description
Forge-20M is a small causal language model built to implement and train a GPT-style transformer from first principles: custom causal self-attention, learned positional embeddings, weight tying, and a full training loop with mixed precision and gradient accumulation. It is trained on educational web text and produces fluent, grammatically coherent completions, but is not factually reliable and should not be used for any application requiring accurate information.
- Developed by: [Your Name]
- Model type: Decoder-only transformer (GPT-2 architecture family)
- Language(s): English
- License: MIT
- Base architecture: nanoGPT by Andrej Karpathy (reimplemented from scratch, not a fork)
Model Sources
- Repository: [link to your GitHub repo]
- Demo: [link to your Gradio Space / Kaggle notebook, if public]
Uses
Direct Use
Text completion for research, educational, and experimental purposes — exploring small-scale language model behavior, studying training dynamics on constrained hardware, or as a teaching example for from-scratch transformer implementation.
Out-of-Scope Use
- Not suitable for factual question-answering. At 20.75M parameters, the model does not have sufficient capacity to reliably store factual knowledge. Outputs that resemble facts are frequently fabricated.
- Not instruction-tuned. This is a base completion model, not a chat or assistant model. It continues text in the style of its training data rather than following instructions or answering questions directly.
- Not suitable for any production, medical, legal, financial, or safety-critical application.
- Not suitable for multi-step reasoning tasks.
Bias, Risks, and Limitations
- Trained exclusively on FineWeb-Edu (English educational web content), so outputs reflect that domain's style and any biases present in that corpus.
- Coherence degrades noticeably past 4-6 sentences of generation.
- At this parameter count, factual content in generated text should be assumed fabricated unless independently verified.
- No safety filtering, no RLHF, no instruction tuning has been applied. Outputs are unfiltered raw completions.
Recommendations
Users should independently verify any factual claims in generated output and should not deploy this model in any context where incorrect information could cause harm. This model is best understood as a research/learning artifact demonstrating from-scratch LLM training on constrained hardware, not as a production-ready text generation tool.
How to Get Started with the Model
import torch
import tiktoken
device = 'cuda' if torch.cuda.is_available() else 'cpu'
# Requires the GPTConfig / GPT class definitions from this repository
checkpoint = torch.load('ckpt_final_iter4931_valloss4.211.pt', map_location=device)
gptconf = GPTConfig(**checkpoint['model_args'])
model = GPT(gptconf)
model.load_state_dict(checkpoint['model'])
model.to(device)
model.eval()
enc = tiktoken.get_encoding("gpt2")
def generate(prompt, max_new_tokens=150, temperature=0.7, top_k=25, repetition_penalty=1.3):
start_ids = enc.encode(prompt)
x = torch.tensor(start_ids, dtype=torch.long, device=device)[None, ...]
idx = x
generated_ids = []
with torch.no_grad():
for _ in range(max_new_tokens):
idx_cond = idx if idx.size(1) <= model.config.block_size else idx[:, -model.config.block_size:]
logits, _ = model(idx_cond)
logits = logits[:, -1, :]
for token_id in set(idx[0].tolist()):
if logits[0, token_id] > 0:
logits[0, token_id] /= repetition_penalty
else:
logits[0, token_id] *= repetition_penalty
logits = logits / temperature
v, _ = torch.topk(logits, min(top_k, logits.size(-1)))
logits[logits < v[:, [-1]]] = -float('Inf')
probs = torch.nn.functional.softmax(logits, dim=-1)
idx_next = torch.multinomial(probs, num_samples=1)
if idx_next.item() == 50256:
break
idx = torch.cat((idx, idx_next), dim=1)
generated_ids.append(idx_next.item())
return enc.decode(generated_ids)
print(generate("The main concept of physics is"))
Training Details
Training Data
- Dataset: HuggingFaceFW/fineweb-edu,
sample-10BTsplit - Total tokens trained on: ~628M cumulative (100M in initial training, 300M additional fresh, non-overlapping tokens in extended training)
- Tokenizer: GPT-2 byte-level BPE (via
tiktoken), vocabulary size 50,304 (padded)
Training Procedure
Two-stage training on a single NVIDIA T4 GPU (16GB VRAM):
Stage 1:
- 100M tokens, 5,000 steps, batch size 32, gradient accumulation 4 (effective batch 128), block size 512
- AdamW (fused), peak LR 5e-4, cosine decay to 5e-5, 100-step warmup
- fp16 mixed precision, ~94 minutes training time
Stage 2 (resumed):
- 300M additional fresh tokens, cumulative to 4,931 steps
- AdamW (fused), peak LR 3e-4, cosine decay to 2e-5
- fp16 mixed precision, ~92 minutes training time
Training Hyperparameters
| Hyperparameter | Value |
|---|---|
| Optimizer | AdamW (fused), betas=(0.9, 0.95), weight_decay=0.1 |
| Batch size | 32 |
| Gradient accumulation steps | 4 |
| Effective batch size | 128 |
| Block size (context length) | 512 |
| Precision | fp16 (autocast + gradient scaler) |
| Gradient clipping | max_norm 1.0 |
Speeds, Sizes, Times
- Total parameters: 20,878,592 (20.75M)
- Total training time:
186 minutes (3.1 hours) across both stages - Hardware: 1x NVIDIA Tesla T4, 16GB VRAM (Kaggle free tier)
- Throughput: ~58,000 tokens/second
Evaluation
Results
| Checkpoint | Train Loss | Val Loss |
|---|---|---|
| Stage 1 final (step 5,000) | 4.18 | — |
| Stage 2, step 1,000 | 5.13 | 5.15 |
| Stage 2, step 2,000 | 4.55 | 4.54 |
| Stage 2, step 3,000 | 4.35 | 4.35 |
| Stage 2, step 4,000 | 4.23 | 4.24 |
| Stage 2 final (step 4,931) | 4.22 | 4.21 |
Train/validation loss gap remained under 0.05 throughout both training stages — no overfitting observed.
Qualitative Observations
- After Stage 1 (100M tokens): correct grammar and sentence structure, but prone to exact-phrase repetition loops and topic drift after 2-3 sentences.
- After Stage 2 (300M additional tokens): repetition substantially reduced (particularly combined with inference-time repetition penalty), coherence extended to 4-5 sentences, richer vocabulary, learned real document formats (Q&A, academic prose, FAQ structure). Factual accuracy and multi-step reasoning remain unreliable at this parameter count.
Environmental Impact
- Hardware Type: NVIDIA Tesla T4
- Hours used: ~3.1 hours
- Cloud Provider: Kaggle (Google Cloud-backed infrastructure)
- Compute Region: Unknown (Kaggle-managed)
Given the small scale (single GPU, ~3 hours total), the carbon footprint of training this model is minimal relative to large-scale LLM training runs.
Technical Specifications
Model Architecture and Objective
Standard GPT-2-style decoder-only transformer trained with a causal (next-token prediction) language modeling objective.
| Component | Detail |
|---|---|
| Layers | 10 |
| Attention heads | 8 |
| Embedding dimension | 256 |
| Context length | 512 |
| Vocabulary size | 50,304 |
| Positional encoding | Learned absolute |
| Normalization | Pre-norm LayerNorm |
| Activation | GELU |
| Weight tying | Yes (token embedding ↔ output head) |
| Bias | Disabled |
| Attention mechanism | Causal multi-head self-attention (flash attention via scaled_dot_product_attention where available) |
Compute Infrastructure
- Hardware: 1x NVIDIA Tesla T4, 16GB VRAM
- Software: PyTorch 2.x, CUDA 12.x, Python 3.12,
tiktokenfor tokenization
Citation
If referencing this project:
@misc{forge20m,
title = {Forge-20M: A 20.75M Parameter GPT Trained From Scratch on FineWeb-Edu},
author = {[Your Name]},
year = {2026},
note = {Trained from scratch on a single NVIDIA T4 GPU}
}