workshop-pretraining-v2

A small (123.6M parameters) decoder-only Transformer language model, pretrained from scratch for only 100 optimizer steps as part of a hands-on LLM pretraining workshop.

This is an educational checkpoint, not a usable language model. It was trained on about 52 million tokens (roughly 0.5% of the 10B-token plan in the workshop notebook), so it produces mostly frequent words with no coherent meaning. See Limitations.

Model details

Item Value
Architecture Decoder-only Transformer (GPT-2 sized, with modern components)
Parameters 123,587,328
Layers / hidden size / heads 12 / 768 / 12 (head dim 64)
Context length 1024 tokens
Vocabulary 50,304 (GPT-2 BPE with 50,257 tokens, padded to a multiple of 64)
Feed-forward 4x expansion, ReLUΒ² activation, no biases
Normalization RMSNorm, pre-norm
Positional encoding Rotary embeddings (RoPE, theta = 10000)
Output head Tied with the input embedding, tanh logit soft-capping (cap = 30)
Dropout 0.1 (training only)
Tokenizer tiktoken GPT-2 encoding (`<

This is a custom architecture, not GPT2LMHeadModel. Loading it requires trust_remote_code=True. The model code is in configuration_gpt2workshop.py and modeling_gpt2workshop.py in this repository.

Training

Item Value
Data HuggingFaceFW/fineweb-edu, sample-10BT subset, streamed and tokenized into uint16 shards
Training length 100 optimizer steps
Tokens per step 524,288 (micro-batch 4 x 1024 tokens x 128 gradient-accumulation steps)
Tokens seen about 52.4M
Optimizer AdamW, betas (0.9, 0.95), weight decay 0.1 (not applied to norms and 1-D parameters), fused
Learning rate Peak 6e-4, linear warmup for 10 steps, then cosine decay to 6e-5
Gradient clipping 1.0
Precision bfloat16 autocast
Hardware 1x NVIDIA GeForce RTX 3080 Laptop GPU (8 GB), about 93 minutes, about 9,000 tokens/s
Framework PyTorch (no torch.compile)

Results

Metric Value
Training loss, step 10 8.58
Training loss, step 50 7.28
Training loss, step 100 6.97
LAMBADA accuracy (steps 50 and 100) 0.0000

For reference, a uniform random guess over this vocabulary gives a loss of about 10.8.

How to use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "kihyounghan/workshop-pretraining-v2"
device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).to(device)
model.eval()

prompt = "The best way to learn programming is"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to(device)

with torch.no_grad():
    output_ids = model.generate(
        input_ids, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=50
    )

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

trust_remote_code=True runs the Python files stored in this repository. Review them first, or pin a specific commit with revision="<commit hash>".

Example output

Prompt: The theory of relativity states that

The theory of relativity states that you, they the research and the high of the following of ...

The model has learned which words are frequent in English, but not grammar or meaning yet.

Limitations

  • Trained for 100 steps only, so outputs are repetitive, ungrammatical and unrelated to the prompt.
  • Not evaluated for safety, factuality or bias. It is not suitable for any real application.
  • Trained on English web text only. The training data may contain biases and errors.
  • Intended for learning and for testing the pretraining, export and upload pipeline.

ν•œκ΅­μ–΄ μš”μ•½

LLM μ‚¬μ „ν•™μŠ΅(pretraining) μ›Œν¬μˆ μ‹€μŠ΅μš©μœΌλ‘œ μ²˜μŒλΆ€ν„° ν•™μŠ΅ν•œ μ•½ 1.24μ–΅ νŒŒλΌλ―Έν„° λͺ¨λΈμž…λ‹ˆλ‹€. 100 μŠ€ν…(μ•½ 5,200만 토큰)만 ν•™μŠ΅ν•΄μ„œ μ‹€μ œλ‘œ μ“Έ 수 μžˆλŠ” μˆ˜μ€€μ΄ μ•„λ‹ˆλ©°, ν•™μŠ΅ β†’ μ €μž₯ β†’ Hugging Face μ—…λ‘œλ“œ 과정을 ν™•μΈν•˜κΈ° μœ„ν•œ 예제 μ²΄ν¬ν¬μΈνŠΈμž…λ‹ˆλ‹€.

Downloads last month
1,248
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train kihyounghan/workshop-pretraining-v2