Model archived by:

DedeProGames
DedeProGames

NanoDex-1M

A 1,062,272-parameter decoder-only language model pre-trained from scratch on fineweb-edu, using the NanoDex Trainer Space.

Architecture

A standard LlamaForCausalLM decoder-only transformer โ€” SiLU MLP, RMSNorm, rotary position embeddings, grouped-query attention, tied embeddings, no biases โ€” scaled down in width and depth to fit the parameter budget.

Parameters 1,062,272
Hidden size 128
Layers 5
Attention heads 8 (KV: 4)
FFN size 288
Context length 512
Vocab 2,048 (custom BPE trained on fineweb-edu)

Training

Tokens seen 999,817,216
Steps 3,814
Tokens / step 262,144
Optimizer AdamW(0.9, 0.95) wd=0.1 clip=1.0
LR schedule warmup 2% + cosine to 10% (peak 3e-03)
Final loss 3.3092 (ppl 27.4)
Wall time 43.8 min
Trained by @DedeProGames

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("DedeProGames/NanoDex-1M")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/NanoDex-1M")

ids = tok("The mitochondria is", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=60, do_sample=True,
                                temperature=0.8, top_k=50)[0]))

Caveats

This is a nano-scale research artifact. At this parameter count and token budget the model learns word shapes, common collocations and a little syntax โ€” it is not a useful assistant and its output is not factual. It exists to make "pre-train a transformer from scratch" something you can actually watch happen.

Downloads last month
723
Safetensors
Model size
1.06M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Dataset used to train SLM-Archive/NanoDex-1M

Space using SLM-Archive/NanoDex-1M 1