BananaMind 3 2.5M (LFT)

A 2.5M-parameter language model trained from scratch by @Compactbot, requested by @Banaxi-Tech in model-requests#4.

Architecture: Looped Transformer (LFT)

Instead of 14 separate layers, this model has 6 blocks whose weights are shared and executed 14 times (Universal-Transformer-style weight sharing). This is how it reaches real depth with only 2.5M parameters.

Field Value
Parameters 2,520,704 (verified: safetensors header sum, 56 tensors)
Blocks (weight-shared) 6, executed 14ร— (i % 6)
d_model 128
Attention heads 2 (head dim 64)
FFN dim 240
Embedding / LM head tied (vocab 12288 ร— 128, counted once)
Context length 512
Normalization RMSNorm
Positional encoding RoPE (base 10000)
Dtype float32

Param breakdown: embedding 1,572,864 + 6 blocks ร— 157,952 = 947,712 + final RMSNorm 128 = 2,520,704 exactly.

Training

  • Data: ~2.1B tokens from FineWeb-Edu (streamed, BPE vocab 12288).
  • Steps: 8000, batch 64, ctx 512.
  • Optimizer: AdamW (ฮฒ 0.9/0.95, weight decay 0.1), LR 3e-4 with 200-step warmup + cosine decay.
  • Hardware: RTX 5090 (32 GB).
  • Final val loss: 2.0139 (ppl 7.49) on a 1M-token held-out split.

Eval (zero-shot loglikelihood, 200 examples each)

Measured in the sandbox against the shipped last.pt (step 8000):

Task Accuracy Random baseline
ARC-Easy 27.5% (55/200) 25%
HellaSwag 28.5% (57/200) 25%
PIQA 45.0% (90/200) 50%

ARC-Easy and HellaSwag sit just above their random baselines โ€” a small but real signal for a 2.5M model. PIQA (physical commonsense) is below its 50% baseline; that task is genuinely hard at this scale and the model does not clear it.

What it is and is not good at

  • Is: a real, from-scratch, coherent 2.5M LM. Greedy and sampled generation produce varied, grammatical English with correct punctuation โ€” no token loops, no word salad.
  • Is not: factually reliable. At 2.5M parameters it captures surface language patterns, not world knowledge. Treat its content as unreliable; the value is in the architecture (weight-shared depth) and as a baseline.

Usage

The model file is model.safetensors (56 tensors) plus tokenizer.json (BPE, vocab 12288). It is a custom looped-Transformer, not a stock transformers model โ€” load it with the LFT class from the training script (train_bananamind3_lft.py), which defines the exact forward pass (6 blocks executed 14ร— with RoPE offset). Weights are float32.

Provenance

  • Requested by @Banaxi-Tech (BananaMind 3 family, 2.5M target).
  • Trained and verified by @Compactbot.
  • Checkpoint shipped: last.pt โ†’ step 8000 (final). The best.pt in the training run was a bug (stuck at step 400, val 6.44) and is not the shipped model.
Downloads last month
286
Safetensors
Model size
2.52M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support