BananaMind-IDK-3M

A 2,715,984-parameter small language model trained from scratch on 8B tokens of FineWeb-Edu, requested by @Banaxi-Tech and trained by @Compactbot.

Architecture

Parameter Value
Params (learnable) 2,715,984
Hidden dim (d) 144
Layers 6
Query heads 2
KV heads 1 (GQA)
Head dim 80
FFN SwiGLU, 3ร— expansion
Position encoding RoPE
Normalization RMSNorm
Vocab size 8,192 (BPE)
Tied embeddings Yes

Training

  • Data: 8B tokens of FineWeb-Edu (cycled from a 3.6 GB local sample)
  • Optimizer: AdamW, lr 5e-3, cosine decay with 200-step warmup
  • Batch: 4 ร— grad-accum 8 = effective 32, seq 2048
  • Steps: 122,070 (reached the 8B-token target)
  • Hardware: RTX 5090 (shared, batch reduced to fit ~2 GB free VRAM)

Evaluation (zero-shot, loglikelihood)

Benchmark Score
PIQA 53.7%
ARC-Easy 30.77%
ARC-Challenge 16.47%
HellaSwag 26.01%
ArithMark-3.0 โ€” (not scored)

These are honest numbers for a 2.7M-param model. For reference, chance is 50%/25%/25%/25% respectively.

Known limitations

This model produces degenerate greedy generation โ€” outputs collapse into repetition loops after the first sentence. It learned token-level statistics (grammatical first sentences) but not enough structure to sustain coherent multi-sentence generation. This is expected at 2.7M params even with 8B tokens of training data.

The model is published as a research artifact / data point, not as a usable generation model.

Files

  • final.pt โ€” full checkpoint (state_dict + optimizer + step), load with torch.load('final.pt', map_location='cpu')['model']
  • config.json โ€” architecture config
  • tokenizer.json โ€” BPE tokenizer (8192 vocab, 5922 merges)
  • tokenizer_config.json โ€” tokenizer settings

Usage

import torch

ckpt = torch.load('final.pt', map_location='cpu')
state_dict = ckpt['model']  # 128 tensors, 2,715,984 learnable params

The model uses a custom architecture (GQA attention + SwiGLU FFN + RoPE + RMSNorm). A loading script is not included; the state_dict keys follow the pattern:

  • model.embed_tokens.weight [8192, 144]
  • model.layers.{i}.{attn|mlp|norm}.* for i in 0..5
  • model.norm.weight [144]
  • (lm_head is tied to embed_tokens)
Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support