SLM-125M base — legal & financial

A 125.8M-parameter Llama-style language model pretrained from scratch (fresh tokenizer, fresh weights) on US court opinions, SEC filings and educational web text.

This is a base model. It continues text; it does not follow instructions or chat. Give it the opening of a passage and it writes the rest in the same register.

Held-out result: loss 2.1269 nats/token · perplexity 8.389 · 0.6572 bits/byte, on a document-level held-out split of 19,558,737 tokens that was never trained on.

Training curve

Use

Prefix prompts with <|eos|>, not <|bos|>. In training every document begins right after an <|eos|>, and <|bos|> never appears in the training data.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("Deepanshu027/slm-125m-base")
model = AutoModelForCausalLM.from_pretrained("Deepanshu027/slm-125m-base")
ids = [tok.convert_tokens_to_ids("<|eos|>")] + tok.encode("The court held that the defendant", add_special_tokens=False)
out = model.generate(torch.tensor([ids]), max_new_tokens=60, min_new_tokens=16, do_sample=True,
                     temperature=0.8, top_k=50, top_p=0.95, repetition_penalty=1.1)
print(tok.decode(out[0, len(ids):], skip_special_tokens=True))

Try it in the browser at slm-125m-legal-sandy.vercel.app, or call the hosted API at https://mdeepanshu2706--slm-125m-serve-api.modal.run (POST /generate; OpenAPI docs at /docs).

Architecture

Parameters 125,848,320 (all trainable; 113,265,408 non-embedding)
Layers / hidden / heads 12 / 768 / 12 (head dim 64, MHA)
MLP SwiGLU, intermediate 3072
Normalisation / positions RMSNorm (pre-norm) / RoPE (theta 10000)
Context 1024 tokens
Embeddings tied input/output
Tokenizer byte-level BPE, trained from scratch on this corpus, 16,384 tokens

Training data

legal-first: all of case-law and SEC (they hold only ~2B clean tokens), plus a 0.5B web slice. The mix is not the often-quoted 70/20/10: the two legal sources simply do not contain more text, so every clean legal token is used and web text is capped.

Source Dataset Train tokens Share Documents
case-law HFforLegal/case-law 715M 35.1% 206,684
sec PleIAs/SEC 861M 42.2% 45,035
fineweb-edu HuggingFaceFW/fineweb-edu 464M 22.8% 418,405

Pipeline: streamed (never downloaded whole) → a deterministic 6-step cleaning chain (line filters, boilerplate, length, repetition, language, OCR-garble gate for scanned opinions) → MinHash-LSH near-dedup (Jaccard 0.8) and exact dedup → decontamination against CaseHOLD / LexGLUE (any document sharing a 13-gram with the benchmark is dropped) → tokenized and packed into 1,024-token windows. Validation is split by document, so no held-out document is ever partly in training.

Training

Tokens seen 8.16B (4 epochs over 2.04B unique tokens; 64.8 tokens/param)
Steps × batch 15,564 × 524,288 tokens
Optimiser AdamW (b1 0.9, b2 0.95, wd 0.1), grad-clip 1.0
LR 0.0006 → 6e-05, 381 warmup steps, linear warmup, one cosine decay over all steps
Precision fp32 master weights, bf16 autocast
Hardware 8 x H100 SXM5 (Modal), 44.9 min, 3.35M tokens/s (32.0% MFU)
Compute cost ≤ $24.78

Evaluation (held-out, whole split)

Step Tokens Loss Perplexity Bits/byte
1,000 0.52B 2.7763 16.06 0.8578
4,000 2.10B 2.3637 10.63 0.7303
7,000 3.67B 2.2546 9.53 0.6966
10,000 5.24B 2.1883 8.92 0.6761
13,000 6.82B 2.1458 8.55 0.6630
15,564 8.16B 2.1269 8.39 0.6572
15,564 8.16B 2.1269 8.39 0.6572

Held-out loss improved at every one of the 16 evaluations, so 4 epochs did not overfit. Gains roughly halved with each epoch, so this corpus is close to exhausted at this size. Bits-per-byte is tokenizer-independent; use it rather than perplexity to compare against models with other vocabularies. Per source, the model is far stronger on legal text (≈0.69 bits/byte on case law) than on general web text (≈1.0).

Limitations

  • Base model, 125M parameters. Fluent in legal and financial register; unreliable on facts, especially general knowledge (e.g. it will confidently misdescribe photosynthesis).
  • May reproduce text from its training data: public court opinions and SEC filings, which contain names of real people and companies. Do not treat output as legal or financial advice.
  • Domain and era skew: SEC text is weighted toward 1990s 10-K filings; case-law is US-only and partly OCR'd.
  • No safety tuning, instruction tuning or RLHF.
Downloads last month
245
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Deepanshu027/slm-125m-base