Tessera-135M-Gate

The 135M architecture-gate checkpoint from Paragon Intelligence Labs: the smallest member of the Tessera family (HuggingFaceTB/SmolLM2-135M backbone with the full Next Concept Prediction path). It is a deliberately overfit gate run that proves the complete concept path trains end to end before any large spend, and carries the same causal concept-path architecture as Tessera-1B-Nano. It is not a language model quality result: it saw a 1M-token TinyStories subset repeated 16 times.

TL;DR

This gate checkpoint exists to de-risk the 1B-token comparison, and the scale story it records is the point of publishing it.

  • All three objectives train (NTP, NCP, VQ fell 41 to 43% over the run).
  • The decoder causally uses the concept channel even at 135M: zeroing the feedback costs +0.0205 nats of held-out NTP loss.
  • The failure mode this gate exposed is real and instructive: effective codebook perplexity stayed near 2.4 (usage ~31-41%), and shuffled feedback cost nothing (+0.000). At 1B tokens the same architecture left the shortcut far behind (perplexity 7.55, usage 85.6%, zero-delta +0.105). Gate small, then scale: some behaviors only appear above a scale threshold.

Evaluation

Metric Value
Held-out NTP loss 1.8751
Held-out perplexity 6.5213
Zero feedback delta 0.0205
Shuffled feedback delta 0.0000
Codebook perplexity / usage 2.4087 / 33.0%
Training tokens 16,777,216
Tracked compute estimate $1.38

Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.

Architecture

  • chunk size: 4
  • product code: 9 segments x 64 entries
  • causal concept blocks: 2
  • injection point: before token decoder block 2
  • NCP target: next continuous concept
  • loss: L_ntp + 1 L_ncp + 1 L_vq

This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.

Training data and provenance

The matched corpus and the training record are pinned in the ncp-smol repository: packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw intervention evaluation JSON (also shipped in this repository as eval.json and metrics.jsonl).

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("REPO_ID")
model = AutoModelForCausalLM.from_pretrained("REPO_ID", trust_remote_code=True)

The Tessera family

Three checkpoints, one story, in reading order:

All cards are generated from the run artifacts by the same build_card; the study repository holds the whitepaper and full records.

References

Downloads last month
277
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yava-code/Tessera-135M-Gate

Finetuned
(945)
this model

Collection including yava-code/Tessera-135M-Gate

Papers for yava-code/Tessera-135M-Gate