Instructions to use yava-code/Tessera-135M-Gate with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yava-code/Tessera-135M-Gate with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yava-code/Tessera-135M-Gate", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-135M-Gate", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yava-code/Tessera-135M-Gate with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yava-code/Tessera-135M-Gate" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-135M-Gate", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/yava-code/Tessera-135M-Gate
- SGLang
How to use yava-code/Tessera-135M-Gate with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-135M-Gate" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-135M-Gate", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-135M-Gate" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-135M-Gate", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use yava-code/Tessera-135M-Gate with Docker Model Runner:
docker model run hf.co/yava-code/Tessera-135M-Gate
Tessera-135M-Gate
The 135M architecture-gate checkpoint from Paragon Intelligence Labs: the smallest
member of the Tessera family (HuggingFaceTB/SmolLM2-135M backbone with the full
Next Concept Prediction path). It is a deliberately overfit gate run that proves
the complete concept path trains end to end before any large spend, and carries the
same causal concept-path architecture as Tessera-1B-Nano. It is not a language
model quality result: it saw a 1M-token TinyStories subset repeated 16 times.
TL;DR
This gate checkpoint exists to de-risk the 1B-token comparison, and the scale story it records is the point of publishing it.
- All three objectives train (NTP, NCP, VQ fell 41 to 43% over the run).
- The decoder causally uses the concept channel even at 135M: zeroing the feedback costs +0.0205 nats of held-out NTP loss.
- The failure mode this gate exposed is real and instructive: effective codebook perplexity stayed near 2.4 (usage ~31-41%), and shuffled feedback cost nothing (+0.000). At 1B tokens the same architecture left the shortcut far behind (perplexity 7.55, usage 85.6%, zero-delta +0.105). Gate small, then scale: some behaviors only appear above a scale threshold.
Evaluation
| Metric | Value |
|---|---|
| Held-out NTP loss | 1.8751 |
| Held-out perplexity | 6.5213 |
| Zero feedback delta | 0.0205 |
| Shuffled feedback delta | 0.0000 |
| Codebook perplexity / usage | 2.4087 / 33.0% |
| Training tokens | 16,777,216 |
| Tracked compute estimate | $1.38 |
Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.
Architecture
- chunk size: 4
- product code: 9 segments x 64 entries
- causal concept blocks: 2
- injection point: before token decoder block 2
- NCP target: next continuous concept
- loss:
L_ntp + 1 L_ncp + 1 L_vq
This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.
Training data and provenance
The matched corpus and the training record are pinned in the ncp-smol repository:
packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw
intervention evaluation JSON (also shipped in this repository as eval.json and
metrics.jsonl).
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("REPO_ID")
model = AutoModelForCausalLM.from_pretrained("REPO_ID", trust_remote_code=True)
The Tessera family
Three checkpoints, one story, in reading order:
- Tessera-135M-Gate - 135M architecture gate (overfit, published for the scale story) (this model)
- Tessera-1B-Nano-Base - 1B-token matched NTP-only control
- Tessera-1B-Nano - 1B-token matched concept arm
All cards are generated from the run artifacts by the same build_card; the
study repository holds the
whitepaper and full records.
References
- Code and study: https://github.com/yava-code/Tessera-1B-Nano
- ConceptLM: https://arxiv.org/abs/2602.08984
- NCP-ArchPreview: https://arxiv.org/abs/2609.10715
- Downloads last month
- 277
Model tree for yava-code/Tessera-135M-Gate
Base model
HuggingFaceTB/SmolLM2-135M