VISTA-24M / README.md
AwakeningOS's picture
Release VISTA-24M: model, architecture diagrams, training recipe and evaluation evidence
9287d39 verified
|
Raw History Blame Contribute Delete
15.7 kB
---
language:
- en
license: mit
library_name: transformers
pipeline_tag: text-generation
tags:
- pytorch
- custom_code
- causal-lm
- small-language-model
- babylm
- strict-small
- swiglu
- research
---
# VISTA-24M
### A Variance-Informed 7-Layer Language Model
**24.28M parameters · 7 layers · 16,384-token vocabulary · English · MIT**
An independent language-model experiment by **Yusuke Maeda**.
VISTA stands for **Variance-Informed SwiGLU Transformer Architecture**. It gives the feed-forward network three things to work with: the current representation, the changes made in the previous layer, and the spread of the values combined by Attention.
The idea is simple: an average tells us what was collected; the spread adds information about what went into that average. VISTA makes both available to the computation that follows.
| At a glance | This release |
|---|---|
| Default model (`main`) | Raw **80M-word** checkpoint |
| Training corpus | BabyLM 2026 English Strict-Small, **10M words** |
| Complete training run | 10 passes, **100M words of exposure** |
| Released checkpoints | **19** revisions: 1–9M, then 10–100M in 10M steps |
| Full zero-shot six-task mean at 80M | **48.07** |
| BabyLM seven-category NLP mean at 80M | **50.12** |
| Intended use | Small-model research, probability scoring, representation analysis |
**80M means training words, not model parameters or subword tokens.** This checkpoint has seen eight passes through the 10M-word corpus. Its training counter records 129,927,852 valid next-token targets.
## How it works
![VISTA architecture: Attention content and variance, plus previous-layer changes, enter SwiGLU](assets/architecture.png)
Each layer first reads the earlier tokens with causal Attention. The resulting update is added to the current representation and RMS-normalized. The FFN then transforms this updated representation, with two additional sources of information:
1. **What changed in the previous layer?** Two full-width vectors describe its Attention update and its FFN update. A small router uses the current representation to weight their independently projected contributions to the next FFN.
2. **How spread out were the values read by this layer's Attention?** VISTA computes a coordinate-wise variance using the same Attention weights as the normal weighted average. A dedicated projection delivers its direction and magnitude to the FFN.
The FFN combines these inputs inside the gate and up branches of **SwiGLU**, then writes its result back to the main representation. Both residual updates preserve approximately unit RMS. The output embedding and input embedding have separate weights.
| Component | Implementation |
|---|---|
| Main representation | 256 dimensions; RMS-normalized after each residual update |
| Attention | 8 query / 8 key-value heads; 32 dimensions per head; RoPE |
| Attention output gate | Learned, elementwise sigmoid gate before output projection |
| FFN | SwiGLU; intermediate width 896 |
| Previous-layer information | Two normalized **256-dimensional** differences; layers 2–7 |
| Difference router | 16-dimensional query/key projections; softmax over two inputs |
| Variance input | 256-dimensional direction + 1 magnitude value |
| Variance projection | Separate 257 → 1,792 projection in every layer |
| Parameters | **24,281,088**; input/output vocabulary matrices are untied |
| Training context | Up to 512 token positions |
The router's 16-dimensional projection is a scoring space. The information delivered to the FFN remains full-width. At each layer, the two difference vectors come from the **immediately preceding layer**. The current representation itself already carries the processing done by earlier layers.
For a mathematical description, see [Architecture](ARCHITECTURE.md). The model core is in [dense.py](dense.py); the historical class names are retained for checkpoint compatibility.
## Load and use
The repository contains custom PyTorch code, loaded through Transformers. Review that code before enabling `trust_remote_code`. The default backend uses PyTorch SDPA and works without installing FlashAttention.
```bash
pip install -r requirements.txt
```
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
repo = "AwakeningOS/VISTA-24M"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True
).eval()
# Keep the saved parameters in FP32. On CUDA, the adapter uses BF16 autocast.
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
inputs = tokenizer("The little dog ran", return_tensors="pt").to(device)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=32,
do_sample=False, use_cache=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```
This is a **base next-token model**. Use text continuations rather than a chat template. The implementation recomputes the prefix during generation; KV caching is currently unsupported. Keep prompt plus continuation within 512 tokens for the trained context range. The release supplies a ByteLevel decoder for readable generated text; input token IDs and all encoding components match the training tokenizer, preserved under `training/original_tokenizer/`.
For likelihood scoring, pass `labels=inputs["input_ids"]` together with the `attention_mask`; the adapter shifts labels internally and excludes padding transitions. `AutoModel` is also registered for downstream classifiers. Its `output_hidden_states` currently exposes only the final representation.
Load another training point with `revision="chck_40M"` in **both** `from_pretrained` calls. For reproducible experiments, pin a commit hash as the revision.
## Performance: what does this model do well?
The full 80M-word evaluation is strongest on grammatical distinctions and weakest on tracking entities through a changing description. World knowledge and conceptual-property judgments remain challenging at this scale.
| Evaluation | What it probes | VISTA-24M, 80M words |
|---|---|---:|
| BLiMP | Grammatical sentence preferences | **67.73** |
| BLiMP Supplement | Additional grammatical and linguistic contrasts | **58.41** |
| EWoK | Physical and social world knowledge | **52.83** |
| Entity Tracking | Keeping track of entities as a description develops | **18.56** |
| COMPS | Concept–property knowledge and property inheritance | **51.32** |
| Global PIQA, parallel | Shared cultural-commonsense questions, English subset | **28.16** |
| Global PIQA, nonparallel | Culture-specific questions, English subset | **51.00** |
| Global PIQA, two-split mean | One category in the NLP aggregate | **39.58** |
| (Super)GLUE | Seven downstream classification tasks, after fine-tuning | **62.43** |
Scores above are percentages; higher is better. The **seven-category NLP mean is 50.12**: BLiMP, Supplement, EWoK, Entity Tracking, COMPS, Global PIQA (counted once), and (Super)GLUE. The six-task zero-shot mean excludes (Super)GLUE and is **48.07**.
Evaluation uses the official [BabyLM 2026 evaluator](https://github.com/babylm-org/babylm-eval/tree/6f825c291e2c4c78ad33b1935fd64d45f52642dc), causal scoring and the full filtered datasets. Sentence-scoring temperature was **1.0**. The zero-shot summary uses the evaluator's task-average result. Global PIQA uses length-normalized conditional scores.
### Learning through the run
![Full six-task zero-shot mean across ten training checkpoints, with 80M highlighted](assets/learning_curve.png)
The **80M-word checkpoint had the highest full six-task mean among the ten checkpoints evaluated**. It is therefore the default download. The 100M-word checkpoint remains available for studying the complete trajectory. Selection used this evaluation curve; these results do not constitute an untouched final test or a multi-seed estimate.
Every point's category scores are available in [learning_curve.csv](evaluation/learning_curve.csv). The chart uses the six-task mean throughout: downstream fine-tuning was run for the selected 80M model, so a seven-category curve is not available.
### Downstream fine-tuning
| Task | Metric | Score (%) |
|---|---|---:|
| BoolQ | Accuracy | 66.42 |
| MNLI | Accuracy | 44.25 |
| MRPC | F1 | 81.96 |
| MultiRC | Accuracy | 57.30 |
| QQP | F1 | 61.47 |
| RTE | Accuracy | 58.27 |
| WSC | Accuracy | 67.31 |
| **Mean** | BabyLM task aggregation | **62.43** |
The official recipe used learning rate 3e-5, seed 42, maximum length 512, causal last-token pooling and left padding. Training lasted 10 epochs per task, except WSC at 30; batch size was 16 for BoolQ/MultiRC and 32 otherwise. These task-specific fine-tuning results are separate from the base weights in this repository.
### Human-alignment measures
| Measure | Reported score | Meaning |
|---|---:|---|
| Reading | **0.325** | Mean of eye-tracking 0.30 and self-paced reading 0.35, in evaluator reporting units |
| Age of Acquisition | **−23.76** | Correlation × 100; raw correlation **−0.2375718** |
AoA uses all 19 checkpoints to compare the order in which model vocabulary becomes predictable with human acquisition data. It describes a **training trajectory**, rather than only the 80M model. The negative correlation is an important weakness of this run. The human-like mean is **−11.72** and the nine-category overall mean is **36.38**. Reading and AoA are included there, and excluded from the NLP-only mean by definition.
For context, inserting the local NLP score into the saved **15 September 2026** Strict-Small public-table snapshot would place it **39th among 130 entries**. This is a snapshot comparison, not an official competition placement. The score, timestamp, aggregation and snapshot hash are recorded in [final_scores.json](evaluation/final_scores.json).
## Training recipe
![Sequence length, learning rate and dropout schedule over ten passes](assets/training_schedule.png)
| Setting | Value |
|---|---|
| Objective | Causal next-token prediction from scratch |
| Corpus | BabyLM 2026 English Strict-Small; 10,000,000 words per pass |
| Total exposure | 100,000,000 words; selected checkpoint at 80,000,000 |
| Tokenizer | 16,384-entry ByteLevel BPE with NFKC normalization |
| Special IDs | `<unk>` 0, `<s>` 1, `</s>` 2, `<pad>` 3, `<mask>` 4 |
| Sequence length by pass | 128, 256, 512, 128, 256, 512, 128, 256, 512, 512 |
| Batch budget | 16,384 packed token slots per update; accumulation 1 |
| Optimizer | Fused AdamW, β=(0.9, 0.95), ε=1e-8 |
| Peak LR by exposure | 0–30M: 8e-4; 30–60M: 7e-4; 60–100M: 6e-4 |
| Warmup | Linear over the first 1.6M input words |
| Weight decay | 0.1 for matrix parameters; 0 for vectors |
| Gradient clipping | Global norm 1.0 |
| FFN dropout | 0, 0.02, 0.05 at the same three exposure boundaries |
| Precision | FP32 parameters and residual geometry; BF16 matrix operations |
| Model seed / data-order seed | 20260907 / 20260904 |
| Training implementation | Regional `torch.compile`; variable-length FlashAttention |
| Hardware | One NVIDIA RTX 3090 in the author's local PC |
Documents are shuffled for each pass. Long documents receive a randomized first-chunk boundary, retaining the text on both sides; packed documents use isolated causal masks. The trainer reads every row with a causal objective. The shared data builder also records unused objective flags; those flags do not introduce masked-language-model training in VISTA.
The six source files contain BNC spoken (762,073 words), CHILDES (2,841,101), Gutenberg (2,557,721), OpenSubtitles (2,282,877), Simple Wikipedia (1,531,437), and Switchboard (24,791). Obtain the corpus through [BabyLM's official data instructions](https://babylm.github.io/guidelines.html); the corpus is not redistributed here. File checksums, preprocessing details and per-pass accounting are in [data_manifest.json](training/data_manifest.json).
The original training-loop timer recorded **20m 14s to 80M words** and **25m 18s to 100M**, with 7.27 GiB peak reserved memory. These timings exclude earlier acceptance/compilation preparation and all benchmark work. See [recorded_timing.json](training/recorded_timing.json) for the exact scope and values. This is a single-run measurement, not a general throughput claim.
The released model is a raw checkpoint. The training code also collected a tail weight average after 80M; that average is separate from these released raw weights.
## Checkpoints and BabyLM release information
The [official checkpoint convention](https://github.com/babylm-org/babylm-eval/blob/6f825c291e2c4c78ad33b1935fd64d45f52642dc/strict/README.md) is reflected in these revisions:
| Revisions | Contents |
|---|---|
| `main` | Selected 80M-word raw model |
| `chck_1M` … `chck_9M` | First optimizer update reaching each nominal milestone |
| `chck_10M`, `chck_20M`, … `chck_100M` | End-of-pass raw models |
The early 1–9M checkpoints were recovered by replaying the first pass. At 10M, the reproduced **model tensors, optimizer, cursor and RNG state matched exactly**. Checkpoint-file bytes differed due to serialization. [checkpoints.json](checkpoints.json) lists actual word counters and exported weight hashes, including the small batch-boundary overshoot of early milestones.
| Public artifact / evaluation item | Included here |
|---|---|
| Transformers likelihood model + tokenizer | Safetensors, custom source, AutoModel/AutoModelForCausalLM |
| Architecture and training specifications | This card, equations, recipe and data checksums |
| 19 named checkpoint revisions | Model weights and actual exposure counters |
| Full zero-shot results, every 10M | Ten checkpoint summaries and a machine-readable curve |
| Selected-model downstream results | Seven task scores and predictions |
| Reading and AoA | Selected-model Reading predictions; 19-point AoA surprisals and score |
| Fast results for every early checkpoint | Not included; the released weights allow this separate evaluation |
| Submission-ready collated competition JSON | Not claimed; this repository is the model and research-evidence release |
Prediction files preserve numeric scores/IDs. AoA context passages are omitted; use the official data with the supplied word/context IDs to reconstruct them. [EVALUATION.md](EVALUATION.md) describes the layout and scoring conventions.
## Reproducibility, use and limitations
See [REPRODUCING.md](REPRODUCING.md) for the source layout and prerequisites. The inference model is directly loadable. The archival training scripts record the original recipe and require environment/data-path setup before training.
The model is useful for studying a small language model's learning dynamics, routing and Attention-derived variance. It has no instruction or safety fine-tuning. Generated text can be inaccurate, biased or inappropriate; world knowledge, entity tracking and long-form reliability are limited. No application-level safety or broad multilingual evaluation has been performed. The effects of individual mechanisms require dedicated comparisons; the combined benchmark result alone does not assign each component a causal contribution.
Code and model weights are released under **MIT**. Training and evaluation datasets retain their respective licenses and access terms.
```bibtex
@misc{maeda2026vista24m,
author = {Maeda, Yusuke},
title = {VISTA-24M: A Variance-Informed 7-Layer Language Model},
year = {2026},
howpublished = {Hugging Face model repository},
url = {https://huggingface.co/AwakeningOS/VISTA-24M}
}
```