--- language: en license: apache-2.0 library_name: transformers pipeline_tag: text-generation inference: false tags: - babylm - babylm-2026 - strict-small - linear-attention - state-tracking - delta-rule - custom_code --- # bind2_0 > ## ⚠️ The official benchmark badge is NOT the point of this repo > > On the official BabyLM 2026 strict-small zero-shot surface this model is **statistically tied** with its > matched monolithic and bind1 controls (mean-4 excl. entity: **56.92** vs 56.60/56.63 at 23.9M; slightly > above the GPT-2 baseline on BLiMP, **66.11** vs 65.08). It wins nothing there, and that is part of the > finding. > > **What this stage actually shows** (three results, all kept): > 1. **The mechanism is trainable**: with direct-task training on a purpose-built synthetic swap-tracking > task (n=800 per eval, 5-way, chance 0.20), the forced-bottleneck architecture reached **0.9988 > accuracy** — with a sharp grokking transition between 5M and 10M training tokens > (0.179 → 0.969 → 0.996) — while matched monolithic, bind1-style, and no-binding controls stayed at > chance (0.2125 / 0.1938 / 0.1938). > 2. **It does not emerge from plain LM pretraining**: after standard BabyLM strict-small pretraining, > zero-shot give-verb state-tracking probes (n=60, chance 0.50) show **no model above chance**; this > build (23.9M) scored 0.200 — significantly below chance, i.e. strong recency capture. The synthetic > grok did not transfer. > 3. **The architecture costs ~nothing on general language** ("no tax, no win"): official zero-shot scores > tied across architectures; raw LM perplexity is slightly worse than the monolithic control > (11.0 vs 9.7 at 23.9M), as expected for a forced bottleneck. > > The full falsification-timeline context (what came before this stage and what it forced next) lives at > the hub: [`SecludedCorner/bind-evolution`](https://huggingface.co/datasets/SecludedCorner/bind-evolution). ## Model description (family level) bind2_0 is a small causal LM (main branch = **23.9M params**; branch `27m` = **27M params**) combining: - **Delta-rule fast-weight memory**: gated delta-rule recurrent value dynamics (GatedDeltaNet), using the third-party MIT-licensed [`flash-linear-attention` (fla)](https://github.com/fla-org/flash-linear-attention) implementation as the core recurrent layer. The delta-rule/fast-weight design follows Yang, Kautz & Hatamizadeh, *Gated Delta Networks: Improving Mamba2 with Delta Rule* (ICLR 2025, arXiv:2412.06464); only the permissively licensed fla implementation is used here. - **A forced bottleneck**: attention is chunk-local; information can cross chunk boundaries **only** through the recurrent state S. This makes the recurrent state the sole carrier of long-range bindings — the design hypothesis under test at this stage. Later family members are not described here; see the hub for the family narrative. ## Training data Official **BabyLM 2026 Strict-Small** corpus (the provided ~10M-word text-only corpus; no custom data). Training: 150M tokens over the 16.3M-token encoding (SEQ256, batch 16, vocab 16k), recurrent state reset per block. Final training perplexity: 11.0 (23.9M build), 10.8 (27M build); matched monolithic control: 9.7. ## Results (official pipeline, strict-small zero-shot, single seed) Main branch (23.9M): | task | bind2_0 | mono control | bind1 control | GPT-2 baseline | |---|---:|---:|---:|---:| | BLiMP | 66.11 | 65.35 | 65.50 | 65.08 | | BLiMP supplement | 58.11 | 58.17 | 58.35 | 57.25 | | EWoK | 51.95 | 51.32 | 51.57 | — | | entity_tracking (filtered) | 19.02 | 21.16 | 19.22 | 21.07 | | COMPS | 51.49 | 51.55 | 51.11 | 51.81 | | mean(4, excl. entity) | **56.92** | 56.60 | 56.63 | — | Branch `27m` (27M; mono control at this tier is 27.4M): | task | bind2_0 | mono control | bind1 control | GPT-2 baseline | |---|---:|---:|---:|---:| | BLiMP | 65.14 | 64.35 | 66.68 | 65.08 | | BLiMP supplement | 60.81 | 58.55 | 60.90 | 57.25 | | EWoK | 51.16 | 50.70 | 51.90 | — | | entity_tracking (filtered) | 20.53 | 19.24 | 20.00 | 21.07 | | COMPS | 50.89 | 51.00 | 51.36 | 51.81 | | mean(4, excl. entity) | 57.00 | 56.15 | 57.71 | — | entity_tracking is ~chance for every model under the current filtered standard (non-discriminative), hence excluded from the mean. The mean-of-4 is NOT the official leaderboard "Overall" (which also weights GLUE, reading, AoA, and more). All numbers single-seed; the across-architecture spread (~1.5pp) is within seed noise. ## Honest limitations — what this stage cannot do - **It does not track state zero-shot.** After plain LM pretraining, give-verb state-tracking probes are at or below chance (0.200 at 23.9M = strong recency capture). Do not use this model expecting emergent entity/state tracking. - **It does not beat its controls on the official benchmark.** Tied within noise; that is the honest reading, not modesty. - **The synthetic grok required direct-task training** — it is evidence the bottleneck can force state into the recurrent path, not evidence of a general capability. - Single seed per build; raw LM perplexity pays a small bottleneck tax (11.0 vs 9.7). ## What this stage forced next The gap between "trainable in principle (0.9988 synthetic grok)" and "does not emerge from LM pretraining (chance zero-shot)" forced the next question on the ladder: split the confound — first prove the mechanism is *causally real at depth* under a pre-registered gate, separately from transfer. That question — including a preregistered NULL we report as NULL, and the causal evidence around it — is answered on the hub: [`SecludedCorner/bind-evolution`](https://huggingface.co/datasets/SecludedCorner/bind-evolution). ## How to load ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("SecludedCorner/bind2_0", revision="<40-char commit SHA>") model = AutoModelForCausalLM.from_pretrained("SecludedCorner/bind2_0", revision="<40-char commit SHA>", trust_remote_code=True) ``` - `trust_remote_code=True` is **required** for the model (the config's `auto_map` points at the inlined `modeling_babylm.py` shipped in this repo). The tokenizer loads without it. - **A CUDA GPU with a working Triton is required for inference.** The fla 0.5.1 GatedDeltaNet path dispatches to Triton kernels; on CPU-only environments the forward pass fails at kernel launch (`RuntimeError: 0 active drivers`). The checkpoint itself loads fine on CPU; only the forward pass needs a GPU. ## Dependency pins Exact versions the checkpoints were trained/exported/validated with (Python 3.11.15; also shipped as `requirements_pins.txt` in this repo): ``` torch==2.12.1+cu126 transformers==5.13.0 triton-windows==3.7.1.post27 flash-linear-attention==0.5.1 fla-core==0.5.1 safetensors==0.8.0 numpy==2.4.6 ``` Notes: - **fla** is required at runtime: the inlined modeling code lazily imports `fla.layers.GatedDeltaNet`, which this architecture instantiates. Installed from PyPI as release **0.5.1** (no commit pin — the environment records the plain PyPI release; `fla-core` 0.5.1 comes with it). - **triton**: the training/export environment is Windows and uses the `triton-windows` fork (3.7.1.post27); on Linux use the upstream `triton` matching your torch build. - **torch** build is CUDA 12.6 (`+cu126`); pick the equivalent build for your platform. ## Export fidelity & known-defect disclosure **Export fidelity (verified twice):** - At grid-evaluation time (2026-07-12) the HF export was verified logit-identical to the training model (wrapper-vs-original logit diff = 0.00e+00). - On 2026-07-15 a dedicated round-trip validation re-checked both exports, as they sit on disk, against the original training checkpoints: **all 186 weight tensors bitwise identical** (max abs diff 0.0, no missing/extra keys) and **logits bitwise identical** (max abs diff 0.00e+00 across 4 deterministic batches of 8×128 tokens, fp32, passing at both atol 1e-4 and atol 1e-5) for **both** the 23.9M and 27M builds. Caveat, disclosed: that re-check ran on CPU, where fla's Triton kernels cannot execute, so three fla components were replaced by math-equivalent pure-PyTorch implementations applied identically to both sides. It therefore validates **export fidelity** (weights and module wiring survive `.pt → safetensors → AutoModelForCausalLM` exactly), not Triton-kernel numerics; a GPU re-run with stock kernels remains the gold check. **Known defect — `attention_mask` is accepted but ignored:** - The exported wrapper accepts `attention_mask` in `forward()` and never uses it — on the causal-LM path and, for this architecture, on the AutoModel (sequence-classification) path as well. Empirically, `attention_mask=ones`, `=zeros`, and omitted all produce bitwise-identical logits on both builds. - **Consequence:** in a batch, right-padding is silently attended over as real tokens — **batched padded inference gives wrong results.** Run unbatched, or length-sorted/unpadded. Per-example inference is unaffected; the published zero-shot numbers above came from the per-example evaluation setting and are unaffected by this defect. ## How to cite this model Always cite at a pinned revision: pass `revision="<40-char commit SHA>"` to `from_pretrained`, or use the `/tree/` URL form. Authoritative per-branch SHAs are recorded at push time in the project `PUBLISH_LEDGER`; the final SHAs are noted in a dated addendum below after publication. ## Card freeze policy **The body of this card is frozen at publish.** Any later information (including the final commit SHAs and resolved links) is added only as clearly dated addendum sections below this line — the text above is never silently edited. ## Branches - `main` — 23.9M-parameter build (the primary artifact) - `27m` — 27M-parameter build (same architecture and recipe, wider)