Mamb2_8B_Recall_SQ3.25
Packed 3.25-bit recurrent state plus a freshly trained Resurface adapter
for the pure nvidia/mamba2-8b-3t-4k
language model. The verified result is 7.855570 WikiText-2 validation PPL
and 271/384 (70.57%) normal multi-key recall, using 27.1797 MiB of
persistent state and table storage per sequence at batch one.
SQ3.25 describes recurrent state. The 8B model weights remain FP16. This
repository distributes the adapter, calibration table, custom runtime and
verification evidence. Download the original NVIDIA base separately. It is
not a standalone 8B checkpoint, a 3.25-bit weight model, or a Transformers/PEFT
from_pretrained package.
This is the SQ3.25 member of the Mamb2_8B_Recall family. Its adapter was trained fresh on the fixed quantized-state base; it is distinct from the original Mamb2_8B_Recall adapter.
Verified quality
| Configuration | WikiText-2 validation PPL | Normal MK | N=16 | N=64 |
|---|---|---|---|---|
| Original S16, archived reference | 7.334322057221965 | 146/384 (38.02%) | — | — |
| Fixed Q3.25 state, no adapter | 8.186186562837207 | 32/384 (8.33%) | 30/192 | 2/192 |
| Q3.25 + released Resurface | 7.855569605864836 | 271/384 (70.57%) | 181/192 | 90/192 |
| Q3.25 after adapter removal | 8.186186562837207 | 32/384 (8.33%) | 30/192 | 2/192 |
PPL covers 130 reset windows of up to 2,048 predicted tokens, with 264,764 targets in the pinned WikiText-2 validation split. MK uses 384 normal prompts and 384 paired target-removed diagnostic prompts. Both Q3.25 configurations match the removed target 0/384 times; this is not an abstention score. The normal-MK denominator excludes those controls.
Against the fixed Q3.25 base, the adapter improves PPL 4.0387% and normal MK 62.2396 percentage points: 242 paired gains and three regressions. The paired bootstrap 95% interval for the MK change is [+57.2917, +67.1875] percentage points. PPL remains 7.11% higher than the archived S16 reference, which was not rerun in this experiment.
All nine predefined quality and integrity gates passed, including the PPL threshold, positive paired recall improvement, exact control replay and unchanged persistent cache. Removing the adapter reproduced all 130 baseline window NLLs and all 768 generated sequences exactly, including reset/cache receipts. An independent CPU audit of the recorded full evidence also passed; that audit does not rerun GPU logits. See full results, raw comparison and full audit.
State format and memory
Each 128-coordinate recurrent-state row stores 32 INT8 coordinates, 32 INT4 coordinates and 64 zero-carry coordinates, plus two FP16 scales. That is 52 bytes per row, or 3.25 bits per coordinate including scales. The custom runtime uses the packed carry for every token, including prompt processing. The selected coordinate table is fixed before adapter training.
| Persistent component, batch one | Bytes | Size |
|---|---|---|
| SSM cache | 23,855,104 | 22.7500 MiB |
| Convolution cache | 4,587,520 | 4.3750 MiB |
| One shared static table | 57,344 | 0.0547 MiB |
| Cache plus table | 28,499,968 | 27.1797 MiB |
| Separate FP16 adapter payload | 2,308,208 | 2.3082 MB |
| Cache, table and adapter | 30,808,176 | 29.3810 MiB |
| Original FP16 model weights, separate | 16,473,999,360 | 16.4740 GB |
The cache/table figure is 76.64% smaller than the archived S16 cache of 116.3750 MiB. Static table and adapter storage can be shared across requests; each is counted once here. These are tensor payloads, excluding temporary computation, allocator reserve and training teacher/optimizer state. They are not total GPU memory requirements. The serialized adapter file is 2,375,743 bytes; its 1,154,104 FP16 parameters occupy 2,308,208 bytes when loaded.
Download and verify
With the Hugging Face CLI installed:
hf download EndlessChasing/Mamb2_8B_Recall_SQ3.25 \
--local-dir Mamb2_8B_Recall_SQ3.25
cd Mamb2_8B_Recall_SQ3.25
python3 scripts/release_state_resurface_v11.py --verify release
The standard-library verifier checks the exact adapter, table, raw evidence,
internal checksums and unchanged measured runtime files. Verification does
not require CUDA or download the base. The original 18-file release bundle
is preserved under release/; runtime code, scripts and notices are supplied
alongside it. This verification checks published evidence rather than running
a new quality evaluation.
Download the separate source checkpoint and tokenizer at the pinned revision:
hf download nvidia/mamba2-8b-3t-4k \
release/mp_rank_00/model_optim_rng.pt \
mt_nlg_plus_multilingual_ja_zh_the_stack_frac_015_256k.model \
--revision b915550c63ba9359f88f44d1f6a600d85af27302 \
--local-dir models/source
The loader verifies the checkpoint and tokenizer SHA-256 values, then casts the official BF16 weights to FP16. Native NVIDIA Megatron BF16 equivalence is not established by these measurements.
Environment and inference
The measured environment was Linux, Python 3.10.12, PyTorch 2.11.0+cu128, Triton 3.6.0, Mamba-SSM 2.3.2.post1, NumPy 1.26.4, datasets 4.8.5 and SentencePiece 0.2.1, on an RTX PRO 6000 Blackwell Server Edition. Install the matching CUDA/PyTorch native Mamba and causal-convolution dependencies first, then install this package from the downloaded root:
python -m pip install --no-deps -e .
python scripts/infer_state_resurface_v11.py \
--bundle release \
--source-dir models/source \
--prompt 'The capital of France is' \
--max-new-tokens 32
The entrypoint checks recorded backend versions, external source hashes,
precision flags and the fixed 16-warp RMSNorm configuration. An incompatible
backend fails before model loading. pyproject.toml contains broad dependency
ranges, not an exact environment lock; other GPU architectures are not validated.
See the release guide.
This wrapper provides greedy batch-one generation with prompts up to 4,096
tokens. Its separate GPU smoke reproduced one archived 263-token MK prompt
and all 12 continuation token IDs, with the same cache bytes. That check
validates the wrapper on that case; it does not repeat or extend the full
PPL/MK evaluation. The example prompt above is a usage example, not a reported
quality measurement. Native InferenceParams, variable-length batching and
a serving integration are not provided.
To check the package through the inference entrypoint without torch or CUDA:
python3 scripts/infer_state_resurface_v11.py --bundle release --verify-only
Training and limitations
The original 8,236,999,680 base parameters remain frozen. After fixing the selected state table, a fresh post-D Resurface-style adapter completed 1,536 successful TRAIN updates, with five numerical overflow retries. Only the final TRAIN-derived FP16 export was evaluated; validation/CONFIRM did not select a checkpoint. Training/export forward parity was checked at 128 and 512 tokens, with separately audited gradient fixtures.
The validation corpus and CONFIRM prompt families have historical exposure. The results establish performance on this specified protocol. They do not establish unseen-template recall, full-4K recall quality, other-language performance or broad language-model generalization. The wrapper's 4K prompt limit is an interface constraint, not a demonstrated 4K recall result. Source identity/version/gradient guards do not constitute a complete post-run bytewise audit of the base model. Adapter/table transfer to other checkpoints requires separate validation.
The tagged repository contains the frozen protocol, full training/evaluation reproduction guide and retained evidence. Base weights and training corpora are downloaded separately.
Provenance and licenses
This release corresponds to GitHub
v0.1.0-q325-resurface,
commit be037e1a2635ad7b540e0a1d2e551febf647c346.
- Adapter SHA-256:
339334b3431027504fb8da4ce3167d63c8c2f309920cfb0076bde1f72a115cb7. - Selected calibration SHA-256:
a77bbd86ae609d8ef7cf90f10d0904cf95b4df8e059c9ea6cb2fd12886436fb8. - Full independent audit SHA-256:
ad2b3f14a075b0d60993f3579e99fc98bfbac8aa527f6a0c8d684cc9493e7d72.
New project code and the fresh adapter are distributed under GPL-3.0. The original NVIDIA model, Mamba implementation and StateQuant reference have their separate Apache-2.0 licenses. The root GPL license does not relicense those upstream assets or WikiText text. The base checkpoint is not included. See LICENSE and third-party notices for source attribution, StateQuant provenance and WikiText licensing details.
Method references: StateQuant and Resurface. The preserved StateQuant reference comes from the user-supplied archive identified in the notices; current availability of that external repository is not required to use this packaged runtime.
Model tree for EndlessChasing/Mamb2_8B_Recall_SQ3.25
Base model
nvidia/mamba2-8b-3t-4k