--- language: - en pipeline_tag: text-generation base_model: nvidia/mamba2-8b-3t-4k base_model_relation: quantized license: other license_name: mixed-model-and-software-licenses license_link: https://huggingface.co/EndlessChasing/Mamba2_3GB_ReCall_Recovered/blob/main/LICENSES.md datasets: - Salesforce/wikitext tags: - mamba2 - quantized - e8 - associative-recall - custom-code - research inference: false --- # Mamba2_3GB_ReCall_Recovered **Research prerelease:** an independently compressed and adapted version of [NVIDIA Mamba-2 8B](https://huggingface.co/nvidia/mamba2-8b-3t-4k), with an enabled soft Resurface-inspired recall adapter. This is a base language model, not an instruction-tuned assistant. It requires the supplied custom loader; it is not a standard Transformers `from_pretrained` checkpoint. This repository mirrors the exact files of [GitHub release v0.2.0-resurface](https://github.com/EndlessChasing/mamba2-8b-e8w5/releases/tag/v0.2.0-resurface) under [`release/`](release/), preserving the original manifest hierarchy. ## Published quality summary | Fixed released configuration | Official WikiText-2 test PPL ↓ | Historical synthetic CONFIRM normal MK ↑ | | --- | ---: | ---: | | Compressed/readapted base, Resurface off | 7.53418476 | 91/384 (23.70%) | | Same frozen base, released Resurface on | **7.50295894** | **340/384 (88.54%)** | The PPL column uses the complete WikiText-2 **test** split: 147 reset windows and 300,963 scored tokens. [The fixed-release comparison](https://github.com/EndlessChasing/mamba2-8b-e8w5/blob/main/evaluation/wt2_test_v1/comparison.json) and [independent audit](https://github.com/EndlessChasing/mamba2-8b-e8w5/blob/main/evaluation/wt2_test_v1/cpu_audit_v1.json) were completed after publication, without selecting a new checkpoint. The MK column is the earlier synthetic CONFIRM result on 384 normal prompts; MK was not rerun for the test update. The two columns have different evaluation data and should not be treated as one newly paired test. ## Model and storage The complete pure Mamba-2 architecture is retained: **56 blocks, width 4096, eight SSM groups, 256K vocabulary, untied embedding/head, 8,236,999,680 base parameters**, plus 1,154,104 adapter parameters. | Component | Representation | | --- | --- | | 112 input/output projections | Rotated E8P12 plus signed-axis residual: 20 index bits per 8 values, nominal 2.5 bits/value; scales/transforms/headers counted separately | | Embedding / output head | Group 128 W4 / W5, with FP16 scales | | 393 other base tensors | Readapted FP16 | | 224 adapter tensors | FP16 soft gated cross-head readout at all 56 layers | Encoded base-plus-adapter data occupy **3,141,468,439 bytes (3.141 GB)**. The original 26-asset GitHub bundle totals **3,165,987,804 bytes (3.166 GB)**, including tokenizer, software, licenses and reports; this Hub card is additional. The adapter file is 2,539,647 bytes. The E8HUF001 container stores raw members in this version, with **no additional Huffman compression**. **“3GB” describes stored model data, not GPU memory.** The reference loader expands weights to FP16. Its archived-source generation smoke peaked at 22.064 GB allocated GPU memory on a 96 GB RTX PRO 6000; this is not a minimum-VRAM test. Batch-one recurrent/conv cache was 122,028,032 bytes (116.375 MiB). There is no compressed-resident GPU kernel or demonstrated ASIC/neuromorphic deployment. ## Measured quality | Complete WikiText-2 validation | PPL | | --- | ---: | | Original NVIDIA weights cast to FP16 | 7.334175947 | | Compressed/readapted base | 7.622396588 | | Same base with released soft adapter | **7.593163114** | The adapter improves paired compressed-base PPL by 0.383521%; candidate PPL remains 3.531237% above original FP16. Evaluation covers 130 reset windows and 264,764 targets. The source number is an aligned historical evaluation; current/adapter/restored runs were paired. This validation corpus informed development and is not an untouched test set. | Independent numeric-binding CONFIRM | Recall | Target-removed false matches | | --- | ---: | ---: | | Compressed/readapted base | 91/384 (23.6979%) | 0/384 | | Same base with released adapter | **340/384 (88.5417%)** | **0/384** | CONFIRM uses new numeric instances, three shared template families and 16/64 records per prompt. The paired gain is 64.84375 percentage points, with 251 gains and 2 losses. The uncompressed source was **not** evaluated on CONFIRM. Historical DEV recall was 167/384 for original FP16 versus 346/384 for the trained candidate; these models had unequal adaptation budgets. The adapter operates **after native `D*x`, before grouped gated RMSNorm**. The same soft policy remains enabled for recall and PPL, without extra recurrent state. It was trained for 1536 updates on synthetic numeric bindings plus prose CE/KL and gate-closure regularization; the prose teacher was the frozen compressed base. All 507 base tensors stayed unchanged during adapter training. This is a post-D Resurface-inspired variant, not an exact original-method port. See the [training protocol](https://github.com/EndlessChasing/mamba2-8b-e8w5/blob/v0.2.0-resurface/docs/RESURFACE_READAPTED_TRAINING_PROTOCOL.md) and [full results](https://github.com/EndlessChasing/mamba2-8b-e8w5/blob/v0.2.0-resurface/docs/RESURFACE_READAPTED_RESULTS.md). An initial DEV run failed historical cross-process output equality; that failure is preserved. A disclosed continuation removed that prerequisite, retaining the same candidate and quality thresholds. Same-process restored controls, complete PPL and subsequent fresh CONFIRM passed. Cross-process bitwise determinism remains unresolved. No unseen-template, longer-context, general task parity, global-smallest or equal-training-budget claim is made. ## Download, verify and run Use the official `hf` CLI. Restore needs Python 3.10+, PyTorch and a C++17 compiler; generation needs compatible Linux/CUDA and the public Mamba runtime. Allow additional disk space for restored data and the temporary joined container. ```sh hf download EndlessChasing/Mamba2_3GB_ReCall_Recovered \ --revision v0.2.0-resurface --local-dir ./model-download mkdir model-software unzip model-download/release/source.zip -d model-software python3 -m pip install -e ./model-software python3 -m pip install 'mamba-ssm==2.3.2.post1' --no-build-isolation python3 model-software/scripts/package_release.py verify \ --release-dir ./model-download/release \ --expected-manifest-sha256 da5931dc8315bf576b773abdf4c77828a2994ccaf4fd14858c7798235b2ef19d python3 model-software/scripts/package_release.py restore \ --release-dir ./model-download/release --output ./restored-model \ --expected-manifest-sha256 da5931dc8315bf576b773abdf4c77828a2994ccaf4fd14858c7798235b2ef19d python3 -m mamba_e8w5.release_generate \ --model-dir ./restored-model/raw \ --prompt "The capital of France is" --max-new-tokens 12 \ --repeat --report ./generation-receipt.json ``` The loader verifies all 507 decoded base and 224 adapter tensor hashes and always installs the soft adapter. Original full-precision weights, Hessians and training data are unnecessary. Keep reports outside `release/` and `restored-model/raw/`. See the [environment and release guide](https://github.com/EndlessChasing/mamba2-8b-e8w5/blob/v0.2.0-resurface/docs/DOWNLOAD.md) for dependency details. Context plus generation must not exceed 4096 tokens. ## Attribution and license scopes - Upstream NVIDIA model weights and tokenizer: **Apache-2.0**, pinned revision `b915550c63ba9359f88f44d1f6a600d85af27302`. This independently modified model is not endorsed by NVIDIA. - Repository software and QuIP#-derived E8/LDLQ components: **GPL-3.0**. Corresponding source and notices accompany the release. The mixed-license metadata does not replace the component license texts or imply that applying a GPL quantizer automatically relicenses every model weight. - Native Mamba runtime: **Apache-2.0**. The new public adapter implementation and trained artifact are distinct from excluded private reference materials. - WikiText was obtained separately under its upstream terms; its text and tokenized passages are not included in this distribution. Read the complete [license scope and attribution](https://github.com/EndlessChasing/mamba2-8b-e8w5/blob/v0.2.0-resurface/licenses/THIRD_PARTY.md) and the original license texts shipped under [`release/licenses/`](release/licenses/).