|
Download README.md from EndlessChasing/Mamb2_8B_Recall: direct link, hf CLI and curl.
- Browser
- Download file 11.5 kB
-
https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/README.md
- Command line
-
hf download hf://EndlessChasing/Mamb2_8B_Recall/README.md
-
curl -L -o README.md https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/README.md
11.5 kB
| license: gpl-3.0 | |
| language: | |
| - en | |
| base_model: nvidia/mamba2-8b-3t-4k | |
| base_model_relation: adapter | |
| pipeline_tag: text-generation | |
| inference: false | |
| tags: | |
| - mamba2 | |
| - state-space-model | |
| - adapter | |
| - resurface | |
| - recall | |
| - custom-runtime | |
| datasets: | |
| - Salesforce/wikitext | |
| # Mamb2_8B_Recall | |
| ## Official WikiText-2 test PPL | |
| | Fixed published model | Official WT2 test PPL ↓ | Historical synthetic CONFIRM normal MK ↑ | | |
| | --- | ---: | ---: | | |
| | Without Resurface | 7.24453218 | 147/384 (38.28%) | | |
| | With published Resurface | **6.96281766** | **365/384 (95.05%)** | | |
| **Official test split · 147 reset windows · 300,963 next-token targets.** | |
| Original native SSD parallel prefill; no state rounding after each token. | |
| [Test results and reproduction](docs/WT2_TEST_V1_RESULTS.md) · | |
| [Paired raw evaluation](reports/wt2_test_v1/comparison.json) · | |
| [CPU audit](reports/wt2_test_v1/cpu_audit_v1.json). | |
| The MK column reports the earlier synthetic CONFIRM evaluation; it was not rerun for this WT2 test update. | |
| Historical validation PPL and synthetic CONFIRM MK results are retained below. | |
| A **post-D Resurface-style readout adapter** for the uncompressed, pure | |
| [`nvidia/mamba2-8b-3t-4k`](https://huggingface.co/nvidia/mamba2-8b-3t-4k) | |
| language model. On the verified evaluation protocol, normal numeric multi-key | |
| recall improved from **147/384 (38.28%) to 365/384 (95.05%)**, while WikiText-2 | |
| validation perplexity improved from **7.3341759473 to 7.0520635156**. | |
| This repository distributes the **2,374,143-byte adapter**, with 1,154,104 FP16 | |
| parameters. Inference requires the separately downloaded 8B base weights and | |
| the custom native runtime. The adapter is not a standalone 2.37 MB language | |
| model, a PEFT checkpoint, or an `AutoModel.from_pretrained` package. | |
| Source and reproduction code: | |
| [`EndlessChasing/mamb2_8B_Recall`](https://github.com/EndlessChasing/mamb2_8B_Recall/tree/4478c034aff9e5dab7af47efa84214287824138d), | |
| pinned at commit `4478c034aff9e5dab7af47efa84214287824138d`. | |
| ## Historical validation and synthetic CONFIRM MK | |
| | Model | WikiText-2 validation PPL | Synthetic CONFIRM MK | N=16 | N=64 | Target removed | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | Base model in FP16 runtime | 7.334175947318572 | 147/384 (38.28%) | 113/192 | 34/192 | 0/384 | | |
| | Base model + this adapter | **7.052063515606311** | **365/384 (95.05%)** | **191/192** | **174/192** | 0/384 | | |
| PPL covers all **130 reset windows**, up to 2,048 next-token targets per window, | |
| and **264,764 targets** in the pinned WikiText-2 validation split. Each window | |
| starts with zero state and uses parallel SSD scan computation. MK uses three | |
| synthetic numeric templates, 16 or 64 key/value bindings, and prompt lengths | |
| of 257–1,236 tokens. Generation is greedy, at most 12 new tokens, over the full | |
| 256K vocabulary; scoring matches the first standalone six-digit integer. | |
| The 384 normal prompts have 384 paired target-removed controls. MK uses native | |
| prefill followed by recurrent decoding with fresh FP16 cache. | |
| A fresh public GitHub clone was replayed on an RTX PRO 6000 Blackwell on | |
| 2026-09-28. Across the two arms, all **260 per-window NLLs and 1,536 MK | |
| generation records matched exactly**. An independent audit recalculated the | |
| metrics, decoded generated token IDs and checked regenerated input identities. | |
| See the [verification report](https://github.com/EndlessChasing/mamb2_8B_Recall/blob/4478c034aff9e5dab7af47efa84214287824138d/docs/VERIFICATION.md) | |
| and its linked raw receipts. | |
| ## Method and training | |
| The base has 56 Mamba2 blocks, 8 SSM groups, 8,236,999,680 parameters, and no | |
| attention or MoE blocks. Its official BF16 weights are cast to FP16 by the | |
| native runtime and frozen. The adapter modifies readout at the gated RMSNorm | |
| inputs, after the D skip term, with a learned head-mixing residual and a soft | |
| sigmoid gate. The same gate operates on prose and recall inputs. It adds no | |
| recurrent cache. | |
| This is a Resurface-style adaptation based on the | |
| [Resurface reference](https://github.com/Oso1106/Resurface-Multi-Binding-Recall-Is-Latent-in-Mamba-s-State), | |
| with a different insertion site. It is not an exact reproduction of the | |
| reference's private checkpoint, data or code. | |
| Training completed 1,536 successful AdamW updates with FP32 adapter masters, | |
| one pass over 1,536 synthetic numeric TRAIN examples, and paired prose from | |
| 448 pinned WikiText-2 TRAIN windows. The objective combines recall answer CE, | |
| prose CE, KL to a separately loaded frozen base teacher, and a prose gate | |
| closure penalty. Six overflow retries occurred within the declared budget. | |
| The final step was the sole candidate. Full settings and source hashes are | |
| in the [frozen protocol](https://github.com/EndlessChasing/mamb2_8B_Recall/blob/4478c034aff9e5dab7af47efa84214287824138d/docs/PROTOCOL.md). | |
| ## Download and use | |
| Use a Linux CUDA environment. The tested environment was Python 3.10.12, | |
| PyTorch 2.11.0+cu128, Mamba-SSM 2.3.2.post1, Triton 3.6.0, NumPy 1.26.4, | |
| datasets 4.8.5 and SentencePiece 0.2.1. Install CUDA-compatible PyTorch first; | |
| building Mamba-SSM requires a matching CUDA toolkit/compiler. | |
| ```bash | |
| # Download the small adapter, runtime, scripts and verification receipts. | |
| hf download EndlessChasing/Mamb2_8B_Recall --local-dir mamb2-recall | |
| cd mamb2-recall | |
| python -m pip install -e . | |
| python -m pip install --no-build-isolation 'mamba-ssm==2.3.2.post1' | |
| # Separately download the original checkpoint and tokenizer (~16.48 GB). | |
| hf download nvidia/mamba2-8b-3t-4k \ | |
| release/mp_rank_00/model_optim_rng.pt \ | |
| mt_nlg_plus_multilingual_ja_zh_the_stack_frac_015_256k.model \ | |
| --revision b915550c63ba9359f88f44d1f6a600d85af27302 \ | |
| --local-dir models/source | |
| ``` | |
| The base loader verifies its pinned checkpoint/tokenizer hashes. Verify the | |
| adapter and use the project's runtime explicitly: | |
| ```python | |
| from pathlib import Path | |
| import hashlib | |
| import torch | |
| from mamba2_recall import runtime, resurface_native | |
| from mamba2_recall.evaluation import generate_greedy | |
| adapter = Path("artifacts/source_resurface_v1/adapter_fp16.pt") | |
| assert hashlib.sha256(adapter.read_bytes()).hexdigest() == ( | |
| "e8b2b4dfe69f8e85dc9e147c9aeaa3297cff14f4043e558bb1795ad476c1fca0" | |
| ) | |
| torch.backends.cuda.matmul.allow_tf32 = False | |
| torch.set_float32_matmul_precision("highest") | |
| tokenizer = runtime.SentencePieceTokenizer("models/source") | |
| model = runtime.load_source_model("models/source") | |
| payload = resurface_native.read_fp16(adapter) | |
| assert payload["binding"]["source_checkpoint_sha256"] == runtime.SOURCE_CHECKPOINT_SHA256 | |
| prompt = "Key-value records:\n123456: 654321\n234567: 765432\n\nLookup the value for key 123456.\nValue:" | |
| with resurface_native.install_fp16( | |
| model, adapter, expected_binding=payload["binding"] | |
| ): | |
| text, token_ids, prompt_tokens, cache_bytes = generate_greedy( | |
| model, tokenizer, prompt, max_new_tokens=12, execution="prefill" | |
| ) | |
| print(text) | |
| ``` | |
| This example shows the loading API; the small example prompt is not a reported | |
| benchmark result. Measured peak PyTorch CUDA allocation during paired | |
| evaluation was **17.08 GB** and during two-model training was **34.59 GB**. | |
| These measurements exclude some device and host overhead and are not minimum | |
| hardware requirements. Full base weights remain resident during inference. | |
| ## Reproduce historical validation PPL and synthetic CONFIRM MK | |
| From the downloaded repository after the commands above: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES= python scripts/prepare_data.py \ | |
| --split confirm --source-dir models/source | |
| confirm_manifest_sha=$(python -c "import hashlib; from pathlib import Path; print(hashlib.sha256(Path('training_data/numeric_v1/confirm/manifest.json').read_bytes()).hexdigest())") | |
| python scripts/evaluate.py \ | |
| --source-dir models/source \ | |
| --adapter artifacts/source_resurface_v1/adapter_fp16.pt \ | |
| --data-root training_data/numeric_v1 \ | |
| --eval-manifest-sha256 "$confirm_manifest_sha" \ | |
| --split confirm \ | |
| --report reports/hf_adapter_evaluation.json | |
| ``` | |
| This evaluates the base and serialized adapter in the same process, with full | |
| WikiText-2 validation PPL and both synthetic CONFIRM MK conditions, then checks baseline restoration after removing | |
| the adapter. The output path must be new. The preparation script and evaluator | |
| validate frozen data content; use the newly generated manifest's SHA because | |
| its Python build metadata may vary. See the | |
| [reproduction guide](https://github.com/EndlessChasing/mamb2_8B_Recall/blob/4478c034aff9e5dab7af47efa84214287824138d/docs/REPRODUCE.md) | |
| for training and the independent saved-report audit. | |
| `resurface_config.json` describes this custom adapter format and its base | |
| binding; it is not a Transformers or PEFT configuration. The release includes | |
| `checksum_manifest.json` for checking its packaged files. | |
| ## Limitations | |
| - The numeric CONFIRM instances/template family and WikiText-2 validation | |
| corpus had been observed during earlier project development. The reported | |
| results establish performance on this specified replay protocol, rather | |
| than an untouched holdout. TRAIN and CONFIRM numeric identities were | |
| separately checked to be disjoint. | |
| - Unseen template families, longer recall distances, full 4K recall, other | |
| languages, broad downstream task quality and untouched prose have not been | |
| established by these measurements. | |
| - Target-removed 0/384 means the absent answer was not recovered; it does not | |
| measure calibrated refusal behavior. | |
| - The runtime casts the official BF16 source to FP16. Native Megatron BF16 | |
| parity and tokenwise FP16-state-rounding PPL have not been established. | |
| - Exact replay was observed in the recorded environment. Other GPU, CUDA, | |
| PyTorch or Mamba kernel versions can change numerical outputs. | |
| - The published adapter is bound to the specified uncompressed source. | |
| Transfer to a quantized or different checkpoint needs separate validation. | |
| ## License and attribution | |
| The `gpl-3.0` card metadata reflects this release's existing GitHub repository | |
| license; see the bundled `LICENSE`. The custom runtime derives from the | |
| associated [E8/W5 project](https://github.com/EndlessChasing/mamba2-8b-e8w5). | |
| The original NVIDIA model is downloaded separately under its Apache-2.0 | |
| license, and its own model card and terms apply to those base weights. | |
| This release does not redistribute those weights. The metadata does not | |
| relicense the NVIDIA model or create an additional adapter-specific license. | |
| Method credit: [Resurface — Multi-Binding Recall Is Latent in Mamba's State](https://github.com/Oso1106/Resurface-Multi-Binding-Recall-Is-Latent-in-Mamba-s-State). | |
| Implementation and experiment details are available in the pinned GitHub | |
| repository linked above. | |
| ## Additional official test files | |
| This main-branch docs overlay adds the new test evidence and helpers. If you | |
| downloaded an original release tag, fetch the additional small files from main: | |
| ```bash | |
| hf download EndlessChasing/Mamb2_8B_Recall \ | |
| --revision main --local-dir . \ | |
| --include 'scripts/*recall*wt2*test*v1.py' \ | |
| --include 'docs/RECALL_NATIVE_PREFILL_WT2_TEST_V1*' \ | |
| --include 'docs/WT2_TEST_V1_RESULTS.md' \ | |
| --include 'reports/wt2_test_v1/*' | |
| ``` | |
| See [test results and reproduction](docs/WT2_TEST_V1_RESULTS.md). Original release manifests | |
| and checksum lists refer to the original tagged bundle. The new | |
| [overlay checksums](reports/wt2_test_v1/overlay_checksums.json) bind the small update. | |
| The YAML metadata contains no model-index benchmark values; PPL tables | |
| explicitly state their split. | |