EndlessChasing's picture
Release complete independent W4A16 Mamba2-8B with Resurface adapter
cf0f656 verified
|
Raw History Blame Contribute Delete
5.67 kB
# Independent W4A16 Mamba2-8B Resurface protocol
Declared 2026-09-28 before quantization, training or quality evaluation of this repository's independent W4 candidate. The user requires redistributable weights. No Quamba weights, rotation, implementation or adapter are dependencies.
## Frozen model and quantization
Source: NVIDIA `nvidia/mamba2-8b-3t-4k`, Apache-2.0, revision
`b915550c63ba9359f88f44d1f6a600d85af27302`, checkpoint SHA256
`47c2766f6aad89d73beafbeaecb334aab902d7370906d081764a90bb7a8bbbcb`;
tokenizer SHA256 `5862e2f71caf762bc9845662be5fec2867deb58d874568235a02a36c5111cd09`.
This is the pure 56-layer Mamba2-8B model, with untied 256K embedding/head.
Independently quantize all 114 large matrices (56 input projections, 56 output
projections, embedding and head) to affine uniform four-bit codes. Groups contain
128 adjacent input-axis weights. Each group stores a FP16 scale and FP16 offset;
two codes are packed per byte. Metadata adds 0.25 bits per quantized weight,
so these matrices require approximately 4.25 bits/weight before file headers.
The other 393 small tensors remain FP16. Do not describe every parameter as INT4.
For each group, fit using weight-space MSE only, over fixed range factors
1.00, .99, .98, .97, .96, .95, .94, .92, .90 around its min/max midpoint.
Compare actual serialized FP16 scale/offset and FP16 reconstruction. No prose,
MK data, validation labels or source activation data enter quantization.
This fixed quantizer defines one candidate; no evaluation-based quantizer search.
Constant/zero groups, nonfinite inputs and truncated files must be handled explicitly.
Write/read actual packed files, verify exact reconstruction and per-file hashes,
and bind the resulting manifest to both training and evaluation.
Quality execution decodes the saved W4 package into frozen FP16 native weights.
This is a W4A16 numerical reference, not a packed-resident GPU kernel. Native
FP16 residual orchestration and FP16 state cache match our existing source control.
The source BF16-to-FP16 runtime is not asserted to reproduce Megatron BF16 exactly.
No rotation, residual weight correction or small-tensor readaptation is included.
Use the same post-D adapter as the compressed experiment at all 56 native
gated RMSNorm inputs: `y + sigmoid(w dot u+b) * g*(V@y)` across 128 heads.
Its 224 tensors contain 1,154,104 FP32 trainable master parameters, exported
as FP16 for inference. Initialize V=0, g=1, w=0, b=-4; use soft sigmoid at
train and eval, no task switch, EMA, or new recurrent cache. The adapter is
external to the frozen base; check its identities and absence of base grads.
## Training
Reuse the *same exact data generator contract and seeds* as the compressed
adapter: 1,536 TRAIN numeric bindings over three templates and N=16/64,
disjoint six-digit key/value intervals; full 256K CE on every answer suffix
token. Use seed 2026092803 and `torch.randperm(1536)` once for example order.
Each step pairs one MK example with a 512-token segment from the already
prepared 448 WikiText-2 TRAIN windows. Their historical manifest and file
SHA256 are pinned by the train command. Heldout and validation windows are
not included in training. Reusing this fixed TRAIN text gives the control the
same training exposures as the compressed arm.
Use a second, independently loaded **unadapted W4-decoded FP16** model as the frozen
prose teacher. This is the same own-unadapted-base teacher policy as our previous experiments. The fixed per-step loss is
`MK answer CE + 0.5 prose CE + 0.5 KL(own unadapted teacher || student) + 3 C`,
where C is the same prose-only router closure, with budget 0.006 and excess
coefficient 10. Temperature 1, 511 prose targets per step, 64-token staged
head chunks, and one optimizer update after both task gradients.
AdamW FP32 masters: V/g LR 1e-4, w/b LR 3e-4, betas (.9,.999), eps 1e-8,
weight decay 0, clip norm 1. Multiplicative schedule
`0.1+0.9*(1+cos(pi*j/1535))/2` at successful index j. Native FP16 forward,
block checkpointing, gradient scale 1024 with growth interval 2000. Overflow
retries the exact same pair without optimizer/master update, at most 8 retries
and 1,544 total attempts. Run 1,536 **successful** updates; the final step is
the only candidate. Small GPU smokes are discarded before formal training.
## Evaluation and interpretation
After training, restore the actual serialized FP16 adapter onto the same
W4-decoded base. Evaluate baseline, adapter enabled and restored baseline
in a single process, with identical tokenizer, MK prompts and full WikiText-2
validation windows. For MK use full 256K greedy generation up to 12 tokens,
first standalone six-digit number, normal and target-removed controls, native
prefill plus recurrent decode with fresh FP16 cache. For PPL use all 130
nonoverlapping reset windows and all 264,764 next-token targets. Record all
raw scores, differences and file hashes; no early selection or task switch.
The earlier compressed arm's independent CONFIRM set has now been observed and
uses the same three template families. Running this W4 experiment on it gives
a matched historical comparison, **not a newly untouched holdout**. The
WikiText-2 validation text also informed earlier development. We will not
promote a W4+adapter result to a general recall conclusion without new
templates, longer distances and a truly untouched corpus.
Report the new W4 base and W4+adapter together, with matched source FP16 historical metrics clearly identified as historical. Verify prompt/window hashes before comparing them. Report PPL and recall together; no packed-runtime speed or memory claim follows from these quality measurements.