CPU loader findings (Transformers 5.16.1) and a complete runnable tiny fixture
Thanks for publishing this fixture and clearly documenting its intentional omissions. We evaluated it for a CPU-only Quant Fidelity Suite capture/reproducibility demo. I wanted to share the actual loader findings and the small alternative we built, in case either helps.
What we tested
Revision: 75206ec3987bcce2ac1ab2ebb7cf8a50d318c9f3
Runtime: PyTorch 2.11.0+cpu, Transformers 5.16.1, using QFS's hf_capture.load_model(..., "cpu", "float32").
The safetensors file downloaded and its SHA-256 verified correctly. The issue was compatibility with the declared native GlmMoeDsaForCausalLM architecture, not file corruption or CPU capacity:
- Stored
self_attn.q_proj,k_proj, andv_projweights were reported as unexpected; the instantiated architecture expects the MLA projections (q_a_proj,q_b_proj,kv_a_proj_with_mqa,kv_b_proj) and associated norms. - Each attention
o_proj.weightwas[64, 64]in the checkpoint versus[64, 1024]expected by the model under this config/version. Loading raised a size-mismatch error. - Required DSA indexer tensors, router correction buffers, and
lm_head.weightwere missing. The config hastie_word_embeddings: false, so the missing head is not automatically resolved by declared weight tying. AutoTokenizercould not instantiate a tokenizer from the placeholder metadata; there is no vocabulary/tokenizer serialization in the repository.
Several omissions are already disclosed in your card. These findings are specifically about using the artifact for a complete forward/capture, not about its usefulness for config parsing, tensor-header inspection, or planner tests. We have not tested every Transformers version.
Why we used a different fixture
For our use case, silently initializing absent weights or ignoring mismatched shapes would mean capturing a different model from the published checkpoint. It would also make a zero-KL self-comparison a misleading success criterion. We needed clean loading, a working tokenizer, and two independent cold captures, followed by forced numerical KL replay.
We therefore built a complete native-class fixture rather than repairing this checkpoint in the loader:
malaiwah/glm-moe-dsa-tiny-random-bf16
- 277,824 parameters; 571 KB checkpoint.
- Four layers, hidden width 64, eight routed experts/top-2, and full/shared/full/shared DSA indexers.
- Complete untied output head and a working 260-token byte-level tokenizer.
- Native
save_pretrainedoutput; no remote code or missing/mismatched-weight overrides. - BF16 artifact, with native router correction buffers retained FP32, to match QFS's BF16 hidden-capture contract. That dtype choice is a QFS requirement, not an assertion that your FP32 storage is intrinsically wrong.
- Clean QFS loading report and two fresh CPU captures over 252 scored positions, with identical capture-content digests and forced numerical KL = 0.0.
The generator, pinned environment, reproduction commands, and full capture/comparison evidence are public. We also re-downloaded the published artifacts anonymously and reproduced the result. Like yours, this is random-init testing material, not a trained model or a model-quality claim.
Sharing as a possible reference for completing the native loading path, or as a reusable fixture if that saves you work. Thanks again for making the original fixture available.