Add Qwen3.6-35B-A3B Jacobian-lens adapter
Adds a Miru Tracer Jacobian-lens adapter for Qwen/Qwen3.6-35B-A3B.
Contributor: Lucas Teske at Teske's Lab.
Fit details
- Base model:
Qwen/Qwen3.6-35B-A3B(Apache-2.0) - Model revision:
995ad96eacd98c81ed38be0c5b274b04031597b0 - Model architecture SHA-256:
18c2b5747dfced7887ec8af29a8cdada3637cf4c88f33f99a78c65c2cb207bf5(miru-semantic-v1) - Model config SHA-256:
8842f9c349baaea67fe949237518f5161c41c680849c188cfdc32e01f79118e8(transformers-json-v1) - Tokenizer SHA-256:
522cccf6b8e6164ea42f0c04a13599c64d784e6e307b4272164eed8c4607ad9e - Calibration corpus: frozen prompt sequence selected from
Salesforce/wikitext,wikitext-103-raw-v1, train split, dataset revisionb08601e04326c79dfdd32d625aee71d232d685c3 - Prompt-sequence SHA-256:
dcb2a738953033139ae35f3665271ccdb795409c6073a85ea46f02a852143cc6 - Processed-prefix SHA-256:
5a85854c042b31dcce791cdddc5fc0d891508d0d01cb15d94b89988d8fddf9af - Miru Tracer:
0.3.2 - Transformers:
5.13.0 - PyTorch:
2.12.1+cu126 - Compute dtype:
bfloat16 - Fit settings:
dim_batch=8,max_seq_len=128,skip_first=16, target layer 39, checkpoint chunk size 5 - Successful prompts: 459 of a 1,000-prompt maximum; 0 skipped
- Convergence: default early-stopping criterion reached after 459 prompts; final rolling 10-prompt
d_mean=0.0019539635108923002(threshold0.002) - Adapter tensors: 39
float16matrices, each2048 x 2048 - Adapter size: 327,244,360 bytes
- Adapter SHA-256:
2c49664cf02764af9d0ce94d6ca68bc5196289e3954c166353d648df2148f0eb
MoE routing audit
Before fitting, the frozen 1,000-prompt calibration sequence was audited at the same token positions used by the Jacobian estimator. Across 111,000 valid token positions and 888,000 routed expert selections per layer:
- every layer reached at least 250 of 256 routed experts (97.65625% minimum coverage);
- 20 of 40 layers reached all 256 experts, and 35 reached at least 99%;
- 44 of 10,240 layer-expert cells were unobserved;
- the minimum normalized routing entropy was 0.803093; and
- the minimum effective expert count was 85.9.
The audit supports using this WikiText sequence for a distribution-conditioned fit. It does not establish equal adapter quality for code, chat, mathematics, or multilingual traffic.
The production fit was a fresh run with the model balanced across two H100 GPUs. The final adapter loads successfully, contains finite matrices, and retains the fitter's embedded model, tokenizer, prompt-sequence, fit, and convergence provenance.
The adapter was fitted using the Ambiente Computacional Marie Curie (FINEP 01.22.181.00) at UFSCar.
Contributor Public-Domain Certification
I certify that I created this contribution or otherwise have the authority
to submit it. To the extent that I own copyright or related rights in the
adapter and its accompanying metadata, I permanently dedicate those rights to
the public domain under the Unlicense. Where a public-domain dedication is
not legally recognized, I make the contribution available under all
permissions and disclaimers stated by the Unlicense. I have disclosed the
base model and calibration sources, and I am not knowingly submitting
material that I lack permission to distribute. If I am contributing as part
of my employment or for another organization, I certify that I am authorized
to make this dedication on its behalf.
Signed-off-by: Lucas Teske (@racerxdl ), 2026-09-08
Automated by MisterMal