Add Qwen3.6-35B-A3B Jacobian-lens adapter

#6
by racerxdl - opened

Adds a Miru Tracer Jacobian-lens adapter for Qwen/Qwen3.6-35B-A3B.

Contributor: Lucas Teske at Teske's Lab.

Fit details

  • Base model: Qwen/Qwen3.6-35B-A3B (Apache-2.0)
  • Model revision: 995ad96eacd98c81ed38be0c5b274b04031597b0
  • Model architecture SHA-256: 18c2b5747dfced7887ec8af29a8cdada3637cf4c88f33f99a78c65c2cb207bf5 (miru-semantic-v1)
  • Model config SHA-256: 8842f9c349baaea67fe949237518f5161c41c680849c188cfdc32e01f79118e8 (transformers-json-v1)
  • Tokenizer SHA-256: 522cccf6b8e6164ea42f0c04a13599c64d784e6e307b4272164eed8c4607ad9e
  • Calibration corpus: frozen prompt sequence selected from Salesforce/wikitext, wikitext-103-raw-v1, train split, dataset revision b08601e04326c79dfdd32d625aee71d232d685c3
  • Prompt-sequence SHA-256: dcb2a738953033139ae35f3665271ccdb795409c6073a85ea46f02a852143cc6
  • Processed-prefix SHA-256: 5a85854c042b31dcce791cdddc5fc0d891508d0d01cb15d94b89988d8fddf9af
  • Miru Tracer: 0.3.2
  • Transformers: 5.13.0
  • PyTorch: 2.12.1+cu126
  • Compute dtype: bfloat16
  • Fit settings: dim_batch=8, max_seq_len=128, skip_first=16, target layer 39, checkpoint chunk size 5
  • Successful prompts: 459 of a 1,000-prompt maximum; 0 skipped
  • Convergence: default early-stopping criterion reached after 459 prompts; final rolling 10-prompt d_mean=0.0019539635108923002 (threshold 0.002)
  • Adapter tensors: 39 float16 matrices, each 2048 x 2048
  • Adapter size: 327,244,360 bytes
  • Adapter SHA-256: 2c49664cf02764af9d0ce94d6ca68bc5196289e3954c166353d648df2148f0eb

MoE routing audit

Before fitting, the frozen 1,000-prompt calibration sequence was audited at the same token positions used by the Jacobian estimator. Across 111,000 valid token positions and 888,000 routed expert selections per layer:

  • every layer reached at least 250 of 256 routed experts (97.65625% minimum coverage);
  • 20 of 40 layers reached all 256 experts, and 35 reached at least 99%;
  • 44 of 10,240 layer-expert cells were unobserved;
  • the minimum normalized routing entropy was 0.803093; and
  • the minimum effective expert count was 85.9.

The audit supports using this WikiText sequence for a distribution-conditioned fit. It does not establish equal adapter quality for code, chat, mathematics, or multilingual traffic.

The production fit was a fresh run with the model balanced across two H100 GPUs. The final adapter loads successfully, contains finite matrices, and retains the fitter's embedded model, tokenizer, prompt-sequence, fit, and convergence provenance.

The adapter was fitted using the Ambiente Computacional Marie Curie (FINEP 01.22.181.00) at UFSCar.

Contributor Public-Domain Certification

I certify that I created this contribution or otherwise have the authority
to submit it. To the extent that I own copyright or related rights in the
adapter and its accompanying metadata, I permanently dedicate those rights to
the public domain under the Unlicense. Where a public-domain dedication is
not legally recognized, I make the contribution available under all
permissions and disclaimers stated by the Unlicense. I have disclosed the
base model and calibration sources, and I am not knowingly submitting
material that I lack permission to distribute. If I am contributing as part
of my employment or for another organization, I certify that I am authorized
to make this dedication on its behalf.

Signed-off-by: Lucas Teske (@racerxdl ), 2026-09-08

Automated by MisterMal

rlaneth changed pull request status to merged

Sign up or log in to comment