GitHub: Build-scripts and docs  |  Forensics study  |  Runtime fork  |  Discussions

Scion-35B-A3B: ternary MoE experts + trained corrections

A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus small trained corrections, with no full-precision masters and no full-model QAT.

2.61 bpw | 11.34 GB (6.3x smaller than the BF16 reference) | best PPL of the 2-bit class | task retention at Q4-class level, at about half of Q4_K_M's size

The name comes from grafting. A scion is the shoot grafted onto a rootstock, and here the trained corrections are grafted onto a 2.125 bpw ternary body.


Quickstart

The one thing to know first: this file uses a container that only exists in the runtime fork. Stock/upstream llama.cpp (and the runtimes built on it, like LM Studio and Ollama) will refuse it with invalid ggml type 142. The Hub's auto-generated "Use this model" panel points at upstream binaries, so ignore it: you need sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime. A pre-built copy of that fork is below, so most people do not need to compile anything.

Step 1 - pick a runtime

Your machine Use this Notes
Windows x64 + NVIDIA llama-v0.4.1-bin-win-cuda-12.4-x64.zip + cudart-llama-bin-win-cuda-12.4-x64.zip Preferred for most drivers. CUDA 13.3 build also available if your driver is very new.
Windows x64, no GPU windows-x64-cpu.zip Slower, works everywhere. MTP still works.
Linux x64 + NVIDIA linux-x64-cuda.tar.gz
Linux x64, no GPU linux-x64-cpu.tar.gz
Linux x64 + AMD linux-x64-hip.tar.gz The verified path is a source build with -DGGML_HIP=ON; see below.
macOS Apple Silicon macos-arm64-metal.tar.gz Metal path is present but has not been benchmarked by the author.
Anything else / you want max control Build from source One command block below.

All release links point at the v0.4.1 release.

Do not use Vulkan. It is not a supported path for this model. The PQ2_0 kernels exist only for CPU, CUDA and Metal; a Vulkan build would leave the expert matmuls on the CPU, which defeats the point. See Limitations.

Step 2 - get the weights

# model body (11.3 GB)
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .

# optional k=1 speculative drafter (50 MB, ~1.2x faster)
hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .

hf download is the current CLI; the older huggingface-cli download works the same.

Step 3 - run

# chat (the model thinks by default, so leave room for the trace)
./llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -ngl 99 -c 4096 -t 8 \
    --temp 0.6 --top-p 0.95 --top-k 20 \
    -p "Explain quantum computing in simple terms." -n 16384

# OpenAI-compatible server with a web UI
./llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t 8 --port 8080

# same, with the drafter auto-detected from the sidecar GGUF
./llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080
  • -ngl 99 offloads every layer (0 is CPU-only).
  • -c is context up to 262144; start at 4096.
  • -t should equal the physical core count (SMT siblings hurt CPU expert throughput).
  • If VRAM is tight, keep -ngl 99 and offload experts with -ncmoe.
  • On Windows, use llama-cli.exe / llama-server.exe, PowerShell or cmd.

Windows CUDA users: extract the main zip, then extract the three DLLs from the matching cudart zip (cudart64_12.dll, cublas64_12.dll, cublasLt64_12.dll) into the same folder as llama-server.exe / llama-cli.exe. Without them the binary will fail to start with a missing-DLL error.

Build from source (any platform)

git clone -b moe-corr-runtime https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh          # must print "RESULT: OK"
rm -rf build

# pick ONE backend flag:
cmake -B build -DGGML_CUDA=ON          # NVIDIA
# cmake -B build -DGGML_HIP=ON         # AMD (verified on gfx1100 / RX 7900 XT)
# cmake -B build                       # CPU only
cmake --build build -j --target llama-cli llama-server
# binaries land in build/bin/

On Windows, build with CMake + Ninja against MSVC (or the LLVM toolchain); a plain cmake -B build -DGGML_CUDA=ON after vcvarsall.bat x64 is enough.

Troubleshooting

Symptom Cause Fix
tensor '...' has invalid ggml type 142. should be in [0, 43) You ran an upstream llama.cpp build, an old binary earlier on PATH, or a stale build/ dir Use the fork binary from above; delete build/ and re-clone if compiling. Run ./verify-container-support.sh.
Missing cudart64_*.dll / cublas64_*.dll CUDA runtime not bundled in the main zip Extract the matching cudart-* zip next to the executables.
GPU not used / slow Wrong build, or experts on CPU Confirm ggml-cuda.dll loaded and -ngl 99; add -ncmoe only if VRAM is tight.
Out of VRAM at long context KV cache growth Lower -c, or offload experts with -ncmoe.
Hangs at start on old CPUs CPU backend picked a variant the CPU lacks Use the fork bundle (ships all CPU variants) or rebuild with -DGGML_NATIVE=OFF.

For AI assistants / agents helping a user

If you are an AI agent wiring Scion up for someone, the requirements are:

Runtime MUST be sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime.
  - Upstream llama.cpp, LM Studio, Ollama and llama.cpp wheels CANNOT load this file.
Model file: Scion-35B-A3B-PQ2_0-corr.gguf
  - Custom tensor type is PQ2_0 (ggml type id 142), group size 128. NOT upstream Q2_0 (type 42).
GPU backends that implement PQ2_0: CUDA, CPU, Metal. Vulkan does NOT.
Windows x64 + NVIDIA, no build required:
  1. Download llama-v0.4.1-bin-win-cuda-12.4-x64.zip and cudart-llama-bin-win-cuda-12.4-x64.zip from
     https://github.com/sky-is-green/prism-ml-llama.cpp/releases/tag/v0.4.1
  2. Extract both into ONE folder (the cudart DLLs must sit next to llama-server.exe / llama-cli.exe).
  3. Run:  llama-server.exe -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t <physical cores> --port 8080
Recommended sampling: --temp 0.6 --top-p 0.95 --top-k 20, and allow -n up to 16384 (reasoning trace).
Optional speedup: -md Scion-35B-A3B-mtp-drafter.gguf (drafter auto-detected, same context).
"invalid ggml type 142" == wrong (upstream) binary; use the fork.

Highlights

  • One 11.34 GB file at 2.61 bpw (expert banks at 2.125 bpw). The BF16 reference of the same model is 71.07 GB; Q4_K_M is 21.71 GB and IQ2_M is 12.56 GB. It is the smallest published build of this model I have seen at this quality level.
  • Task retention inside the Q4 and BF16 noise band: HellaSwag 400 gives 79.00% (BF16 81.25, Q4_K_M 80.00) and Winogrande gives 76.25% (BF16 76.00, Q4_K_M 76.00). It has the joint-best Winogrande row in the table and the best PPL of the 2-bit class (8.354 against IQ2_M 8.413 and Q2_K 8.473).
  • Trained, not calibrated: rank-512 correction branches on the attention output and the MoE block output, plus router deltas, trained by output-KD against the BF16 teacher with the deployed quantizer in the loop (ternary Lloyd g128). No imatrix and no calibration corpus, which is what separates this build from the imatrix-calibrated quants on the chart.
  • One file, no --lora: the corrections are embedded (adapter.embedded=true) and attached at load. There is no adapter plumbing.
  • A k=1 speculative drafter ships alongside (Scion-35B-A3B-mtp-drafter, 50 MB): it drafts the next token from the model's own hidden state and gives 1.14-1.37x faster generation on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).
  • The gap is stated, not hidden: full-vocabulary KLD against BF16 is 0.269 mean, a strong 2-bit-class result but still behind Q4_K_M at 0.031. Tail-aware training was attempted at full scale and did not transfer (TAIL-EXPERIMENT-PLAN.md).

Resources

  • GitHub sky-is-green/scion: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.
  • Bonsai 2 ternary forensics: the dense-model study this method grew out of (format recovery, the trained-weight residual, the calibration-artifact result).
  • Runtime: sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime, a fork of Prism ML's llama.cpp with the PQ2_0 container, the ffn_moe_out virtual target and embedded-adapter support.
  • Retention grid: this release against every community quant of the same model, same protocol.
  • Discussions: questions, test reports and failures are all welcome.

Model Overview

Item Specification
Base model empero-ai/Qwen3.8-35B-A3B-Distill, a Qwen3.8-line reasoning distill built on Qwen/Qwen3.6-35B-A3B (Apache-2.0)
Parameters 34.9B total, about 3B active per token (256 experts, top-8 plus shared)
Architecture qwen3_5_moe (qwen35moe in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward
Context length 262,144 tokens (inherited from the base model)
Weight format Ternary PQ2_0 g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), Q8_0 for the rest, embedded corrections in the legacy q1_0_g128 container (rank-512)
Low-bit coverage Expert banks only; attention, embeddings and the head stay Q8_0, and norms, routers and the output stay F32
Deployed size 10.558 GiB / 11.337 GB (text only, single file)
Backends llama.cpp fork; verified on CPU and ROCm/gfx1100. Pre-built Windows x64 CUDA binaries are on the v0.4.1 release but the CUDA path is not yet runtime-verified end-to-end. Metal untested. Vulkan unsupported (no PQ2_0 kernels)
License Apache-2.0 (inherited from the base)

Weight Representation: ternary PQ2_0 + trained corrections

Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the model is Q8_0 (attention, embeddings, LM head) and F32 (norms, routers, output). The corrections are rank-512 low-rank branches on the attention output and the MoE block output, plus exact router deltas. They ship in the compact legacy q1_0_g128 container (2-bit codes plus one fp16 group scale per 128) and are merged into the file. Effective overall: 2.61 bpw.

Memory Requirement

Format bpw Size vs BF16
BF16 (reference) 16.38 71.07 GB 1.0x
Q8_0 8.72 37.80 GB 1.9x
Q4_K_M 5.01 21.71 GB 3.3x
IQ2_M 2.90 12.56 GB 5.7x
Scion-35B-A3B 2.61 11.34 GB 6.3x

Sizes are the published on-disk files of the same model, measured under one protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see Benchmarks).

Shipped Components

Component Pack Size Residency
Language model (this repo) PQ2_0 experts, Q8_0 rest, embedded corrections 11.34 GB resident; the whole model
k=1 MTP drafter (separate repo) fp16 head (fc1/gelu/fc2) over the model's own hidden + next-token embedding 50 MB transient; ~1 ms/eval on GPU
Uncorrected body (not uploaded) PQ2_0 experts and Q8_0 rest 10.46 GiB for swap tests
Corrections (not uploaded) rank-512 branches and router deltas (q1_0_g128) 98 MiB for swap tests

The language model is the single released body file; the drafter is a separate 50 MB file in its own repo (Scion-35B-A3B-mtp-drafter). The two-file variant (body plus separate adapter) exists for reproducing the merge and swapping corrections at runtime. Ask in Discussions if you want it.

About the Hub's quant chip. The Hub parses file names and labels this file Q2_0; the same happens on Prism ML's own PQ2_0 releases. The container is Prism's PQ2_0 (legacy name Q1_0_g128, type 142/43, identical byte layout): 2-bit codes with one fp16 group scale per 128 weights. It is not upstream llama.cpp's g64 Q2_0 (type 42). No single quant name fits the file anyway, because it is a mix: PQ2_0 expert banks, the embedded corrections in the legacy q1_0_g128 container, and Q8_0 for the rest (norms and routers in F32).

Best Practices

Generation Parameters

Recommended values, from the base model card:

  • temperature=0.6, top_p=0.95, top_k=20

This is a reasoning distill, and answers open with a long thinking segment. Allow generous max_new_tokens (for example -n 16384); a small cap ends generation mid-thought, before any answer.

System Prompt

A simple prompt works, for example You are a helpful assistant. The base is a reasoning SFT distill and does not require a special system prompt.

Choosing Context and Offload

  • The weights fit a 12 GB card; about 16 GB is comfortable with context.
  • -ngl 99 on a single card. When VRAM is tight, offload experts to CPU with -ncmoe, which is the VRAM-budget dial; the ternary container roughly halves the CPU-tail penalty compared to an f16 expert bank.
  • Threads should equal physical cores (-t 8 on an 8C/16T CPU; SMT siblings collapse CPU expert throughput).
  • Do not layer-split across two cards when one card fits; the proxy measurements showed a 40% generation loss.

Speculative decoding (k=1 drafter)

The drafter lives in its own repo: Scion-35B-A3B-mtp-drafter (50 MB). It is a small frozen-body head that predicts the model's own next token from its post-norm hidden state plus the next token's embedding; it reuses the model's own output_norm/output/token_embd, so it adds no second vocabulary projection. It is not a stock llama.cpp drafter: the fork's draft-mtp-sidecar implementation loads it (auto-detected from the GGUF) and runs it entirely against the target context, with no draft model.

hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .

# server (no extra flags; the sidecar is auto-detected)
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080

# CLI
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8

Measured on one RX 7900 XT (greedy):

Workload Baseline + drafter Speedup Draft acceptance
Bench prompt, 400 tokens 12.41 ms/tok 9.07 ms/tok 1.37x 0.738
Open-ended CLI generation 49.0 tok/s 57.2 tok/s 1.17x 0.46-0.74
Server confirmation - - 1.14-1.21x 0.44-0.70

Teacher-forced acceptance: fineweb 0.483 / wikitext 0.360. Speedup is prompt-dependent - treat ~1.2x as the typical figure and 1.37x as the benchmark-prompt best. Drafting leaves the model's outputs unchanged apart from rare near-tie flips on batched verification.

Benchmarks

Community protocol, the same for every row: wikitext-2 PPL (c512, 580 chunks); KLD against BF16 logits over 50 chunks (25.5k tokens); HellaSwag 400 and Winogrande 400 zero-shot (about +/-2% CI). BF16 and all quants were measured in one H100 session; this release was measured on the local card and cross-checked against the pod (IQ2_M KLD local vs pod: 0.2%).

Task Retention (400 tasks each)

Variant Size bpw HellaSwag Winogrande
BF16 (reference) 71.07 GB 16.38 81.25 76.00
Q4_K_M 21.71 GB 5.01 80.00 76.00
IQ2_M 12.56 GB 2.90 79.00 75.75
Q2_K 13.84 GB 3.19 76.50 73.25
Scion-35B-A3B 11.34 GB 2.61 79.00 76.25

Within +/-2% noise of Q4_K_M and BF16; tied with IQ2_M on HellaSwag and ahead of it on Winogrande at 1.2 GB less. Treat sub-1% differences as ties.

Additional tasks (re-measured 2026-09-30)

The 2026-09-27 grid covered HellaSwag and Winogrande. A later local run added the remaining standard multiple-choice tasks on the same release and the same fork build (one RX 7900 XT, llama-perplexity --multiple-choice, same seed):

Task Scion-35B-A3B
ARC-Challenge (299) 57.19 +/-2.87
ARC-Easy (570) 80.00 +/-1.68
MMLU (2000, seed 1234) 40.15 +/-1.10
TruthfulQA (817) 32.56 +/-1.64
wikitext-2 PPL (c512, 580 chunks) 8.2731 +/-0.054

The fresh PPL uses --no-warmup and -np 8, so its level sits ~1% below the grid's 8.3539; the model is unchanged.

Distributional Fidelity: the known gap

Variant PPL (lower better) KLD mean vs BF16 (lower better) KLD 99.9%
BF16 7.1595 - -
Q4_K_M 7.2354 0.0314 1.141
IQ2_M 8.4133 0.1636 3.963
Q2_K 8.4729 0.1493 3.276
Scion-35B-A3B 8.3539 0.2694 4.758

The corrections improved mean token likelihood (PPL 11.60 uncorrected, 8.35 after) more than they improved the full-distribution tail. PPL ranks this build first of the 2-bit class and KLD ranks it last, and that disagreement is the research result: the training matches the teacher's top-50 logits, and the rest of the distribution is unconstrained. Tail-aware training was attempted at full scale (a tail-conditional KD term plus router bias plus a hard-window curriculum, the cur05 recipe) and did not transfer: on a 40-layer body the teacher puts ~98% of its mass inside the top-512 cache, so the tail terms are nearly inert, and the retrain ties this release on every task while PPL regresses ~2%. The gap stands, and the negative is recorded in the project's register.

Full Grid

Scion-35B-A3B against every community quant of the same model, plus BF16

All ten community quants plus BF16; up is better, left is smaller. Full table and method notes: QUANT-RETENTION-35B.md.

Full per-variant table
model size GB bpw PPL KLD mean KLD 99.9% HellaSwag Winogrande
Scion-35B-A3B 11.34 2.61 8.354 0.269 4.758 79.00 76.25
IQ2_M 12.56 2.90 8.413 0.164 3.963 79.00 75.75
Q2_K 13.84 3.19 8.473 0.149 3.276 76.50 73.25
IQ3_M 16.34 3.77 7.520 0.057 1.225 80.00 74.25
Q3_K_M 17.66 4.07 7.432 0.058 1.611 79.75 76.00
IQ4_XS 19.63 4.53 7.264 0.022 0.636 80.75 75.75
Q4_K_M 21.71 5.01 7.235 0.031 1.141 80.00 76.00
Q5_K_M 25.35 5.84 7.273 0.015 0.717 80.50 76.25
Q6_K 29.21 6.73 7.153 0.008 0.360 80.75 74.75
Q8_0 37.80 8.72 7.160 0.004 0.209 80.25 75.50
BF16 71.07 16.38 7.160 - - 81.25 76.00

Use Cases

  • A 35B-A3B on one consumer GPU: 11.3 GB of weights fit a 20 GB card with room for context, and can be tiered further with CPU expert offload.
  • Low-bit research and testing: a reference point for "ternary experts plus trained corrections, no full-precision masters, no full-model QAT". The build-scripts and every negative result are public.
  • Local-first serving: the model was built for an offline assistant stack, so a single file with embedded corrections keeps deployment to one download and one set of flags.

Limitations

  • KLD tail (stated above): distributional fidelity is strong-2-bit, not Q4-class. PPL, HellaSwag and Winogrande look Q4-class; KLD is where the gap lives. A full-scale tail-aware retrain was attempted and did not transfer.
  • Text only: the base model has a vision tower, and this file carries no vision tensors (Q8_0 language path only).
  • Reasoning distill: long thinking traces; budget max_new_tokens accordingly.
  • Protocol caveats: the KLD figure is 50 chunks; task numbers are 400-task runs (about +/-2% CI); the 2-bit and 3-bit competitors are imatrix-calibrated on Wikipedia-like data, which flatters their wikitext KLD, and the corrections were trained on fineweb.
  • Platform coverage: CPU and ROCm (gfx1100) verified. Windows x64 CUDA binaries are now published on the v0.4.1 release, but the CUDA path has not been runtime-verified end-to-end yet; reports welcome. Metal is present but untested. Vulkan is unsupported: ggml-vulkan has no PQ2_0 kernels, so the expert matmuls would fall back to CPU.
  • The Hub chip says Q2_0: it is filename-derived; see the note above, and use the fork.
  • Not affiliated with Prism ML, empero-ai, or Alibaba Cloud. It builds on Prism ML's engine work (fork and PQ2_0 container) and community GGUF conversions.

Citation

@misc{scion35b2026,
    title  = {Scion-35B-A3B: ternary MoE experts with trained corrections},
    author = {SkyIsNotGreen},
    year   = {2026},
    month  = {September},
    url    = {https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B}
}

License and Attribution

Apache-2.0, inherited from the base model; see LICENSE and NOTICE. Base weights: empero-ai and Qwen / Alibaba Cloud (Apache-2.0); BF16 conversion: MrFuzzihead; container and kernels: Prism ML (MIT) with TAARDIS conventions (MIT); engine: llama.cpp (MIT).

Downloads last month
2,133
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SkyIsNotGreen/Scion-35B-A3B

Quantized
(22)
this model
Quantizations
1 model