Instructions to use SkyIsNotGreen/Scion-35B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SkyIsNotGreen/Scion-35B-A3B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0 # Run inference directly in the terminal: llama cli -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0 # Run inference directly in the terminal: llama cli -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Use Docker
docker model run hf.co/SkyIsNotGreen/Scion-35B-A3B:Q2_0
- LM Studio
- Jan
- vLLM
How to use SkyIsNotGreen/Scion-35B-A3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SkyIsNotGreen/Scion-35B-A3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SkyIsNotGreen/Scion-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SkyIsNotGreen/Scion-35B-A3B:Q2_0
- Ollama
How to use SkyIsNotGreen/Scion-35B-A3B with Ollama:
ollama run hf.co/SkyIsNotGreen/Scion-35B-A3B:Q2_0
- Unsloth Desktop
- Pi
How to use SkyIsNotGreen/Scion-35B-A3B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SkyIsNotGreen/Scion-35B-A3B:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use SkyIsNotGreen/Scion-35B-A3B with Docker Model Runner:
docker model run hf.co/SkyIsNotGreen/Scion-35B-A3B:Q2_0
- Lemonade
How to use SkyIsNotGreen/Scion-35B-A3B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SkyIsNotGreen/Scion-35B-A3B:Q2_0
Run and chat with the model
lemonade run user.Scion-35B-A3B-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use SkyIsNotGreen/Scion-35B-A3B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SkyIsNotGreen/Scion-35B-A3B:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use SkyIsNotGreen/Scion-35B-A3B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SkyIsNotGreen/Scion-35B-A3B:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SkyIsNotGreen/Scion-35B-A3B:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GitHub: Build-scripts and docs | Forensics study | Runtime fork | Discussions
Scion-35B-A3B: ternary MoE experts + trained corrections
A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus small trained corrections, with no full-precision masters and no full-model QAT.
2.61 bpw | 11.34 GB (6.3x smaller than the BF16 reference) | best PPL of the 2-bit class | task retention at Q4-class level, at about half of Q4_K_M's size
The name comes from grafting. A scion is the shoot grafted onto a rootstock, and here the trained corrections are grafted onto a 2.125 bpw ternary body.
Quickstart
The one thing to know first: this file uses a container that only exists in
the runtime fork. Stock/upstream llama.cpp (and the runtimes built on it, like
LM Studio and Ollama) will refuse it with invalid ggml type 142. The Hub's
auto-generated "Use this model" panel points at upstream binaries, so ignore it:
you need sky-is-green/prism-ml-llama.cpp,
branch moe-corr-runtime. A pre-built copy of that fork is below, so most people
do not need to compile anything.
Step 1 - pick a runtime
| Your machine | Use this | Notes |
|---|---|---|
| Windows x64 + NVIDIA | llama-v0.4.1-bin-win-cuda-12.4-x64.zip + cudart-llama-bin-win-cuda-12.4-x64.zip |
Preferred for most drivers. CUDA 13.3 build also available if your driver is very new. |
| Windows x64, no GPU | windows-x64-cpu.zip |
Slower, works everywhere. MTP still works. |
| Linux x64 + NVIDIA | linux-x64-cuda.tar.gz |
|
| Linux x64, no GPU | linux-x64-cpu.tar.gz |
|
| Linux x64 + AMD | linux-x64-hip.tar.gz |
The verified path is a source build with -DGGML_HIP=ON; see below. |
| macOS Apple Silicon | macos-arm64-metal.tar.gz |
Metal path is present but has not been benchmarked by the author. |
| Anything else / you want max control | Build from source | One command block below. |
All release links point at the v0.4.1 release.
Do not use Vulkan. It is not a supported path for this model. The
PQ2_0kernels exist only for CPU, CUDA and Metal; a Vulkan build would leave the expert matmuls on the CPU, which defeats the point. See Limitations.
Step 2 - get the weights
# model body (11.3 GB)
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
# optional k=1 speculative drafter (50 MB, ~1.2x faster)
hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .
hf download is the current CLI; the older huggingface-cli download works the
same.
Step 3 - run
# chat (the model thinks by default, so leave room for the trace)
./llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-ngl 99 -c 4096 -t 8 \
--temp 0.6 --top-p 0.95 --top-k 20 \
-p "Explain quantum computing in simple terms." -n 16384
# OpenAI-compatible server with a web UI
./llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t 8 --port 8080
# same, with the drafter auto-detected from the sidecar GGUF
./llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080
-ngl 99offloads every layer (0is CPU-only).-cis context up to 262144; start at 4096.-tshould equal the physical core count (SMT siblings hurt CPU expert throughput).- If VRAM is tight, keep
-ngl 99and offload experts with-ncmoe. - On Windows, use
llama-cli.exe/llama-server.exe, PowerShell or cmd.
Windows CUDA users: extract the main zip, then extract the three DLLs from
the matching cudart zip (cudart64_12.dll, cublas64_12.dll,
cublasLt64_12.dll) into the same folder as llama-server.exe / llama-cli.exe.
Without them the binary will fail to start with a missing-DLL error.
Build from source (any platform)
git clone -b moe-corr-runtime https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh # must print "RESULT: OK"
rm -rf build
# pick ONE backend flag:
cmake -B build -DGGML_CUDA=ON # NVIDIA
# cmake -B build -DGGML_HIP=ON # AMD (verified on gfx1100 / RX 7900 XT)
# cmake -B build # CPU only
cmake --build build -j --target llama-cli llama-server
# binaries land in build/bin/
On Windows, build with CMake + Ninja against MSVC (or the LLVM toolchain); a
plain cmake -B build -DGGML_CUDA=ON after vcvarsall.bat x64 is enough.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
tensor '...' has invalid ggml type 142. should be in [0, 43) |
You ran an upstream llama.cpp build, an old binary earlier on PATH, or a stale build/ dir |
Use the fork binary from above; delete build/ and re-clone if compiling. Run ./verify-container-support.sh. |
Missing cudart64_*.dll / cublas64_*.dll |
CUDA runtime not bundled in the main zip | Extract the matching cudart-* zip next to the executables. |
| GPU not used / slow | Wrong build, or experts on CPU | Confirm ggml-cuda.dll loaded and -ngl 99; add -ncmoe only if VRAM is tight. |
| Out of VRAM at long context | KV cache growth | Lower -c, or offload experts with -ncmoe. |
| Hangs at start on old CPUs | CPU backend picked a variant the CPU lacks | Use the fork bundle (ships all CPU variants) or rebuild with -DGGML_NATIVE=OFF. |
For AI assistants / agents helping a user
If you are an AI agent wiring Scion up for someone, the requirements are:
Runtime MUST be sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime.
- Upstream llama.cpp, LM Studio, Ollama and llama.cpp wheels CANNOT load this file.
Model file: Scion-35B-A3B-PQ2_0-corr.gguf
- Custom tensor type is PQ2_0 (ggml type id 142), group size 128. NOT upstream Q2_0 (type 42).
GPU backends that implement PQ2_0: CUDA, CPU, Metal. Vulkan does NOT.
Windows x64 + NVIDIA, no build required:
1. Download llama-v0.4.1-bin-win-cuda-12.4-x64.zip and cudart-llama-bin-win-cuda-12.4-x64.zip from
https://github.com/sky-is-green/prism-ml-llama.cpp/releases/tag/v0.4.1
2. Extract both into ONE folder (the cudart DLLs must sit next to llama-server.exe / llama-cli.exe).
3. Run: llama-server.exe -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t <physical cores> --port 8080
Recommended sampling: --temp 0.6 --top-p 0.95 --top-k 20, and allow -n up to 16384 (reasoning trace).
Optional speedup: -md Scion-35B-A3B-mtp-drafter.gguf (drafter auto-detected, same context).
"invalid ggml type 142" == wrong (upstream) binary; use the fork.
Highlights
- One 11.34 GB file at 2.61 bpw (expert banks at 2.125 bpw). The BF16 reference of the same model is 71.07 GB; Q4_K_M is 21.71 GB and IQ2_M is 12.56 GB. It is the smallest published build of this model I have seen at this quality level.
- Task retention inside the Q4 and BF16 noise band: HellaSwag 400 gives 79.00% (BF16 81.25, Q4_K_M 80.00) and Winogrande gives 76.25% (BF16 76.00, Q4_K_M 76.00). It has the joint-best Winogrande row in the table and the best PPL of the 2-bit class (8.354 against IQ2_M 8.413 and Q2_K 8.473).
- Trained, not calibrated: rank-512 correction branches on the attention output and the MoE block output, plus router deltas, trained by output-KD against the BF16 teacher with the deployed quantizer in the loop (ternary Lloyd g128). No imatrix and no calibration corpus, which is what separates this build from the imatrix-calibrated quants on the chart.
- One file, no
--lora: the corrections are embedded (adapter.embedded=true) and attached at load. There is no adapter plumbing. - A k=1 speculative drafter ships alongside (
Scion-35B-A3B-mtp-drafter, 50 MB): it drafts the next token from the model's own hidden state and gives 1.14-1.37x faster generation on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter). - The gap is stated, not hidden: full-vocabulary KLD against BF16 is 0.269 mean, a strong 2-bit-class result but still behind Q4_K_M at 0.031. Tail-aware training was attempted at full scale and did not transfer (
TAIL-EXPERIMENT-PLAN.md).
Resources
- GitHub
sky-is-green/scion: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register. - Bonsai 2 ternary forensics: the dense-model study this method grew out of (format recovery, the trained-weight residual, the calibration-artifact result).
- Runtime:
sky-is-green/prism-ml-llama.cpp, branchmoe-corr-runtime, a fork of Prism ML's llama.cpp with the PQ2_0 container, theffn_moe_outvirtual target and embedded-adapter support. - Retention grid: this release against every community quant of the same model, same protocol.
- Discussions: questions, test reports and failures are all welcome.
Model Overview
| Item | Specification |
|---|---|
| Base model | empero-ai/Qwen3.8-35B-A3B-Distill, a Qwen3.8-line reasoning distill built on Qwen/Qwen3.6-35B-A3B (Apache-2.0) |
| Parameters | 34.9B total, about 3B active per token (256 experts, top-8 plus shared) |
| Architecture | qwen3_5_moe (qwen35moe in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward |
| Context length | 262,144 tokens (inherited from the base model) |
| Weight format | Ternary PQ2_0 g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), Q8_0 for the rest, embedded corrections in the legacy q1_0_g128 container (rank-512) |
| Low-bit coverage | Expert banks only; attention, embeddings and the head stay Q8_0, and norms, routers and the output stay F32 |
| Deployed size | 10.558 GiB / 11.337 GB (text only, single file) |
| Backends | llama.cpp fork; verified on CPU and ROCm/gfx1100. Pre-built Windows x64 CUDA binaries are on the v0.4.1 release but the CUDA path is not yet runtime-verified end-to-end. Metal untested. Vulkan unsupported (no PQ2_0 kernels) |
| License | Apache-2.0 (inherited from the base) |
Weight Representation: ternary PQ2_0 + trained corrections
Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per
group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the
model is Q8_0 (attention, embeddings, LM head) and F32 (norms, routers, output).
The corrections are rank-512 low-rank branches on the attention output and the
MoE block output, plus exact router deltas. They ship in the compact legacy
q1_0_g128 container (2-bit codes plus one fp16 group scale per 128) and are
merged into the file. Effective overall: 2.61 bpw.
Memory Requirement
| Format | bpw | Size | vs BF16 |
|---|---|---|---|
| BF16 (reference) | 16.38 | 71.07 GB | 1.0x |
| Q8_0 | 8.72 | 37.80 GB | 1.9x |
| Q4_K_M | 5.01 | 21.71 GB | 3.3x |
| IQ2_M | 2.90 | 12.56 GB | 5.7x |
| Scion-35B-A3B | 2.61 | 11.34 GB | 6.3x |
Sizes are the published on-disk files of the same model, measured under one protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see Benchmarks).
Shipped Components
| Component | Pack | Size | Residency |
|---|---|---|---|
| Language model (this repo) | PQ2_0 experts, Q8_0 rest, embedded corrections | 11.34 GB | resident; the whole model |
| k=1 MTP drafter (separate repo) | fp16 head (fc1/gelu/fc2) over the model's own hidden + next-token embedding |
50 MB | transient; ~1 ms/eval on GPU |
| Uncorrected body (not uploaded) | PQ2_0 experts and Q8_0 rest | 10.46 GiB | for swap tests |
| Corrections (not uploaded) | rank-512 branches and router deltas (q1_0_g128) |
98 MiB | for swap tests |
The language model is the single released body file; the drafter is a separate
50 MB file in its own repo
(Scion-35B-A3B-mtp-drafter).
The two-file variant (body plus separate adapter) exists for reproducing the
merge and swapping corrections at runtime. Ask in Discussions if you want it.
About the Hub's quant chip. The Hub parses file names and labels this file
Q2_0; the same happens on Prism ML's ownPQ2_0releases. The container is Prism'sPQ2_0(legacy nameQ1_0_g128, type 142/43, identical byte layout): 2-bit codes with one fp16 group scale per 128 weights. It is not upstream llama.cpp's g64Q2_0(type 42). No single quant name fits the file anyway, because it is a mix:PQ2_0expert banks, the embedded corrections in the legacyq1_0_g128container, andQ8_0for the rest (norms and routers in F32).
Best Practices
Generation Parameters
Recommended values, from the base model card:
temperature=0.6,top_p=0.95,top_k=20
This is a reasoning distill, and answers open with a long thinking segment.
Allow generous max_new_tokens (for example -n 16384); a small cap ends
generation mid-thought, before any answer.
System Prompt
A simple prompt works, for example You are a helpful assistant. The base is a
reasoning SFT distill and does not require a special system prompt.
Choosing Context and Offload
- The weights fit a 12 GB card; about 16 GB is comfortable with context.
-ngl 99on a single card. When VRAM is tight, offload experts to CPU with-ncmoe, which is the VRAM-budget dial; the ternary container roughly halves the CPU-tail penalty compared to an f16 expert bank.- Threads should equal physical cores (
-t 8on an 8C/16T CPU; SMT siblings collapse CPU expert throughput). - Do not layer-split across two cards when one card fits; the proxy measurements showed a 40% generation loss.
Speculative decoding (k=1 drafter)
The drafter lives in its own repo:
Scion-35B-A3B-mtp-drafter
(50 MB). It is a small frozen-body head that predicts the model's own next
token from its post-norm hidden state plus the next token's embedding; it reuses
the model's own output_norm/output/token_embd, so it adds no second
vocabulary projection. It is not a stock llama.cpp drafter: the fork's
draft-mtp-sidecar implementation loads it (auto-detected from the GGUF) and
runs it entirely against the target context, with no draft model.
hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .
# server (no extra flags; the sidecar is auto-detected)
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080
# CLI
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8
Measured on one RX 7900 XT (greedy):
| Workload | Baseline | + drafter | Speedup | Draft acceptance |
|---|---|---|---|---|
| Bench prompt, 400 tokens | 12.41 ms/tok | 9.07 ms/tok | 1.37x | 0.738 |
| Open-ended CLI generation | 49.0 tok/s | 57.2 tok/s | 1.17x | 0.46-0.74 |
| Server confirmation | - | - | 1.14-1.21x | 0.44-0.70 |
Teacher-forced acceptance: fineweb 0.483 / wikitext 0.360. Speedup is prompt-dependent - treat ~1.2x as the typical figure and 1.37x as the benchmark-prompt best. Drafting leaves the model's outputs unchanged apart from rare near-tie flips on batched verification.
Benchmarks
Community protocol, the same for every row: wikitext-2 PPL (c512, 580 chunks);
KLD against BF16 logits over 50 chunks (25.5k tokens); HellaSwag 400 and
Winogrande 400 zero-shot (about +/-2% CI). BF16 and all quants were measured in one
H100 session; this release was measured on the local card and cross-checked
against the pod (IQ2_M KLD local vs pod: 0.2%).
Task Retention (400 tasks each)
| Variant | Size | bpw | HellaSwag | Winogrande |
|---|---|---|---|---|
| BF16 (reference) | 71.07 GB | 16.38 | 81.25 | 76.00 |
| Q4_K_M | 21.71 GB | 5.01 | 80.00 | 76.00 |
| IQ2_M | 12.56 GB | 2.90 | 79.00 | 75.75 |
| Q2_K | 13.84 GB | 3.19 | 76.50 | 73.25 |
| Scion-35B-A3B | 11.34 GB | 2.61 | 79.00 | 76.25 |
Within +/-2% noise of Q4_K_M and BF16; tied with IQ2_M on HellaSwag and ahead of it on Winogrande at 1.2 GB less. Treat sub-1% differences as ties.
Additional tasks (re-measured 2026-09-30)
The 2026-09-27 grid covered HellaSwag and Winogrande. A later local run added
the remaining standard multiple-choice tasks on the same release and the same
fork build (one RX 7900 XT, llama-perplexity --multiple-choice, same seed):
| Task | Scion-35B-A3B |
|---|---|
| ARC-Challenge (299) | 57.19 +/-2.87 |
| ARC-Easy (570) | 80.00 +/-1.68 |
| MMLU (2000, seed 1234) | 40.15 +/-1.10 |
| TruthfulQA (817) | 32.56 +/-1.64 |
| wikitext-2 PPL (c512, 580 chunks) | 8.2731 +/-0.054 |
The fresh PPL uses --no-warmup and -np 8, so its level sits ~1% below the
grid's 8.3539; the model is unchanged.
Distributional Fidelity: the known gap
| Variant | PPL (lower better) | KLD mean vs BF16 (lower better) | KLD 99.9% |
|---|---|---|---|
| BF16 | 7.1595 | - | - |
| Q4_K_M | 7.2354 | 0.0314 | 1.141 |
| IQ2_M | 8.4133 | 0.1636 | 3.963 |
| Q2_K | 8.4729 | 0.1493 | 3.276 |
| Scion-35B-A3B | 8.3539 | 0.2694 | 4.758 |
The corrections improved mean token likelihood (PPL 11.60 uncorrected, 8.35
after) more than they improved the full-distribution tail. PPL ranks this build
first of the 2-bit class and KLD ranks it last, and that disagreement is the
research result: the training matches the teacher's top-50 logits, and the rest
of the distribution is unconstrained. Tail-aware training was attempted at
full scale (a tail-conditional KD term plus router bias plus a hard-window
curriculum, the cur05 recipe) and did not transfer: on a 40-layer body the
teacher puts ~98% of its mass inside the top-512 cache, so the tail terms are
nearly inert, and the retrain ties this release on every task while PPL regresses
~2%. The gap stands, and the negative is recorded in the project's register.
Full Grid
All ten community quants plus BF16; up is better, left is smaller. Full table
and method notes:
QUANT-RETENTION-35B.md.
Full per-variant table
| model | size GB | bpw | PPL | KLD mean | KLD 99.9% | HellaSwag | Winogrande |
|---|---|---|---|---|---|---|---|
| Scion-35B-A3B | 11.34 | 2.61 | 8.354 | 0.269 | 4.758 | 79.00 | 76.25 |
| IQ2_M | 12.56 | 2.90 | 8.413 | 0.164 | 3.963 | 79.00 | 75.75 |
| Q2_K | 13.84 | 3.19 | 8.473 | 0.149 | 3.276 | 76.50 | 73.25 |
| IQ3_M | 16.34 | 3.77 | 7.520 | 0.057 | 1.225 | 80.00 | 74.25 |
| Q3_K_M | 17.66 | 4.07 | 7.432 | 0.058 | 1.611 | 79.75 | 76.00 |
| IQ4_XS | 19.63 | 4.53 | 7.264 | 0.022 | 0.636 | 80.75 | 75.75 |
| Q4_K_M | 21.71 | 5.01 | 7.235 | 0.031 | 1.141 | 80.00 | 76.00 |
| Q5_K_M | 25.35 | 5.84 | 7.273 | 0.015 | 0.717 | 80.50 | 76.25 |
| Q6_K | 29.21 | 6.73 | 7.153 | 0.008 | 0.360 | 80.75 | 74.75 |
| Q8_0 | 37.80 | 8.72 | 7.160 | 0.004 | 0.209 | 80.25 | 75.50 |
| BF16 | 71.07 | 16.38 | 7.160 | - | - | 81.25 | 76.00 |
Use Cases
- A 35B-A3B on one consumer GPU: 11.3 GB of weights fit a 20 GB card with room for context, and can be tiered further with CPU expert offload.
- Low-bit research and testing: a reference point for "ternary experts plus trained corrections, no full-precision masters, no full-model QAT". The build-scripts and every negative result are public.
- Local-first serving: the model was built for an offline assistant stack, so a single file with embedded corrections keeps deployment to one download and one set of flags.
Limitations
- KLD tail (stated above): distributional fidelity is strong-2-bit, not Q4-class. PPL, HellaSwag and Winogrande look Q4-class; KLD is where the gap lives. A full-scale tail-aware retrain was attempted and did not transfer.
- Text only: the base model has a vision tower, and this file carries no vision tensors (Q8_0 language path only).
- Reasoning distill: long thinking traces; budget
max_new_tokensaccordingly. - Protocol caveats: the KLD figure is 50 chunks; task numbers are 400-task runs (about +/-2% CI); the 2-bit and 3-bit competitors are imatrix-calibrated on Wikipedia-like data, which flatters their wikitext KLD, and the corrections were trained on fineweb.
- Platform coverage: CPU and ROCm (gfx1100) verified. Windows x64 CUDA
binaries are now published on the
v0.4.1release, but the CUDA path has not been runtime-verified end-to-end yet; reports welcome. Metal is present but untested. Vulkan is unsupported:ggml-vulkanhas no PQ2_0 kernels, so the expert matmuls would fall back to CPU. - The Hub chip says
Q2_0: it is filename-derived; see the note above, and use the fork. - Not affiliated with Prism ML, empero-ai, or Alibaba Cloud. It builds on Prism ML's engine work (fork and PQ2_0 container) and community GGUF conversions.
Citation
@misc{scion35b2026,
title = {Scion-35B-A3B: ternary MoE experts with trained corrections},
author = {SkyIsNotGreen},
year = {2026},
month = {September},
url = {https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B}
}
License and Attribution
Apache-2.0, inherited from the base model; see LICENSE and
NOTICE.
Base weights: empero-ai and
Qwen / Alibaba Cloud (Apache-2.0);
BF16 conversion: MrFuzzihead;
container and kernels: Prism ML (MIT) with
TAARDIS conventions (MIT);
engine: llama.cpp (MIT).
- Downloads last month
- 2,133
2-bit
