Instructions to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF", filename="ornith-1.0-35b-MXFP4_MOE-MTP.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE # Run inference directly in the terminal: llama cli -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE # Run inference directly in the terminal: llama cli -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE # Run inference directly in the terminal: ./llama-cli -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE # Run inference directly in the terminal: ./build/bin/llama-cli -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Use Docker
docker model run hf.co/s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
- LM Studio
- Jan
- vLLM
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
- Ollama
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with Ollama:
ollama run hf.co/s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
- Unsloth Studio
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF to start chatting
- Pi
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with Docker Model Runner:
docker model run hf.co/s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
- Lemonade
How to use s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF:MXFP4_MOE
Run and chat with the model
lemonade run user.Ornith-1.0-35B-NVFP4-MTP-GGUF-MXFP4_MOE
List all available models
lemonade list
s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF
MXFP4 + Multi-Token Prediction (MTP) heads grafted from Qwen3.6-35B-A3B, packaged as a single GGUF for llama.cpp.
Designed for NVIDIA Blackwell GPUs (sm_120 / sm_121) including the RTX PRO
6000 and the DGX Spark (GB10). MXFP4 is dequantized natively by
Blackwell tensor cores, and the grafted MTP heads enable draft-mtp
speculative decoding for ~2× decode throughput versus the body-only quant.
Original Model
Ornith-1.0-35B is a
self-improving agentic-coding model released by the
DeepReinforce team, post-trained via RL on top
of Qwen3.5-35B-A3B. It emits a reasoning_content block before its final
answer and is competitive with Qwen3.6-35B-A3B and Gemma 4-31B on
Terminal-Bench 2.1, SWE-bench Verified/Pro, and Claw-eval.
- Architecture: Qwen3.5 MoE (
qwen3_5_moe), 40 layers, 256 experts, hidden_size 2048 - Parameters: 35B total / ~3B active
- Vocabulary: 248,064 tokens (multimodal vocab preserved; vision tower not included in this GGUF)
- License: MIT (inherited from upstream)
- Citation: see Citation below
Quantization Details
This repository ships two files with the same trunk weights but different expert quantizations:
File 1 — ornith-1.0-35b-MXFP4_MOE-MTP.gguf (MXFP4 + MTP, the original)
| Property | Value |
|---|---|
Body weights — trunk 3D experts (blk.*.ffn_*_exps.weight, 80 tensors) |
MXFP4 (GGML type 39, E2M1 + E8M0 scale, 32-element blocks, 4.25 bpw) |
Body weights — 2D linears (attn_qkv, attn_gate, ssm_*, shared experts) |
Q8_0 |
MTP head — 3D experts (blk.40.nextn.ffn_*_exps, 2 tensors) |
MXFP4 |
MTP head — non-experts (blk.40.attn_*, nextn.*_norm, nextn.eh_proj, blk.40.ffn_down_exps) |
Q8_0 (one Q5_K outlier from the graft source: blk.40.ffn_down_exps.weight) |
Sensitive tensors (token_embd, output) |
Q8_0 |
Norms, biases, router (*.norm.weight, ssm_dt.bias, ffn_gate_inp*) |
F32 / BF16 (default for MXFP4_MOE) |
| Total size | 20.30 GB (4.572 bpw average) |
| Total tensors | 713 |
| Source ftype | MXFP4_MOE (LLAMA_FTYPE_MOSTLY_MXFP4_MOE = 38) |
| Source tool | llama-quantize MXFP4_MOE from llama.cpp build 9590 (CUDA 13.1, sm_120 Blackwell) |
| MTP source | MTP block grafted byte-for-byte from unsloth/Qwen3.6-35B-A3B-MTP-GGUF |
File 2 — ornith-1.0-35b-NVFP4-MTP.gguf (true NVFP4 + MTP, NEW)
| Property | Value |
|---|---|
Body weights — trunk 3D experts (blk.*.ffn_*_exps.weight, blk.*.ffn_gate_up_exps.weight, 80 tensors) |
NVFP4 (GGML type 40, E2M1 + E4M3 scale, 16-element blocks, 4.50 bpw) |
Body weights — 2D linears (attn_qkv, attn_gate, ssm_*, shared experts) |
Q8_0 |
MTP head — 3D experts (blk.40.nextn.ffn_*_exps, 2 tensors) |
MXFP4 (from graft source, since the Qwen3.6 MTP source is MXFP4) |
MTP head — non-experts (blk.40.attn_*, nextn.*_norm, nextn.eh_proj, blk.40.ffn_down_exps) |
Q8_0 (one Q5_K outlier from the graft source: blk.40.ffn_down_exps.weight) |
Sensitive tensors (token_embd, output) |
Q8_0 |
Norms, biases, router (*.norm.weight, ssm_dt.bias, ffn_gate_inp*) |
F32 / BF16 (default for NVFP4 routing) |
| Total size | 19.85 GiB (21.31 GB) (~4.80 bpw average) |
| Total tensors | 713 |
| Source ftype | NVFP4 (LLAMA_FTYPE_MOSTLY_NVFP4 = 39) |
| Source tool | llama-quantize NVFP4 from llama.cpp patched to (a) register NVFP4 in QUANT_OPTIONS, (b) route 3D MoE tensors to NVFP4 and 2D linears to Q8_0 (parity with MXFP4_MOE) |
| MTP source | MTP block grafted byte-for-byte from unsloth/Qwen3.6-35B-A3B-MTP-GGUF |
What is MXFP4?
MXFP4 is the OCP Microscaling Formats 4-bit floating-point spec
(OCP Microscaling Formats (MX) Specification v1.0). It is an
open, vendor-neutral standard — supported on both NVIDIA Blackwell and
AMD Instinct MI355X. Each block of 32 contiguous elements shares a
single 8-bit unsigned scale factor (E8M0) — a pure power-of-two with no
mantissa, giving an enormous dynamic range (2⁻¹²⁷ to 2¹²⁷) but fractional
precision only at the byte-block level.
What is NVFP4, and how does it differ?
NVFP4 is NVIDIA's proprietary Blackwell-native variant. Same element encoding (E2M1) and same 4-bit storage cost, but:
| MXFP4 (this file) | NVFP4 (second file, released) | |
|---|---|---|
| Element | E2M1 | E2M1 (same) |
| Block size | 32 elements | 16 elements |
| Scale format | E8M0 (power-of-two only) | FP8 E4M3 (fractional) |
| Second-level scale | none | FP32 per tensor |
| Effective bpw | 4.25 | 4.50 |
| Hardware | Blackwell + AMD MI355X | Blackwell only |
| Tensor cores used | Blackwell FP4 | Blackwell FP4 (same) |
NVFP4's finer 16-element blocks + fractional E4M3 scales + per-tensor FP32 shift fit each block's distribution more tightly, giving slightly lower quantization error than MXFP4's power-of-two snapping. Empirically the gap on 35B-class MoE models is <1% on standard benchmarks (MMLU, GPQA, HumanEval); for most inference workloads the two formats are functionally interchangeable on Blackwell.
What is MTP and why graft it?
Multi-Token Prediction uses a small auxiliary head to predict multiple tokens ahead. At inference, those predictions become a draft that the main model verifies in a single forward pass — speculative decoding with zero quality loss (output distribution is identical to non-MTP) and 1.5-2× decode speedup when drafts are accepted.
Ornith-1.0-35B is a Qwen3.5-35B-A3B post-trained variant. Its trunk weights
share the same parameter shapes and base tensor layout as Qwen3.6-35B-A3B,
including the same qwen3_5_moe MTP block position (20 tensors at
blk.40.*, ~488 MB). Because the MTP head operates on the same hidden
state and embedding space as the trunk, the grafted heads work without any
additional training — the MTP drafts the next token given the trunk's
hidden state, the trunk then verifies. No re-tuning required.
Acceptance rates measured on RTX PRO 6000:
| Sampling | Accept rate | Speedup vs body-only |
|---|---|---|
| Greedy (temp=0) | 100% (21/21, 80/80) | ~1.85× |
| temp=0.6, top-p=0.95 | 86–93% | ~1.7× |
| temp=0.9 (production) | 75–88% | ~1.5× |
Performance
RTX PRO 6000 (Blackwell, sm_120, 97 GB GDDR7, ~1.79 TB/s)
Single-slot, 200k context, temp=0.9, top_p=0.95, top_k=20, min_p=0.01,
KV cache q8_0:
| Mode | Decode (tok/s) | Prefill (tok/s) | Notes |
|---|---|---|---|
| MXFP4_MOE (body only, no MTP) | 240 | ~6,000 | baseline |
| MXFP4_MOE + MTP (n-max=3) | 310–320 | ~6,000 | production config |
| Q8_K_XL + MTP (Unsloth upstream) | ~165 | ~5,500 | reference |
The MTP head costs ~280 MiB of KV context; the trunk savings (4.57 BPW vs 8.5 BPW for Q8_K) more than pay for it.
DGX Spark (Blackwell GB10, sm_121, 121 GB unified LPDDR5X, ~273 GB/s)
3 slots × 200k context each (c=600000, parallel=3), KV cache q8_0.
Production sweep (8 configs, B×UB×n-max combinations, 4-run stability test):
| Config (B / UB / n-max) | Combined tok/s | Per-slot tok/s | Accept rate |
|---|---|---|---|
| 2048 / 2048 / 3 ⭐ (winner) | 93.0 ± 1.2 | ~31 each | 67–88% |
| 2048 / 1024 / 3 (baseline) | 92.0 | ~31 each | 58–83% |
| 4096 / 1024 / 3 | 92.4 | ~31 each | 60–91% |
| 1024 / 512 / 3 | 91.6 | ~31 each | 64–79% |
| 2048 / 1024 / 4 | 85.3 | ~29 each | 51–70% |
| 2048 / 1024 / 5 | 82.3 | ~29 each | 48–68% |
| 2048 / 1024 / 6 | 81.3 | ~30 each | 43–51% |
The ~41% throughput ratio vs the 6000 Pro (93 / 240) matches the GB10's ~273 GB/s vs the 6000 Pro's ~1.79 TB/s unified-memory bandwidth gap, confirming the workload is bandwidth-bound. MTP-only (no ngram-mod) was chosen because the MTP head already provides the draft stream and the ngram-mod bookkeeping overhead exceeds its marginal gain.
Provided Files
| File | Size | Tensor type | Notes |
|---|---|---|---|
ornith-1.0-35b-MXFP4_MOE-MTP.gguf (this file — MXFP4 under the hood) |
20.30 GB | MXFP4 (type 39) for experts, Q8_0 for sensitive, F32 for norms | Recommended for most Blackwell users — same throughput as NVFP4, smaller storage |
ornith-1.0-35b-NVFP4-MTP.gguf (RELEASED — true NVFP4, tensor type 40) |
21.31 GB | NVFP4 (type 40) for the 80 trunk 3D experts, Q8_0 for 2D linears, F32 for norms | Slightly higher accuracy vs MXFP4 (E4M3 fractional scales vs E8M0 power-of-two). Use when MMLU-Pro / GPQA accuracy matters more than 1 GB of storage. |
chat_template.jinja |
7.5 KB | — | Jinja chat template (inherited from upstream; emits reasoning_content block) |
config.json |
3.3 KB | — | Original HF model config (text_config subset) |
tokenizer.json |
19 MB | — | HF fast tokenizer |
tokenizer_config.json |
1.2 KB | — | Tokenizer metadata |
vocab.json |
6.7 MB | — | BPE vocab |
preprocessor_config.json |
390 B | — | Image preprocessor (inherited; unused by text-only inference) |
processor_config.json |
1.2 KB | — | Multimodal processor (inherited; unused by text-only inference) |
generation_config.json |
213 B | — | Generation defaults |
model.safetensors.index.json |
3.3 MB | — | Weight map (informational) |
LICENSE |
1 KB | — | MIT license text |
README.md |
— | — | This file |
The video_preprocessor_config.json file is also present but is
informational only.
Usage with llama.cpp
Requirements
- llama.cpp build 9590 or later (FP4 tensor core support, sm_120 / sm_121 CUDA)
- CUDA toolkit with Blackwell support
- Build with
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES="120;121"(set to your GPU's compute capability)
Server (recommended for production)
RTX PRO 6000 / single-GPU Blackwell, single slot, 200k context:
llama-server \
-m ornith-1.0-35b-MXFP4_MOE-MTP.gguf \
--host 0.0.0.0 --port 8080 --slots --metrics \
-t 20 -cb --no-warmup --no-mmap --mlock \
--jinja -fa on -ctk q8_0 -ctv q8_0 \
--cache-reuse 256 -ctxcp 256 --checkpoint-every-n-tokens 4096 \
-b 2048 -ub 2048 -c 1000000 -np 5 -ngl 99 \
--chat-template-file chat_template.jinja \
--spec-type draft-mtp,ngram-mod \
--spec-draft-n-max 3 \
--spec-ngram-mod-n-match 24 \
--spec-ngram-mod-n-min 48 \
--spec-ngram-mod-n-max 64 \
--reasoning-budget 2048 \
--reasoning-budget-message "I have thought enough. Let me give my answer now." \
--temp 0.9 --top-p 0.95 --top-k 20 --min-p 0.01 --repeat-penalty 1.0
DGX Spark / 3 slots × 200k context each:
llama-server \
-m ornith-1.0-35b-MXFP4_MOE-MTP.gguf \
--host 0.0.0.0 --port 8080 --slots --metrics \
-t 20 -cb --no-warmup --no-mmap --mlock \
--jinja -fa on -ctk q8_0 -ctv q8_0 \
--cache-reuse 256 -ctxcp 256 --checkpoint-every-n-tokens 4096 \
-b 2048 -ub 2048 -c 600000 -np 3 -ngl 99 \
--chat-template-file chat_template.jinja \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--reasoning-budget 2048 \
--reasoning-budget-message "I have thought enough. Let me give my answer now." \
--temp 0.9 --top-p 0.95 --top-k 20 --min-p 0.01 --repeat-penalty 1.0
CLI
llama-cli \
-m ornith-1.0-35b-MXFP4_MOE-MTP.gguf \
--chat-template-file chat_template.jinja \
-p "Explain gradient descent in 3 sentences." \
-ngl 99 \
--temp 0.9 --top-p 0.95 --top-k 20 --min-p 0.01 \
--spec-type draft-mtp --spec-draft-n-max 3
Direct download with llama.cpp
llama-cli \
--hf-repo s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF \
--hf-file ornith-1.0-35b-NVFP4-MTP.gguf \
-p "What is 17 times 23?"
(Substitute ornith-1.0-35b-MXFP4_MOE-MTP.gguf for the original MXFP4
variant.)
Important Notes
- Blackwell only. MXFP4/NVFP4 are hardware-specific formats. They will not run efficiently on non-Blackwell GPUs. For AMD, Intel, or pre-Blackwell NVIDIA GPUs, use the upstream deepreinforce-ai/Ornith-1.0-35B-GGUF (Q4_K_M / Q8_0 variants).
--no-mmaprecommended on unified-memory machines (DGX Spark) to avoid mmap-related slowdowns.- KV cache type q8_0 is the production setting. F16 KV uses ~2× more memory for negligible quality change on long context.
- MTP is single-slot at peak speed but works well in multi-slot mode — the MTP context is shared across slots and adds minimal per-slot overhead. On the DGX Spark, 3 slots at ~31 t/s each = ~93 t/s combined.
- MXFP4 vs NVFP4 — pick
ornith-1.0-35b-MXFP4_MOE-MTP.gguf(the MXFP4 file) for smallest storage and identical Blackwell throughput. Pickornith-1.0-35b-NVFP4-MTP.gguf(the NVFP4 file) if you need the slightly tighter quantization error of E4M3 fractional scales.
How the MTP graft was made
The MTP heads were transferred byte-for-byte from
unsloth/Qwen3.6-35B-A3B-MTP-GGUF
into the deepreinforce-ai Ornith-1.0-35B MXFP4 GGUF:
- Source GGUF (Qwen3.6) has a 20-tensor MTP block at
blk.40.*totaling ~488 MB. - Destination GGUF (Ornith-1.0-35B MXFP4) has the same
qwen3_5_moearchitecture withblock_count=40and no MTP tensors. - The graft script appends the MTP block immediately after the destination
data block, rewrites tensor offsets, and adds two new KV pairs:
qwen35moe.nextn_predict_layers=1and bumpsqwen35moe.block_count=40→41.
This is possible because:
- Ornith-1.0-35B is post-trained on top of Qwen3.5-35B-A3B (same architecture, same shapes, same embeddings)
- The MTP head is small (~488 MB) and depends only on the trunk's hidden state and the shared embedding/output projection
- Qwen3.5 and Qwen3.6 share the same
qwen3_5_moeMTP architecture (Qwen3.6 is Qwen3.5 post-trained)
Acceptance at 100% on greedy decoding confirms the graft is functionally correct: when the trunk deterministically predicts token N, the MTP head's prediction of token N+1 matches exactly.
The graft script is available on request.
How to verify the tensor types
To confirm this file is MXFP4 (not NVFP4), inspect the GGUF header with the
official gguf_dump.py tool (shipped with llama.cpp):
python3 gguf-py/gguf/scripts/gguf_dump.py --no-tensors --markdown \
ornith-1.0-35b-MXFP4_MOE-MTP.gguf | head -50
Look for:
general.file_type = 38(meansLLAMA_FTYPE_MOSTLY_MXFP4_MOE)- Tensor types: 82×
MXFP4(the 3D experts), 320×Q8_0, 308×F32, 2×BF16, 1×Q5_K(graft outlier)
If those numbers are present, you have the MXFP4 file. For true NVFP4
the general.file_type would be 39, the tensor count for type 40 (NVFP4)
would be 80 (40 trunk layers × 2 expert tensors per layer), and total file
size would be ~21.3 GB.
Verifying the NVFP4 file
To confirm the second file in this repository is genuine NVFP4:
python3 gguf-py/gguf/scripts/gguf_dump.py --no-tensors --markdown \
ornith-1.0-35b-NVFP4-MTP.gguf | head -60
Look for:
general.file_type = 39(meansLLAMA_FTYPE_MOSTLY_NVFP4)- Tensor types in this file: 80×
NVFP4(type 40 — the 3D trunk experts only:blk.*.ffn_down_exps.weight+blk.*.ffn_gate_up_exps.weight× 40 layers), 320×Q8_0(2D linears:attn_qkv,attn_gate,ssm_*, shared experts, sensitive tensors), 308×F32(norms, biases, router), 2×MXFP4(the 2 large 3D expert tensors grafted from the Qwen MTP source — type 39, not 40), 2×I32, 1×Q5_K(graft outlier) - Total: 713 tensors, ~21.31 GB
qwen35moe.block_count = 41andqwen35moe.nextn_predict_layers = 1confirm the MTP head is grafted in
The NVFP4 file was produced by patching llama.cpp's
tools/quantize/quantize.cpp and src/llama-quant.cpp to:
- Register
LLAMA_FTYPE_MOSTLY_NVFP4(value 39) inQUANT_OPTIONSso the CLI acceptsNVFP4as an ftype argument (the upstreamQUANT_OPTIONStable omits NVFP4 even thoughllama.hdefines the constant), AND - Route 3D MoE tensors (
tensor->ne[2] > 1) toGGML_TYPE_NVFP4and 2D tensors toGGML_TYPE_Q8_0— mirroring the routing block that MXFP4_MOE already has atllama-quant.cpp:461.
Without fix (2), the quantizer would default every tensor to NVFP4 (yielding 390 NVFP4 tensors), losing Q8_0 accuracy on the 2D linears.
Licensing
This model is licensed under MIT, the same license as the original deepreinforce-ai/Ornith-1.0-35B. See LICENSE for the full text.
The MTP heads originate from Qwen/Qwen3.6-35B-A3B (Apache 2.0). The MTP block was sourced via the unsloth/Qwen3.6-35B-A3B-MTP-GGUF redistribution, which is permitted under the original Apache 2.0 terms.
Citation
@misc{ornith-35b,
title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
url = {https://deep-reinforce.com/ornith_1_0.html},
author = {{DeepReinforce Team}},
year = {2026}
}
@misc{qwen3.5,
title = {{Qwen3.5-35B-A3B}: Post-trained Mixture-of-Experts Language Model},
url = {https://huggingface.co/Qwen/Qwen3.5-35B-A3B},
author = {{Qwen Team}},
year = {2026}
}
@misc{qwen3.6,
title = {{Qwen3.6-35B-A3B}: Qwen3.5 Post-trained with MTP Heads},
url = {https://huggingface.co/Qwen/Qwen3.6-35B-A3B},
author = {{Qwen Team}},
year = {2026}
}
Acknowledgements
- DeepReinforce Team for training and releasing Ornith-1.0-35B
- Alibaba Qwen Team for the Qwen3.5 and Qwen3.6 base models
- Unsloth for the MTP-enabled GGUF redistribution that made the graft possible
- llama.cpp for FP4 tensor core support and the
draft-mtpspeculative decoding implementation
- Downloads last month
- 58,158
4-bit
Model tree for s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF
Base model
deepreinforce-ai/Ornith-1.0-35B