FROM ./Janus-35B-A3B.Q4_K_M.gguf # The bundled GGUF carries no MTP / NextN block. The base # (llmfan46/Qwen3.6-35B-A3B-uncensored-heretic) publishes its GGUFs already # MTP-clean — blk.0…blk.39 only, block_count 40, no nextn_predict_layers key # (read from the bundled file's own header, 2026-09-18) — so scripts/strip_mtp.py # is a defensive no-op here, not a required build step. The separate # llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF variant # does keep the MTP head (its Q4_K_M has block_count 41, NextN block at index # 40) and would need the strip; this repo does not ship it. If you want the MTP # head, run the upstream safetensors under vLLM/SGLang, or take that variant's # GGUF. # Chat template — Qwen 3.6 ChatML in Ollama Go-template form, with the # tool-calling blocks Ollama's capability detector looks for. Without a # TEMPLATE that references .Tools and .ToolCalls, an Ollama that uses this # template anyway (OLLAMA_GO_TEMPLATE=1 - per Ollama's source, not tested live - # or one that predates the template comparison described below) rejects any # request carrying a `tools` array with ` does not support tools`; # Ollama 0.33.3 would instead switch to the GGUF's embedded template. Same template as the dense # 27B sibling (FoolDev/Thanatos-27B-HERETIC) — the ChatML wire format is # identical across Qwen 3.6 and 3.8. # # Thinking IS replayed across turns: every assistant message that carries # reasoning renders ..., earlier turns included, where Qwen's # stock condition kept only the turn in progress (a tool-call chain included). # The client must send the reasoning back - per Ollama's source, /api/chat reads # each assistant message's `thinking` field and /v1/chat/completions reads only # `reasoning` (it drops `reasoning_content` and `thinking` silently). Every # retained trace stays in the prompt and prefill grows to match, which is the # slow part on a CPU-only host. The prompt-token counts that used to sit here # were measured on the dense 27B this repo shipped in the interim; they are # removed rather than carried over, pending re-measurement. # # The thinking block in the assistant branch is load-bearing for a second # reason. Ollama infers a Go template's "thinking" capability from a `.Thinking` field inside # `range .Messages` wrapped in /, and Ollama 0.33.3 and 0.34.0 (verified; # the selection code is unchanged through 0.34.2, checked in source) # compare that capability list with the GGUF's embedded Jinja template at load # time, switching to the embedded template if that one advertises more. Delete # the block and Ollama switches: in one test on the dense 27B this repo shipped # in the interim a single tool call then came back twice # (a separate raw generation showed the model drafting the call inside # first; that the parser matched such a draft is inferred, not confirmed). Per Ollama's source (not tested live), a # setup forcing this template (OLLAMA_GO_TEMPLATE=1) would also reject thinking # requests with HTTP 400. To stop replaying earlier turns' reasoning, change the # condition back to Qwen's stock # `(and $.IsThinkSet (and .Thinking (or $last (gt $i $lastUserIdx))))` - the # $lastUserIdx loop at the top of the template is kept for it; do not delete the # block. # # Two more properties of this template are read from its PARSE TREE, not its # output, so no render test can see them break. Ollama takes the tool-call tag # from the first `{{ if }}` whose condition names .ToolCalls and then the first # literal text in that block, so (a) the `range .ToolCalls` body must start with # text, never an action - that was 0.9.5 - and (b) no earlier condition may name # .ToolCalls, which is why the assistant branch aliases it as # `{{ $calls := .ToolCalls }}` and branches on the variable. The thinking tags # come from the first and last nodes of the list holding {{ .Thinking }}, both of # which must be literal text, so nothing conditional may end that block - the # blank-line condition sits after it. scripts/check_go_template.py emulates both # of Ollama's readers and asserts what they derive. # # Reasoning effort: the block at the top of the template injects an effort # instruction keyed on Ollama's think level ($.ThinkLevel), not the request # value. It is Ollama-only now — this repo no longer ships chat_template.jinja, # and the Qwen 3.6 base's own embedded template, which governs the llama.cpp # path from here on, has no reasoning_effort handling at all. # Ollama folds `reasoning_effort` into four levels first: # high/xhigh/max/ultra -> high or max (xhigh line), low/minimal -> low (low # line), medium and unset -> medium (no line), none -> thinking off. The # outer `if $.ThinkLevel` keeps the template rendering on Ollama builds that # predate the field. TEMPLATE """{{- $lastUserIdx := -1 -}} {{- range $idx, $msg := .Messages -}} {{- if eq $msg.Role "user" }}{{ $lastUserIdx = $idx }}{{ end -}} {{- end }} {{- $effort := "" }} {{- if $.ThinkLevel }} {{- if or (eq $.ThinkLevel "high") (eq $.ThinkLevel "max") }}{{ $effort = "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." }} {{- else if eq $.ThinkLevel "low" }}{{ $effort = "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." }} {{- end }} {{- end }} {{- if or .System .Tools $effort }}<|im_start|>system {{ if $effort }}{{ $effort }}{{ if or .System .Tools }} {{ end }}{{ end }}{{ if .System }}{{ .System }}{{ if .Tools }} {{ end }}{{ end }} {{- if .Tools }}# Tools You may call one or more functions to assist with the user query. You are provided with function signatures within XML tags: {{- range .Tools }} {"type": "function", "function": {{ json .Function }}} {{- end }} For each function call, return a json object with function name and arguments within XML tags: {"name": , "arguments": } {{- end -}}<|im_end|> {{ end }} {{- $prevIsTool := false }} {{- range $i, $_ := .Messages }} {{- $rest := slice $.Messages $i }} {{- $last := eq (len $rest) 1 -}} {{- $nextIsTool := and (gt (len $rest) 1) (eq (index $rest 1).Role "tool") -}} {{- if eq .Role "user" }}<|im_start|>user {{ .Content }}<|im_end|> {{ else if eq .Role "assistant" }}{{ $calls := .ToolCalls }}{{ $blank := or .Content (not $calls) }}<|im_start|>assistant {{ if $.IsThinkSet -}} {{ .Thinking }} {{ end -}} {{ if and $.IsThinkSet $blank }} {{ end -}} {{ if .Content }}{{ .Content }}{{ if $calls }} {{ end }}{{ end }} {{- if .ToolCalls }} {{- range .ToolCalls }} {"name": "{{ .Function.Name }}", "arguments": {{ .Function.Arguments }}} {{- end }} {{- end }}{{ if not $last }}<|im_end|> {{ end }} {{- else if eq .Role "tool" }} {{- if $prevIsTool }} {{ else }}<|im_start|>user {{ end }} {{ .Content }} {{- if not $nextIsTool }}<|im_end|> {{ end }} {{- end }} {{- $prevIsTool = eq .Role "tool" }} {{- if and (ne .Role "assistant") $last }}<|im_start|>assistant {{ if and $.IsThinkSet (not $.Think) -}} {{ else -}} {{ end -}} {{ end }} {{- end }}""" # Sampling tuned for reasoning + general use. See README "Recommended sampling" # for creative/RP alternatives. PARAMETER temperature 1.0 PARAMETER top_p 0.95 PARAMETER top_k 0 PARAMETER repeat_penalty 1.05 PARAMETER num_ctx 262144 # Stop tokens. Without these, Ollama only honors <|im_end|> from the GGUF # metadata; the model occasionally emits <|endoftext|> instead and Ollama # keeps generating past it (synthesising a fake new user turn). Listing # both — plus <|im_start|> as a belt-and-braces guard against the same # loop — keeps responses cleanly terminated. Same fix the sibling # (FoolDev/Thanatos-27B-HERETIC) shipped in commit 6672746. PARAMETER stop "<|im_end|>" PARAMETER stop "<|endoftext|>" PARAMETER stop "<|im_start|>" SYSTEM """You are Janus, a precise and capable assistant for reasoning, writing, coding, and long-form dialogue. Behavior rules: - Answer the user's actual request directly. - Be accurate, complete, and structured. - Think before answering, but do not get stuck in repetitive loops or meta-commentary. - If the request is ambiguous or incomplete, state what is missing and make the smallest reasonable assumption needed to continue. - If the user wants creative writing, preserve tone, continuity, and character consistency. - If the user wants analysis or technical help, prefer concrete steps, examples, and decisions over fluff. - Finish with a usable answer, not just planning.""" # Hardware notes # -------------- # This Q4_K_M is 21,233,608,512 bytes on disk — 21.23 GB decimal, 19.78 GiB. # Measured 2026-09-18 on Ollama 0.33.3's CPU backend, from llama.cpp's own # allocation lines, in an isolated store at num_ctx 8192 / 32768 / 65536: # weights 19.77 GiB (CPU model buffer 5361.34 MiB + CPU_REPACK # 14878.12 MiB); constant # KV cache only 10 of the 40 layers are full-attention (indices # 3, 7, ... 39), and llama.cpp logs the cache as "10 # layers": 20,480 B/token at f16 = 160 / 640 / 1280 MiB # at 8K / 32K / 64K, i.e. 0.625 GiB per 32K, and # 10,880 B/token at q8_0 = 340 / 680 MiB at 32K / 64K, # i.e. 0.332 GiB per 32K. Exactly linear in num_ctx -> # 5.0 GiB at the 262144 default with f16, 2.66 with q8_0 # recurrent state 62.81 MiB for the 30 Gated-DeltaNet layers (R f32 2.81 # + S f32 60.00) — identical at every context # compute buffer 160.04 / 208.04 / 544.07 MiB at 8K / 32K / 64K; grows # faster than linearly, so it is not extrapolated # totals 20.7 GiB at 32768 and 21.6 GiB at 65536 (both sums of # measured parts); ~25.4 GiB at the 262144 default with # the f16 cache and ~23.0 GiB with # OLLAMA_KV_CACHE_TYPE=q8_0 — those two carry the 64K # compute buffer forward, so treat them as floors. # No YaRN rope-scaling is baked in this GGUF, so output # past ~262K degrades; treat 1.01M as an advertised # ceiling. # # This is a 35B-A3B MoE across 40 layers: 256 experts, 8 routed per token plus # one shared expert, ~34.7B parameters total and ~3B active per token. The active # count cuts compute per token, not the memory floor — every expert has to be # resident, so budget for the whole 19.77 GiB of weights. # # Working configurations (rows assume the 262144 default: ~25.4 GiB with the f16 # cache, ~23.0 GiB with q8_0): # ✓ Single H100 80GB / A100 80GB — full GPU offload # ✓ RTX 5090 32GB / RTX 4090 24GB + 32GB RAM — partial offload # ✓ Mac Studio M2/M3 Ultra 64GB+ — unified memory # ✓ Linux box with 48GB+ RAM (CPU-only) — CPU inference # ⚠ 32GB hosts (e.g. ASUS ROG Flow Z13) — set OLLAMA_KV_CACHE_TYPE=q8_0 # or lower num_ctx # # Throughput, measured 2026-09-18 on CPU only (Ryzen AI Max+ 395, Ollama 0.33.3, # isolated store): ./scripts/bench.sh aggregated 24.97 tok/s over its three-prompt # mix (4,543 tokens / 181,890 ms; 25.88 / 25.57 / 24.87 individually), about 5x the # 4.97 tok/s the dense 27B managed on the same CPU — that is the ~3B active # parameters per token. OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1 # aggregated 24.85 tok/s over the same mix, 0.5% below f16 — it buys memory, not # speed. Generation only: # this is a reasoning-first model, so an answer costs many more tokens than its # visible length suggests. # # To run on a 32 GB unified-memory laptop, override these in your local # Modelfile copy (or via `/set parameter` in the interactive `ollama run` REPL): # PARAMETER num_ctx 4096 # PARAMETER num_batch 256 # # If you have ≥48 GB RAM but want partial GPU offload, set: # PARAMETER num_gpu 24 # offload most layers (model has 40)