File size: 12,477 Bytes
f48fa07
64b629a
f48fa07
 
 
 
 
 
 
 
 
 
c0d1a26
f48fa07
64b629a
9dfadfe
 
 
 
f48fa07
 
 
be5b45f
936e88c
 
 
 
 
 
f48fa07
 
 
 
9ead6fb
936e88c
 
6708e8f
 
752eff3
9dfadfe
f48fa07
 
9dfadfe
 
 
936e88c
 
 
 
 
2e2b8f8
edf99d7
 
 
 
 
 
 
 
 
 
 
 
f48fa07
 
 
 
 
 
2e2b8f8
 
 
 
9dfadfe
 
 
 
2e2b8f8
 
 
 
 
 
 
ee4c9b0
2e2b8f8
ee4c9b0
64b629a
ee4c9b0
64b629a
 
 
 
 
 
 
752eff3
64b629a
 
 
 
 
 
 
 
 
3950347
64b629a
3950347
 
 
64b629a
 
edf99d7
3950347
 
 
 
9ead6fb
edf99d7
 
 
 
64b629a
3070964
 
 
64b629a
 
 
 
3950347
 
 
 
64b629a
3950347
 
64b629a
3950347
 
64b629a
f4db68f
64b629a
f4db68f
 
 
 
 
 
64b629a
 
 
 
 
6deeb72
64b629a
6deeb72
64b629a
f707944
64b629a
 
 
 
 
c0d1a26
39c0d86
64b629a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f48fa07
aa37e74
 
 
 
f48fa07
aa37e74
 
 
 
 
 
 
 
 
 
 
 
 
 
 
03b00ae
 
 
c0d1a26
f48fa07
aa37e74
 
 
c0d1a26
aa37e74
 
64b629a
03b00ae
c0d1a26
f48fa07
03b00ae
f48fa07
64b629a
09c1a7b
 
 
 
 
 
 
 
 
64b629a
 
f4db68f
64b629a
 
 
 
f48fa07
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
FROM ./Janus-35B-A3B.Q4_K_M.gguf

# The bundled GGUF carries no MTP / NextN block. The base
# (llmfan46/Qwen3.6-35B-A3B-uncensored-heretic) publishes its GGUFs already
# MTP-clean β€” blk.0…blk.39 only, block_count 40, no nextn_predict_layers key
# (read from the bundled file's own header, 2026-09-18) β€” so scripts/strip_mtp.py
# is a defensive no-op here, not a required build step. The separate
# llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF variant
# does keep the MTP head (its Q4_K_M has block_count 41, NextN block at index
# 40) and would need the strip; this repo does not ship it. If you want the MTP
# head, run the upstream safetensors under vLLM/SGLang, or take that variant's
# GGUF.

# Chat template β€” Qwen 3.6 ChatML in Ollama Go-template form, with the
# tool-calling blocks Ollama's capability detector looks for. Without a
# TEMPLATE that references .Tools and .ToolCalls, an Ollama that uses this
# template anyway (OLLAMA_GO_TEMPLATE=1 - per Ollama's source, not tested live -
# or one that predates the template comparison described below) rejects any
# request carrying a `tools` array with `<model> does not support tools`;
# Ollama 0.33.3 would instead switch to the GGUF's embedded template. Same template as the dense
# 27B sibling (FoolDev/Thanatos-27B-HERETIC) β€” the ChatML wire format is
# identical across Qwen 3.6 and 3.8.
#
# Thinking IS replayed across turns: every assistant message that carries
# reasoning renders <think>...</think>, earlier turns included, where Qwen's
# stock condition kept only the turn in progress (a tool-call chain included).
# The client must send the reasoning back - per Ollama's source, /api/chat reads
# each assistant message's `thinking` field and /v1/chat/completions reads only
# `reasoning` (it drops `reasoning_content` and `thinking` silently). Every
# retained trace stays in the prompt and prefill grows to match, which is the
# slow part on a CPU-only host. The prompt-token counts that used to sit here
# were measured on the dense 27B this repo shipped in the interim; they are
# removed rather than carried over, pending re-measurement.
#
# The thinking block in the assistant branch is load-bearing for a second
# reason. Ollama infers a Go template's "thinking" capability from a `.Thinking` field inside
# `range .Messages` wrapped in <think>/</think>, and Ollama 0.33.3 and 0.34.0 (verified;
# the selection code is unchanged through 0.34.2, checked in source)
# compare that capability list with the GGUF's embedded Jinja template at load
# time, switching to the embedded template if that one advertises more. Delete
# the block and Ollama switches: in one test on the dense 27B this repo shipped
# in the interim a single tool call then came back twice
# (a separate raw generation showed the model drafting the call inside <think>
# first; that the parser matched such a draft is inferred, not confirmed). Per Ollama's source (not tested live), a
# setup forcing this template (OLLAMA_GO_TEMPLATE=1) would also reject thinking
# requests with HTTP 400. To stop replaying earlier turns' reasoning, change the
# condition back to Qwen's stock
# `(and $.IsThinkSet (and .Thinking (or $last (gt $i $lastUserIdx))))` - the
# $lastUserIdx loop at the top of the template is kept for it; do not delete the
# block.
#
# Two more properties of this template are read from its PARSE TREE, not its
# output, so no render test can see them break. Ollama takes the tool-call tag
# from the first `{{ if }}` whose condition names .ToolCalls and then the first
# literal text in that block, so (a) the `range .ToolCalls` body must start with
# text, never an action - that was 0.9.5 - and (b) no earlier condition may name
# .ToolCalls, which is why the assistant branch aliases it as
# `{{ $calls := .ToolCalls }}` and branches on the variable. The thinking tags
# come from the first and last nodes of the list holding {{ .Thinking }}, both of
# which must be literal text, so nothing conditional may end that block - the
# blank-line condition sits after it. scripts/check_go_template.py emulates both
# of Ollama's readers and asserts what they derive.
#
# Reasoning effort: the block at the top of the template injects an effort
# instruction keyed on Ollama's think level ($.ThinkLevel), not the request
# value. It is Ollama-only now β€” this repo no longer ships chat_template.jinja,
# and the Qwen 3.6 base's own embedded template, which governs the llama.cpp
# path from here on, has no reasoning_effort handling at all.
# Ollama folds `reasoning_effort` into four levels first:
# high/xhigh/max/ultra -> high or max (xhigh line), low/minimal -> low (low
# line), medium and unset -> medium (no line), none -> thinking off. The
# outer `if $.ThinkLevel` keeps the template rendering on Ollama builds that
# predate the field.
TEMPLATE """{{- $lastUserIdx := -1 -}}
{{- range $idx, $msg := .Messages -}}
{{- if eq $msg.Role "user" }}{{ $lastUserIdx = $idx }}{{ end -}}
{{- end }}
{{- $effort := "" }}
{{- if $.ThinkLevel }}
{{- if or (eq $.ThinkLevel "high") (eq $.ThinkLevel "max") }}{{ $effort = "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." }}
{{- else if eq $.ThinkLevel "low" }}{{ $effort = "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." }}
{{- end }}
{{- end }}
{{- if or .System .Tools $effort }}<|im_start|>system
{{ if $effort }}{{ $effort }}{{ if or .System .Tools }}

{{ end }}{{ end }}{{ if .System }}{{ .System }}{{ if .Tools }}

{{ end }}{{ end }}
{{- if .Tools }}# Tools

You may call one or more functions to assist with the user query.

You are provided with function signatures within <tools></tools> XML tags:
<tools>
{{- range .Tools }}
{"type": "function", "function": {{ json .Function }}}
{{- end }}
</tools>

For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call>
{{- end -}}<|im_end|>
{{ end }}
{{- $prevIsTool := false }}
{{- range $i, $_ := .Messages }}
{{- $rest := slice $.Messages $i }}
{{- $last := eq (len $rest) 1 -}}
{{- $nextIsTool := and (gt (len $rest) 1) (eq (index $rest 1).Role "tool") -}}
{{- if eq .Role "user" }}<|im_start|>user
{{ .Content }}<|im_end|>
{{ else if eq .Role "assistant" }}{{ $calls := .ToolCalls }}{{ $blank := or .Content (not $calls) }}<|im_start|>assistant
{{ if $.IsThinkSet -}}
<think>
{{ .Thinking }}
</think>
{{ end -}}
{{ if and $.IsThinkSet $blank }}
{{ end -}}
{{ if .Content }}{{ .Content }}{{ if $calls }}
{{ end }}{{ end }}
{{- if .ToolCalls }}
{{- range .ToolCalls }}
<tool_call>
{"name": "{{ .Function.Name }}", "arguments": {{ .Function.Arguments }}}
</tool_call>
{{- end }}
{{- end }}{{ if not $last }}<|im_end|>
{{ end }}
{{- else if eq .Role "tool" }}
{{- if $prevIsTool }}
{{ else }}<|im_start|>user
{{ end }}<tool_response>
{{ .Content }}
</tool_response>
{{- if not $nextIsTool }}<|im_end|>
{{ end }}
{{- end }}
{{- $prevIsTool = eq .Role "tool" }}
{{- if and (ne .Role "assistant") $last }}<|im_start|>assistant
{{ if and $.IsThinkSet (not $.Think) -}}
<think>

</think>

{{ else -}}
<think>
{{ end -}}
{{ end }}
{{- end }}"""

# Sampling tuned for reasoning + general use. See README "Recommended sampling"
# for creative/RP alternatives.
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 0
PARAMETER repeat_penalty 1.05
PARAMETER num_ctx 262144

# Stop tokens. Without these, Ollama only honors <|im_end|> from the GGUF
# metadata; the model occasionally emits <|endoftext|> instead and Ollama
# keeps generating past it (synthesising a fake new user turn). Listing
# both β€” plus <|im_start|> as a belt-and-braces guard against the same
# loop β€” keeps responses cleanly terminated. Same fix the sibling
# (FoolDev/Thanatos-27B-HERETIC) shipped in commit 6672746.
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
PARAMETER stop "<|im_start|>"

SYSTEM """You are Janus, a precise and capable assistant for reasoning, writing, coding, and long-form dialogue.

Behavior rules:
- Answer the user's actual request directly.
- Be accurate, complete, and structured.
- Think before answering, but do not get stuck in repetitive loops or meta-commentary.
- If the request is ambiguous or incomplete, state what is missing and make the smallest reasonable assumption needed to continue.
- If the user wants creative writing, preserve tone, continuity, and character consistency.
- If the user wants analysis or technical help, prefer concrete steps, examples, and decisions over fluff.
- Finish with a usable answer, not just planning."""

# Hardware notes
# --------------
# This Q4_K_M is 21,233,608,512 bytes on disk β€” 21.23 GB decimal, 19.78 GiB.
# Measured 2026-09-18 on Ollama 0.33.3's CPU backend, from llama.cpp's own
# allocation lines, in an isolated store at num_ctx 8192 / 32768 / 65536:
#   weights               19.77 GiB (CPU model buffer 5361.34 MiB + CPU_REPACK
#                         14878.12 MiB); constant
#   KV cache              only 10 of the 40 layers are full-attention (indices
#                         3, 7, ... 39), and llama.cpp logs the cache as "10
#                         layers": 20,480 B/token at f16 = 160 / 640 / 1280 MiB
#                         at 8K / 32K / 64K, i.e. 0.625 GiB per 32K, and
#                         10,880 B/token at q8_0 = 340 / 680 MiB at 32K / 64K,
#                         i.e. 0.332 GiB per 32K. Exactly linear in num_ctx ->
#                         5.0 GiB at the 262144 default with f16, 2.66 with q8_0
#   recurrent state       62.81 MiB for the 30 Gated-DeltaNet layers (R f32 2.81
#                         + S f32 60.00) β€” identical at every context
#   compute buffer        160.04 / 208.04 / 544.07 MiB at 8K / 32K / 64K; grows
#                         faster than linearly, so it is not extrapolated
#   totals                20.7 GiB at 32768 and 21.6 GiB at 65536 (both sums of
#                         measured parts); ~25.4 GiB at the 262144 default with
#                         the f16 cache and ~23.0 GiB with
#                         OLLAMA_KV_CACHE_TYPE=q8_0 β€” those two carry the 64K
#                         compute buffer forward, so treat them as floors.
#                         No YaRN rope-scaling is baked in this GGUF, so output
#                         past ~262K degrades; treat 1.01M as an advertised
#                         ceiling.
#
# This is a 35B-A3B MoE across 40 layers: 256 experts, 8 routed per token plus
# one shared expert, ~34.7B parameters total and ~3B active per token. The active
# count cuts compute per token, not the memory floor β€” every expert has to be
# resident, so budget for the whole 19.77 GiB of weights.
#
# Working configurations (rows assume the 262144 default: ~25.4 GiB with the f16
# cache, ~23.0 GiB with q8_0):
#   βœ“ Single H100 80GB / A100 80GB              β€” full GPU offload
#   βœ“ RTX 5090 32GB / RTX 4090 24GB + 32GB RAM  β€” partial offload
#   βœ“ Mac Studio M2/M3 Ultra 64GB+              β€” unified memory
#   βœ“ Linux box with 48GB+ RAM (CPU-only)       β€” CPU inference
#   ⚠ 32GB hosts (e.g. ASUS ROG Flow Z13)       β€” set OLLAMA_KV_CACHE_TYPE=q8_0
#                                                 or lower num_ctx
#
# Throughput, measured 2026-09-18 on CPU only (Ryzen AI Max+ 395, Ollama 0.33.3,
# isolated store): ./scripts/bench.sh aggregated 24.97 tok/s over its three-prompt
# mix (4,543 tokens / 181,890 ms; 25.88 / 25.57 / 24.87 individually), about 5x the
# 4.97 tok/s the dense 27B managed on the same CPU β€” that is the ~3B active
# parameters per token. OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1
# aggregated 24.85 tok/s over the same mix, 0.5% below f16 β€” it buys memory, not
# speed. Generation only:
# this is a reasoning-first model, so an answer costs many more tokens than its
# visible length suggests.
#
# To run on a 32 GB unified-memory laptop, override these in your local
# Modelfile copy (or via `/set parameter` in the interactive `ollama run` REPL):
#   PARAMETER num_ctx 4096
#   PARAMETER num_batch 256
#
# If you have β‰₯48 GB RAM but want partial GPU offload, set:
#   PARAMETER num_gpu 24    # offload most layers (model has 40)