Text Generation
GGUF
qwen36
Mixture of Experts
conversational
multimodal
agent
ollama
heretic
uncensored
reasoning
distillation
Instructions to use FoolDev/Janus-35B-HERETIC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FoolDev/Janus-35B-HERETIC with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use FoolDev/Janus-35B-HERETIC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FoolDev/Janus-35B-HERETIC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FoolDev/Janus-35B-HERETIC", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Ollama
How to use FoolDev/Janus-35B-HERETIC with Ollama:
ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Unsloth Desktop
- Pi
How to use FoolDev/Janus-35B-HERETIC with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FoolDev/Janus-35B-HERETIC:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FoolDev/Janus-35B-HERETIC with Docker Model Runner:
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Lemonade
How to use FoolDev/Janus-35B-HERETIC with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FoolDev/Janus-35B-HERETIC:Q4_K_M
Run and chat with the model
lemonade run user.Janus-35B-HERETIC-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use FoolDev/Janus-35B-HERETIC with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FoolDev/Janus-35B-HERETIC:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FoolDev/Janus-35B-HERETIC with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FoolDev/Janus-35B-HERETIC:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 12,477 Bytes
f48fa07 64b629a f48fa07 c0d1a26 f48fa07 64b629a 9dfadfe f48fa07 be5b45f 936e88c f48fa07 9ead6fb 936e88c 6708e8f 752eff3 9dfadfe f48fa07 9dfadfe 936e88c 2e2b8f8 edf99d7 f48fa07 2e2b8f8 9dfadfe 2e2b8f8 ee4c9b0 2e2b8f8 ee4c9b0 64b629a ee4c9b0 64b629a 752eff3 64b629a 3950347 64b629a 3950347 64b629a edf99d7 3950347 9ead6fb edf99d7 64b629a 3070964 64b629a 3950347 64b629a 3950347 64b629a 3950347 64b629a f4db68f 64b629a f4db68f 64b629a 6deeb72 64b629a 6deeb72 64b629a f707944 64b629a c0d1a26 39c0d86 64b629a f48fa07 aa37e74 f48fa07 aa37e74 03b00ae c0d1a26 f48fa07 aa37e74 c0d1a26 aa37e74 64b629a 03b00ae c0d1a26 f48fa07 03b00ae f48fa07 64b629a 09c1a7b 64b629a f4db68f 64b629a f48fa07 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 | FROM ./Janus-35B-A3B.Q4_K_M.gguf
# The bundled GGUF carries no MTP / NextN block. The base
# (llmfan46/Qwen3.6-35B-A3B-uncensored-heretic) publishes its GGUFs already
# MTP-clean β blk.0β¦blk.39 only, block_count 40, no nextn_predict_layers key
# (read from the bundled file's own header, 2026-09-18) β so scripts/strip_mtp.py
# is a defensive no-op here, not a required build step. The separate
# llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF variant
# does keep the MTP head (its Q4_K_M has block_count 41, NextN block at index
# 40) and would need the strip; this repo does not ship it. If you want the MTP
# head, run the upstream safetensors under vLLM/SGLang, or take that variant's
# GGUF.
# Chat template β Qwen 3.6 ChatML in Ollama Go-template form, with the
# tool-calling blocks Ollama's capability detector looks for. Without a
# TEMPLATE that references .Tools and .ToolCalls, an Ollama that uses this
# template anyway (OLLAMA_GO_TEMPLATE=1 - per Ollama's source, not tested live -
# or one that predates the template comparison described below) rejects any
# request carrying a `tools` array with `<model> does not support tools`;
# Ollama 0.33.3 would instead switch to the GGUF's embedded template. Same template as the dense
# 27B sibling (FoolDev/Thanatos-27B-HERETIC) β the ChatML wire format is
# identical across Qwen 3.6 and 3.8.
#
# Thinking IS replayed across turns: every assistant message that carries
# reasoning renders <think>...</think>, earlier turns included, where Qwen's
# stock condition kept only the turn in progress (a tool-call chain included).
# The client must send the reasoning back - per Ollama's source, /api/chat reads
# each assistant message's `thinking` field and /v1/chat/completions reads only
# `reasoning` (it drops `reasoning_content` and `thinking` silently). Every
# retained trace stays in the prompt and prefill grows to match, which is the
# slow part on a CPU-only host. The prompt-token counts that used to sit here
# were measured on the dense 27B this repo shipped in the interim; they are
# removed rather than carried over, pending re-measurement.
#
# The thinking block in the assistant branch is load-bearing for a second
# reason. Ollama infers a Go template's "thinking" capability from a `.Thinking` field inside
# `range .Messages` wrapped in <think>/</think>, and Ollama 0.33.3 and 0.34.0 (verified;
# the selection code is unchanged through 0.34.2, checked in source)
# compare that capability list with the GGUF's embedded Jinja template at load
# time, switching to the embedded template if that one advertises more. Delete
# the block and Ollama switches: in one test on the dense 27B this repo shipped
# in the interim a single tool call then came back twice
# (a separate raw generation showed the model drafting the call inside <think>
# first; that the parser matched such a draft is inferred, not confirmed). Per Ollama's source (not tested live), a
# setup forcing this template (OLLAMA_GO_TEMPLATE=1) would also reject thinking
# requests with HTTP 400. To stop replaying earlier turns' reasoning, change the
# condition back to Qwen's stock
# `(and $.IsThinkSet (and .Thinking (or $last (gt $i $lastUserIdx))))` - the
# $lastUserIdx loop at the top of the template is kept for it; do not delete the
# block.
#
# Two more properties of this template are read from its PARSE TREE, not its
# output, so no render test can see them break. Ollama takes the tool-call tag
# from the first `{{ if }}` whose condition names .ToolCalls and then the first
# literal text in that block, so (a) the `range .ToolCalls` body must start with
# text, never an action - that was 0.9.5 - and (b) no earlier condition may name
# .ToolCalls, which is why the assistant branch aliases it as
# `{{ $calls := .ToolCalls }}` and branches on the variable. The thinking tags
# come from the first and last nodes of the list holding {{ .Thinking }}, both of
# which must be literal text, so nothing conditional may end that block - the
# blank-line condition sits after it. scripts/check_go_template.py emulates both
# of Ollama's readers and asserts what they derive.
#
# Reasoning effort: the block at the top of the template injects an effort
# instruction keyed on Ollama's think level ($.ThinkLevel), not the request
# value. It is Ollama-only now β this repo no longer ships chat_template.jinja,
# and the Qwen 3.6 base's own embedded template, which governs the llama.cpp
# path from here on, has no reasoning_effort handling at all.
# Ollama folds `reasoning_effort` into four levels first:
# high/xhigh/max/ultra -> high or max (xhigh line), low/minimal -> low (low
# line), medium and unset -> medium (no line), none -> thinking off. The
# outer `if $.ThinkLevel` keeps the template rendering on Ollama builds that
# predate the field.
TEMPLATE """{{- $lastUserIdx := -1 -}}
{{- range $idx, $msg := .Messages -}}
{{- if eq $msg.Role "user" }}{{ $lastUserIdx = $idx }}{{ end -}}
{{- end }}
{{- $effort := "" }}
{{- if $.ThinkLevel }}
{{- if or (eq $.ThinkLevel "high") (eq $.ThinkLevel "max") }}{{ $effort = "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." }}
{{- else if eq $.ThinkLevel "low" }}{{ $effort = "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." }}
{{- end }}
{{- end }}
{{- if or .System .Tools $effort }}<|im_start|>system
{{ if $effort }}{{ $effort }}{{ if or .System .Tools }}
{{ end }}{{ end }}{{ if .System }}{{ .System }}{{ if .Tools }}
{{ end }}{{ end }}
{{- if .Tools }}# Tools
You may call one or more functions to assist with the user query.
You are provided with function signatures within <tools></tools> XML tags:
<tools>
{{- range .Tools }}
{"type": "function", "function": {{ json .Function }}}
{{- end }}
</tools>
For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call>
{{- end -}}<|im_end|>
{{ end }}
{{- $prevIsTool := false }}
{{- range $i, $_ := .Messages }}
{{- $rest := slice $.Messages $i }}
{{- $last := eq (len $rest) 1 -}}
{{- $nextIsTool := and (gt (len $rest) 1) (eq (index $rest 1).Role "tool") -}}
{{- if eq .Role "user" }}<|im_start|>user
{{ .Content }}<|im_end|>
{{ else if eq .Role "assistant" }}{{ $calls := .ToolCalls }}{{ $blank := or .Content (not $calls) }}<|im_start|>assistant
{{ if $.IsThinkSet -}}
<think>
{{ .Thinking }}
</think>
{{ end -}}
{{ if and $.IsThinkSet $blank }}
{{ end -}}
{{ if .Content }}{{ .Content }}{{ if $calls }}
{{ end }}{{ end }}
{{- if .ToolCalls }}
{{- range .ToolCalls }}
<tool_call>
{"name": "{{ .Function.Name }}", "arguments": {{ .Function.Arguments }}}
</tool_call>
{{- end }}
{{- end }}{{ if not $last }}<|im_end|>
{{ end }}
{{- else if eq .Role "tool" }}
{{- if $prevIsTool }}
{{ else }}<|im_start|>user
{{ end }}<tool_response>
{{ .Content }}
</tool_response>
{{- if not $nextIsTool }}<|im_end|>
{{ end }}
{{- end }}
{{- $prevIsTool = eq .Role "tool" }}
{{- if and (ne .Role "assistant") $last }}<|im_start|>assistant
{{ if and $.IsThinkSet (not $.Think) -}}
<think>
</think>
{{ else -}}
<think>
{{ end -}}
{{ end }}
{{- end }}"""
# Sampling tuned for reasoning + general use. See README "Recommended sampling"
# for creative/RP alternatives.
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 0
PARAMETER repeat_penalty 1.05
PARAMETER num_ctx 262144
# Stop tokens. Without these, Ollama only honors <|im_end|> from the GGUF
# metadata; the model occasionally emits <|endoftext|> instead and Ollama
# keeps generating past it (synthesising a fake new user turn). Listing
# both β plus <|im_start|> as a belt-and-braces guard against the same
# loop β keeps responses cleanly terminated. Same fix the sibling
# (FoolDev/Thanatos-27B-HERETIC) shipped in commit 6672746.
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
PARAMETER stop "<|im_start|>"
SYSTEM """You are Janus, a precise and capable assistant for reasoning, writing, coding, and long-form dialogue.
Behavior rules:
- Answer the user's actual request directly.
- Be accurate, complete, and structured.
- Think before answering, but do not get stuck in repetitive loops or meta-commentary.
- If the request is ambiguous or incomplete, state what is missing and make the smallest reasonable assumption needed to continue.
- If the user wants creative writing, preserve tone, continuity, and character consistency.
- If the user wants analysis or technical help, prefer concrete steps, examples, and decisions over fluff.
- Finish with a usable answer, not just planning."""
# Hardware notes
# --------------
# This Q4_K_M is 21,233,608,512 bytes on disk β 21.23 GB decimal, 19.78 GiB.
# Measured 2026-09-18 on Ollama 0.33.3's CPU backend, from llama.cpp's own
# allocation lines, in an isolated store at num_ctx 8192 / 32768 / 65536:
# weights 19.77 GiB (CPU model buffer 5361.34 MiB + CPU_REPACK
# 14878.12 MiB); constant
# KV cache only 10 of the 40 layers are full-attention (indices
# 3, 7, ... 39), and llama.cpp logs the cache as "10
# layers": 20,480 B/token at f16 = 160 / 640 / 1280 MiB
# at 8K / 32K / 64K, i.e. 0.625 GiB per 32K, and
# 10,880 B/token at q8_0 = 340 / 680 MiB at 32K / 64K,
# i.e. 0.332 GiB per 32K. Exactly linear in num_ctx ->
# 5.0 GiB at the 262144 default with f16, 2.66 with q8_0
# recurrent state 62.81 MiB for the 30 Gated-DeltaNet layers (R f32 2.81
# + S f32 60.00) β identical at every context
# compute buffer 160.04 / 208.04 / 544.07 MiB at 8K / 32K / 64K; grows
# faster than linearly, so it is not extrapolated
# totals 20.7 GiB at 32768 and 21.6 GiB at 65536 (both sums of
# measured parts); ~25.4 GiB at the 262144 default with
# the f16 cache and ~23.0 GiB with
# OLLAMA_KV_CACHE_TYPE=q8_0 β those two carry the 64K
# compute buffer forward, so treat them as floors.
# No YaRN rope-scaling is baked in this GGUF, so output
# past ~262K degrades; treat 1.01M as an advertised
# ceiling.
#
# This is a 35B-A3B MoE across 40 layers: 256 experts, 8 routed per token plus
# one shared expert, ~34.7B parameters total and ~3B active per token. The active
# count cuts compute per token, not the memory floor β every expert has to be
# resident, so budget for the whole 19.77 GiB of weights.
#
# Working configurations (rows assume the 262144 default: ~25.4 GiB with the f16
# cache, ~23.0 GiB with q8_0):
# β Single H100 80GB / A100 80GB β full GPU offload
# β RTX 5090 32GB / RTX 4090 24GB + 32GB RAM β partial offload
# β Mac Studio M2/M3 Ultra 64GB+ β unified memory
# β Linux box with 48GB+ RAM (CPU-only) β CPU inference
# β 32GB hosts (e.g. ASUS ROG Flow Z13) β set OLLAMA_KV_CACHE_TYPE=q8_0
# or lower num_ctx
#
# Throughput, measured 2026-09-18 on CPU only (Ryzen AI Max+ 395, Ollama 0.33.3,
# isolated store): ./scripts/bench.sh aggregated 24.97 tok/s over its three-prompt
# mix (4,543 tokens / 181,890 ms; 25.88 / 25.57 / 24.87 individually), about 5x the
# 4.97 tok/s the dense 27B managed on the same CPU β that is the ~3B active
# parameters per token. OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1
# aggregated 24.85 tok/s over the same mix, 0.5% below f16 β it buys memory, not
# speed. Generation only:
# this is a reasoning-first model, so an answer costs many more tokens than its
# visible length suggests.
#
# To run on a 32 GB unified-memory laptop, override these in your local
# Modelfile copy (or via `/set parameter` in the interactive `ollama run` REPL):
# PARAMETER num_ctx 4096
# PARAMETER num_batch 256
#
# If you have β₯48 GB RAM but want partial GPU offload, set:
# PARAMETER num_gpu 24 # offload most layers (model has 40)
|