Instructions to use Flexingmeow/wolfram-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Flexingmeow/wolfram-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Flexingmeow/wolfram-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Flexingmeow/wolfram-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Flexingmeow/wolfram-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Flexingmeow/wolfram-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Flexingmeow/wolfram-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Flexingmeow/wolfram-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Flexingmeow/wolfram-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Flexingmeow/wolfram-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Flexingmeow/wolfram-GGUF:Q4_K_M
- Ollama
How to use Flexingmeow/wolfram-GGUF with Ollama:
ollama run hf.co/Flexingmeow/wolfram-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Flexingmeow/wolfram-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Flexingmeow/wolfram-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Flexingmeow/wolfram-GGUF with Docker Model Runner:
docker model run hf.co/Flexingmeow/wolfram-GGUF:Q4_K_M
- Lemonade
How to use Flexingmeow/wolfram-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Flexingmeow/wolfram-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.wolfram-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Flexingmeow/wolfram-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Flexingmeow/wolfram-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Flexingmeow/wolfram-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Flexingmeow/wolfram-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Flexingmeow/wolfram-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Wolfram โ Cyber-SME Analysis Model (HoneySentinel)
Wolfram is a locally-deployed 9B language model that acts as the cybersecurity subject-matter-expert / analysis stage of the HoneySentinel honeypot ("Chimera" stage 2). It reads captured attacker activity โ shell transcripts, uploaded payloads, logs โ and produces structured analysis: intent, MITRE ATT&CK technique mapping, IOCs, de-obfuscation, and sophistication assessment.
- Author: Hadi
- Date: 2026-10-07
- Base model:
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B(MIT) - Type: decoder-only causal LM, text-only (vision tower removed)
- Lineage: base โ cyber LoRA fine-tune โ merge โ abliteration (refusal removal)
- Status: research artifact โ published for educational / defensive-security research
What it is / what was done
Wolfram is built in three stages:
- Cyber domain fine-tune (QLoRA). A LoRA adapter was trained on a curated ~7,300-example cybersecurity mix, then merged into the base.
- Abliteration (refusal removal), applied last. A norm-preserving orthogonal edit removes the model's refusal direction so it will analyze malicious-looking input instead of declining. Done after the fine-tune on purpose โ see "Ordering".
- Quantization to GGUF (Q4_K_M / Q6_K / Q8_0) for local CPU deployment.
The result: a model that reliably emits the honeypot's strict ATT&CK-JSON analysis format, retains the base's general reasoning, and does not refuse security-analysis tasks.
Architecture
qwen3_5 hybrid attention: 32 transformer layers (linear/gated-delta-rule attention,
with full attention every 4th layer), hidden size 4096, intermediate 12288, vocab
248,320. The base also ships a vision encoder and declares an MTP (multi-token /
NextN speculative-decode) layer; neither is present in Wolfram โ the vision tower
is dropped (text-only), and the base never shipped MTP weights, so speculative
decoding is not available for this checkpoint.
Training
Fine-tune (QLoRA): 2ร NVIDIA T4 (Kaggle), 1 epoch, 458 optimizer steps,
train_loss 0.74 (plateaued 0.65โ0.70). LoRA rank 16 / ฮฑ 32, 43 M trainable params, 0.48%), 4-bit NF4 base, fp16 compute, paged AdamW 8-bit,
LR 2e-4 cosine. Loss computed on the assistant turn; all-linear targets
(<think> reasoning traces kept.
Data mix (~7,328 examples):
| Source | License | Count | Role |
|---|---|---|---|
| Primus-Reasoning | ODC-BY | 2,500 | cyber tasks with reasoning traces |
| Primus-Instruct | ODC-BY | 835 | cyber instruction/chat |
| Primus-Seed (subset) | ODC-BY | 1,000 | cyber knowledge (attack/defense/threat-intel) |
| uka-cyber | Apache-2.0 | 1,200 | malware RE, pentest, exploitation |
| dolly-15k (subset) | CC-BY-SA | 1,500 | general-domain replay (anti-forgetting) |
| in-domain (generated) | โ | 150 ร2 | honeypot command โ ATT&CK-JSON (the deployment task) |
Abliteration (applied to the merged model): refusal direction extracted from the
response region (the generation, where a reasoning model actually decides to
refuse โ not the prompt token), then norm-preserving biprojected orthogonalization
(grimjim-style MPOA) across layers 6โ31 of down_proj / attention output / token
embeddings (ฮฑ 1.0, k 1, no biprojection; 53 tensors edited).
Ordering note: the cyber fine-tune raised refusals (78% โ 97%) because the training corpora are safety-flavored. Abliterating after the fine-tune is what brings refusals to 0%; doing it first would have been undone by the SFT.
Evaluation (held-out)
Benchmarked base vs. the fine-tuned merge vs. final Wolfram. Cyber MCQ = CyberMetric + SecEval (650 Q); general = MMLU (200 Q); in-domain = ATT&CK-JSON on held-out command sessions.
| Metric | base | +cyber (merged) | Wolfram (final) |
|---|---|---|---|
| Refusal rate | 78% | 97% | 0% |
| Math accuracy | 100% | 100% | 100% |
| General (MMLU) | 0.47 | 0.52 | 0.505 |
| Cyber MCQ | 0.528 | 0.537 | 0.537 |
| In-domain JSON-valid | 0.80 | 1.00 | 1.00 |
| In-domain ATT&CK F1 | 0.148 | 0.898 | 0.89 |
Reading it: the fine-tune's win is the deployment task โ ATT&CK technique F1 rose 6ร (0.15 โ 0.90) with perfect JSON formatting. General ability was retained (no forgetting; perplexity moved <5%). Standalone cyber-knowledge MCQ did not improve (the base already knows the facts; moving that needs more/harder knowledge data, not more epochs). Abliteration is competence-neutral (Wolfram โ merged) while taking refusals to zero.
Formats & performance
wolfram-final/โ bf16 safetensors (merged + abliterated, HF format)wolfram-Q4_K_M.gguf(5.6 GB) ยทwolfram-Q6_K.gguf(7.4 GB) ยทwolfram-Q8_0.gguf(9.5 GB)- CPU on an i7-9700 (8 cores), Q4_K_M,
-t 8: prompt eval โ 38 tok/s, generation โ 3.4 tok/s. Generation is memory-bandwidth-bound โ it does not improve with more threads or flash-attention (benchmarked); the only way higher is a smaller quant (quality cost) or a GPU (none can hold a 9B here). Q4 is the fastest of the three. GGUF conversion required a one-line converter patch (no_mtp=True), since this checkpoint has no MTP head.
Intended use
Defensive honeypot analysis of already-captured, contained attacker artifacts: intent summarization, ATT&CK mapping, IOC extraction, payload/command de-obfuscation. Authorized academic security-research context (ADU capstone), run locally.
Out of scope / responsible use
- Refusals are removed. Wolfram will engage with offensive-security content; it is not safety-aligned. Deploy it behind application-level content controls โ specifically a filter for the mass-casualty / CBRN cluster (explosives, bio, chem, poisoning). Applications remain responsible for their own content policy; gate it where request/response text can actually be inspected, not in the inference engine.
- Not a general assistant; tuned/evaluated for the honeypot analysis task.
- No speculative decoding (no MTP weights in the checkpoint).
Limitations
- In-domain F1 (0.90) is measured on synthetic held-out sessions from the same generator family as training โ it proves the format + mapping were learned, but is somewhat in-distribution. Validate on real captures before relying on it.
- Cyber-knowledge MCQ is flat vs. base; the improvement is task/format competence.
- QLoRA (4-bit) training and Q4 inference both cap fidelity vs. full-precision.
- Abliteration removes the dominant linear refusal direction; a small residual of other-mechanism refusals can remain.
Reproducibility
Pipeline (on the HoneySentinel server, ~/ablit): response-region direction
extraction โ direction scoring โ MPOA ablation โ capability/refusal eval; GGUF via a
patched convert_hf_to_gguf.py + llama.cpp llama-quantize. Fine-tune ran from a
Kaggle kernel; adapter checkpointed to HF Hub. Method write-up:
MiMo-Abliteration-README.md.
License & attribution
Base model MIT. Training data under their respective licenses (ODC-BY, Apache-2.0, CC-BY-SA-3.0; SecEval CC-BY-NC used for evaluation only).
- Downloads last month
- 87
4-bit
6-bit
8-bit
Model tree for Flexingmeow/wolfram-GGUF
Base model
Qwen/Qwen3.5-9B-Base