Wolfram โ€” Cyber-SME Analysis Model (HoneySentinel)

Wolfram is a locally-deployed 9B language model that acts as the cybersecurity subject-matter-expert / analysis stage of the HoneySentinel honeypot ("Chimera" stage 2). It reads captured attacker activity โ€” shell transcripts, uploaded payloads, logs โ€” and produces structured analysis: intent, MITRE ATT&CK technique mapping, IOCs, de-obfuscation, and sophistication assessment.

  • Author: Hadi
  • Date: 2026-10-07
  • Base model: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (MIT)
  • Type: decoder-only causal LM, text-only (vision tower removed)
  • Lineage: base โ†’ cyber LoRA fine-tune โ†’ merge โ†’ abliteration (refusal removal)
  • Status: research artifact โ€” published for educational / defensive-security research

What it is / what was done

Wolfram is built in three stages:

  1. Cyber domain fine-tune (QLoRA). A LoRA adapter was trained on a curated ~7,300-example cybersecurity mix, then merged into the base.
  2. Abliteration (refusal removal), applied last. A norm-preserving orthogonal edit removes the model's refusal direction so it will analyze malicious-looking input instead of declining. Done after the fine-tune on purpose โ€” see "Ordering".
  3. Quantization to GGUF (Q4_K_M / Q6_K / Q8_0) for local CPU deployment.

The result: a model that reliably emits the honeypot's strict ATT&CK-JSON analysis format, retains the base's general reasoning, and does not refuse security-analysis tasks.

Architecture

qwen3_5 hybrid attention: 32 transformer layers (linear/gated-delta-rule attention, with full attention every 4th layer), hidden size 4096, intermediate 12288, vocab 248,320. The base also ships a vision encoder and declares an MTP (multi-token / NextN speculative-decode) layer; neither is present in Wolfram โ€” the vision tower is dropped (text-only), and the base never shipped MTP weights, so speculative decoding is not available for this checkpoint.

Training

Fine-tune (QLoRA): 2ร— NVIDIA T4 (Kaggle), 1 epoch, 458 optimizer steps, train_loss 0.74 (plateaued 0.65โ€“0.70). LoRA rank 16 / ฮฑ 32, all-linear targets (43 M trainable params, 0.48%), 4-bit NF4 base, fp16 compute, paged AdamW 8-bit, LR 2e-4 cosine. Loss computed on the assistant turn; <think> reasoning traces kept.

Data mix (~7,328 examples):

Source License Count Role
Primus-Reasoning ODC-BY 2,500 cyber tasks with reasoning traces
Primus-Instruct ODC-BY 835 cyber instruction/chat
Primus-Seed (subset) ODC-BY 1,000 cyber knowledge (attack/defense/threat-intel)
uka-cyber Apache-2.0 1,200 malware RE, pentest, exploitation
dolly-15k (subset) CC-BY-SA 1,500 general-domain replay (anti-forgetting)
in-domain (generated) โ€” 150 ร—2 honeypot command โ†’ ATT&CK-JSON (the deployment task)

Abliteration (applied to the merged model): refusal direction extracted from the response region (the generation, where a reasoning model actually decides to refuse โ€” not the prompt token), then norm-preserving biprojected orthogonalization (grimjim-style MPOA) across layers 6โ€“31 of down_proj / attention output / token embeddings (ฮฑ 1.0, k 1, no biprojection; 53 tensors edited).

Ordering note: the cyber fine-tune raised refusals (78% โ†’ 97%) because the training corpora are safety-flavored. Abliterating after the fine-tune is what brings refusals to 0%; doing it first would have been undone by the SFT.

Evaluation (held-out)

Benchmarked base vs. the fine-tuned merge vs. final Wolfram. Cyber MCQ = CyberMetric + SecEval (650 Q); general = MMLU (200 Q); in-domain = ATT&CK-JSON on held-out command sessions.

Metric base +cyber (merged) Wolfram (final)
Refusal rate 78% 97% 0%
Math accuracy 100% 100% 100%
General (MMLU) 0.47 0.52 0.505
Cyber MCQ 0.528 0.537 0.537
In-domain JSON-valid 0.80 1.00 1.00
In-domain ATT&CK F1 0.148 0.898 0.89

Reading it: the fine-tune's win is the deployment task โ€” ATT&CK technique F1 rose 6ร— (0.15 โ†’ 0.90) with perfect JSON formatting. General ability was retained (no forgetting; perplexity moved <5%). Standalone cyber-knowledge MCQ did not improve (the base already knows the facts; moving that needs more/harder knowledge data, not more epochs). Abliteration is competence-neutral (Wolfram โ‰ˆ merged) while taking refusals to zero.

Formats & performance

  • wolfram-final/ โ€” bf16 safetensors (merged + abliterated, HF format)
  • wolfram-Q4_K_M.gguf (5.6 GB) ยท wolfram-Q6_K.gguf (7.4 GB) ยท wolfram-Q8_0.gguf (9.5 GB)
  • CPU on an i7-9700 (8 cores), Q4_K_M, -t 8: prompt eval โ‰ˆ 38 tok/s, generation โ‰ˆ 3.4 tok/s. Generation is memory-bandwidth-bound โ€” it does not improve with more threads or flash-attention (benchmarked); the only way higher is a smaller quant (quality cost) or a GPU (none can hold a 9B here). Q4 is the fastest of the three. GGUF conversion required a one-line converter patch (no_mtp=True), since this checkpoint has no MTP head.

Intended use

Defensive honeypot analysis of already-captured, contained attacker artifacts: intent summarization, ATT&CK mapping, IOC extraction, payload/command de-obfuscation. Authorized academic security-research context (ADU capstone), run locally.

Out of scope / responsible use

  • Refusals are removed. Wolfram will engage with offensive-security content; it is not safety-aligned. Deploy it behind application-level content controls โ€” specifically a filter for the mass-casualty / CBRN cluster (explosives, bio, chem, poisoning). Applications remain responsible for their own content policy; gate it where request/response text can actually be inspected, not in the inference engine.
  • Not a general assistant; tuned/evaluated for the honeypot analysis task.
  • No speculative decoding (no MTP weights in the checkpoint).

Limitations

  • In-domain F1 (0.90) is measured on synthetic held-out sessions from the same generator family as training โ€” it proves the format + mapping were learned, but is somewhat in-distribution. Validate on real captures before relying on it.
  • Cyber-knowledge MCQ is flat vs. base; the improvement is task/format competence.
  • QLoRA (4-bit) training and Q4 inference both cap fidelity vs. full-precision.
  • Abliteration removes the dominant linear refusal direction; a small residual of other-mechanism refusals can remain.

Reproducibility

Pipeline (on the HoneySentinel server, ~/ablit): response-region direction extraction โ†’ direction scoring โ†’ MPOA ablation โ†’ capability/refusal eval; GGUF via a patched convert_hf_to_gguf.py + llama.cpp llama-quantize. Fine-tune ran from a Kaggle kernel; adapter checkpointed to HF Hub. Method write-up: MiMo-Abliteration-README.md.

License & attribution

Base model MIT. Training data under their respective licenses (ODC-BY, Apache-2.0, CC-BY-SA-3.0; SecEval CC-BY-NC used for evaluation only).

Downloads last month
87
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Flexingmeow/wolfram-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(57)
this model