SecOpsBot-0.6B 🛡️

A 0.6B-parameter grounded cybersecurity controls assistant. Give it an excerpt from a security standard (NIST SP 800-53, 800-171, 800-61, FIPS, ...) and a task, and it produces a summary, key facts, actionable checklist, lifecycle steps, rationale or definition, citing control identifiers only when they appear in the excerpt.

Small enough to run fully offline on a laptop (see the GGUF version), which matters for compliance and security teams who can't send internal material to cloud LLMs. Works well as the "reader" in a RAG pipeline over security documentation.

Results

Evaluated on 60 rows from the test split, whose source documents were never seen in training.

Model ROUGE-L ↑ Control-ID hallucination ↓ Control-ID recall ↑ Avg words
Qwen3-0.6B (base) 0.279 4.2% 32.7% 98
SecOpsBot-0.6B 0.913 2.0% 95.9% 123
  • Control-ID hallucination: share of control IDs in the answer (e.g. AC-2, SC-7(3)) that do not appear in the source excerpt.
  • Control-ID recall: share of control IDs in the reference answer that the model also cites.

Usage

The model is trained with this system prompt. Always use it:

You are SecOpsBot, a cybersecurity controls and incident response assistant. Answer only using information from the provided source excerpt. Be precise, cite control identifiers when the excerpt names them, and be actionable. If the excerpt does not contain the answer, say so.

Example user message format (from the dataset):

Based on this excerpt about assessing security and privacy controls, what should an organization do to PREPARE for, DETECT, and RESPOND to this issue?

SOURCE:
POTENTIAL ASSESSMENT METHODS AND OBJECTS: SA-08(20)-Examine [SELECT FROM: System and services acquision policy; procedures addressing the security design principle of metadata management used in the specificaon, design, development, implementaon, and modificaon of the system; system design documentaon; security and privacy requirements and specificaons for the system; system security and privacy architecture; system security plan; other relevant documents or records]. SA-08(20)-Interview [SELECT FROM: Organizaonal personnel wit
...
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("nuhmanpk/secopsbot-0.6b")
model = AutoModelForCausalLM.from_pretrained("nuhmanpk/secopsbot-0.6b", device_map="auto")

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": "<source excerpt + task, formatted as above>"},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False,
                                 return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False, repetition_penalty=1.05)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Run locally with Ollama:

ollama run hf.co/nuhmanpk/secopsbot-0.6b-GGUF:Q4_K_M

Training

  • Base: unsloth/Qwen3-0.6B, LoRA r=16 on all attention + MLP projections, merged to 16-bit
  • Data: 13,106 training rows from nuhmanpk/cybersecurity-controls-instructions (train split, 1 epoch), loss on assistant turns only
  • lr 2e-4 cosine, effective batch 16, max length 4096, trained on a Colab T4 with Unsloth

Limitations

  • Answers are only as good as the excerpt you provide; this is a reader, not a knowledge base.
  • Inherits PDF text-extraction noise from the source documents.
  • Not a substitute for official NIST guidance, a qualified assessor, or legal/compliance advice. Verify before acting.
  • English only; evaluated on a small held-out sample.

Author

Built by Nuhman PK

Citation

@misc{secopsbot2026,
  author = {Nuhman PK},
  title  = {SecOpsBot-0.6B: A Tiny Grounded Cybersecurity Controls Assistant},
  year   = {2026},
  url    = {https://huggingface.co/nuhmanpk/secopsbot-0.6b}
}
Downloads last month
304
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nuhmanpk/secopsbot-0.6b

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1332)
this model

Dataset used to train nuhmanpk/secopsbot-0.6b