RedSage-K-SFT

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
(NeurIPS 2026 Evaluations and Datasets Track)
Authors: Pengfei Li1*, Naufal Suryanto1*, Sicheng Zhang1, Muzammal Naseer1,2
1Khalifa University, 2University of Western Australia
*Equal contribution

RISys-Lab on Hugging Face
🌐 Project Page  |   💻 GitHub Code  |   🤗 Datasets & Models

Model summary

RedSage-K-SFT is an 8B cybersecurity model for translating natural-language requests into Kali/Linux commands. It is fine-tuned from RedSage-Qwen3-8B-Ins using supervised fine-tuning on KaliBench. It corresponds to “RedSage-K (SFT)” on our paper and project page.

Property Value
Developer RISys-Lab, Khalifa University
Architecture Qwen3ForCausalLM, 36 layers
Release format Merged LoRA weights, BF16 Safetensors
Language English
Output format <output>command</output>

Training

KaliBench contains 8,504 verified query-command pairs spanning 1,642 sub-tools and 23 capability dimensions. SFT uses 3,504 training pairs, each presented in three modes, for 10,512 examples. The remaining 5,000 pairs form the test split.

Mode Model input
Unrestricted Query only
Restricted Query and candidate tools
Hinted Query, target tool, and usage documentation

Training applies response-only negative log-likelihood to reference commands, without reasoning traces. Appendix G.2 reports:

Setting Value
Hardware / duration One NVIDIA H200 (141 GB) / approximately 32 minutes
Epochs / effective batch size 2 / 32 (4 per device × 8 accumulation steps)
Optimizer / learning rate 8-bit AdamW / 5e-6
Schedule / warmup / weight decay Linear / 10% / 1e-3
LoRA rank / alpha / dropout 64 / 128 / 0
LoRA targets q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Precision / gradient checkpointing BF16 / enabled
Packing / padding-free training Disabled / enabled
Software Transformers 4.57.6, Unsloth 2026.4.6, PEFT 0.19.1, Python 3.12.13

See the training guide for reproduction. Set --max-seq-length 16384 to match the paper; the script defaults to 12,288.

Evaluation

Table 1 results on the 5,000-example KaliBench test split, in percent:

Mode Exact match Tool accuracy Optional F1 Positional F1 Total Score
Unrestricted 30.6 78.1 52.8 66.9 65.9
Restricted 34.4 95.9 55.2 68.4 73.2
Hinted 77.1 97.1 90.5 92.0 93.2

Average Total Score: 77.4%, up from 71.7% for RedSage-Ins (+5.7 percentage points).

Total Score averages tool accuracy, optional-argument F1, and positional-argument F1. Exact match uses canonicalization and alias-aware scoring. Evaluation uses vLLM in BF16, temperature 0.2, and a 16,384-token budget (8,192 input + 8,192 output), without thinking. See the evaluation guide for the full protocol.

Usage

pip install "transformers==4.57.6" accelerate torch safetensors

Load the merged checkpoint directly. Before publication, replace the model ID with its local directory.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "RISys-Lab/RedSage-K-SFT"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.float16, device_map="auto"
).eval()

messages = [
    {"role": "system", "content": (
        "Translate the request into a single accurate Kali/Linux command. "
        "Return only <output>command</output>."
    )},
    {"role": "user", "content": "In list mode, display the privileges of user 'eve' as they would apply to the command 'cat /etc/shadow', using non-interactive mode."},
]
inputs = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs, max_new_tokens=256, do_sample=False, use_cache=True,
        temperature=0.2,
        pad_token_id=tokenizer.pad_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )
completion = outputs[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(completion, skip_special_tokens=True))
# It should return: "<output>sudo --list --user eve --command 'cat /etc/shadow' --non-interactive</output>"

Precision: Stored tensors are BF16, while config.json declares FP16. The example explicitly loads FP16; use torch.bfloat16 on compatible hardware to match the paper's evaluation precision.

Intended use and limitations

Designed for cybersecurity research, education, and command assistance in authorized environments.

  • Commands may contain incorrect tools, flags, or arguments. Review them before execution; behavior also depends on tool versions and the local environment.
  • KaliBench measures single-command generation, not execution success or multi-step agent performance. Verification can accept environment-related runtime failures and timeouts.
  • Synthetic labels may contain errors, and alias-aware scoring may miss valid alternatives. Training and test sets share tools.
  • Results do not establish multilingual, long-context, general-chat, or misuse-resistance performance.

Citation

If you use RedSage-K-SFT, please cite:

@inproceedings{li2026kalibench,
  title={KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards},
  author={Pengfei Li and Naufal Suryanto and Sicheng Zhang and Muzammal Naseer},
  booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
  year={2026},
  url={https://openreview.net/forum?id=BUajyUxKK6}
}
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RISys-Lab/RedSage-K-SFT

Finetuned
(3)
this model
Finetunes
1 model

Dataset used to train RISys-Lab/RedSage-K-SFT

Collection including RISys-Lab/RedSage-K-SFT

Evaluation results