Instructions to use RISys-Lab/RedSage-K-SFT-GRPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use RISys-Lab/RedSage-K-SFT-GRPO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="RISys-Lab/RedSage-K-SFT-GRPO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("RISys-Lab/RedSage-K-SFT-GRPO") model = AutoModelForCausalLM.from_pretrained("RISys-Lab/RedSage-K-SFT-GRPO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use RISys-Lab/RedSage-K-SFT-GRPO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "RISys-Lab/RedSage-K-SFT-GRPO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RISys-Lab/RedSage-K-SFT-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/RISys-Lab/RedSage-K-SFT-GRPO
- SGLang
How to use RISys-Lab/RedSage-K-SFT-GRPO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "RISys-Lab/RedSage-K-SFT-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RISys-Lab/RedSage-K-SFT-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "RISys-Lab/RedSage-K-SFT-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RISys-Lab/RedSage-K-SFT-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use RISys-Lab/RedSage-K-SFT-GRPO with Docker Model Runner:
docker model run hf.co/RISys-Lab/RedSage-K-SFT-GRPO
RedSage-K-SFT-GRPO
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
(NeurIPS 2026 Evaluations and Datasets Track)
Authors: Pengfei Li1*, Naufal Suryanto1*, Sicheng Zhang1, Muzammal Naseer1,2
1Khalifa University, 2University of Western Australia
*Equal contribution
📄 arXiv Paper |
🌐 Project Page |
💻 GitHub Code |
🤗 Datasets & Models
Model summary
RedSage-K-SFT-GRPO is an 8B cybersecurity model for translating natural-language requests into Kali/Linux commands. It applies Group Relative Policy Optimization (GRPO) to RedSage-K-SFT, combining supervised fine-tuning with reinforcement learning using verifiable rewards on KaliBench. It corresponds to “RedSage-K (SFT+GRPO)” on the paper and project page.
| Property | Value |
|---|---|
| Developer | RISys-Lab, Khalifa University |
| Architecture | Qwen3ForCausalLM, 36 layers |
| Release format | Merged LoRA weights, BF16 Safetensors |
| Language | English |
| Output format | <think>...</think> followed by <output>command</output> |
Training
KaliBench contains 8,504 verified query-command pairs spanning 1,642 sub-tools and 23 capability dimensions. GRPO uses 3,504 training pairs, each presented in three modes, for 10,512 prompts. The remaining 5,000 pairs form the test split.
| Mode | Model input |
|---|---|
| Unrestricted | Query only |
| Restricted | Query and candidate tools |
| Hinted | Query, target tool, and usage documentation |
GRPO rewards output format, tool selection, optional-argument F1, positional-argument F1, and exact command match. Rewards are computed from command structure without executing generated commands. Prompts request reasoning in <think> tags before the final command. Appendix G.2 reports the following GRPO settings:
| Setting | Value |
|---|---|
| Hardware / duration | One NVIDIA H200 (141 GB) / approximately 17 hours for GRPO, following approximately 32 minutes of SFT |
| Epochs / effective batch size | 2 / 32 (4 per device × 8 accumulation steps) |
| Optimizer / learning rate | 8-bit AdamW / 5e-6 |
| Schedule / warmup / weight decay | Linear / 10% / 1e-3 |
| LoRA rank / alpha / dropout | 64 / 128 / 0 |
| LoRA targets | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Precision / gradient checkpointing | BF16 / enabled |
| Generations per prompt / generation batch size | 8 / 32 |
| Maximum prompt / completion length | 5,549 / 6,739 tokens |
| Loss / clipping / KL coefficient | BNPO-style GRPO / 0.2 / 0.001 |
| Reward scaling | Group-based |
| Generation backend | Colocated vLLM, tensor parallel size 1, GPU memory utilization 0.3 |
See the training guide and GRPO implementation for reproduction.
Evaluation
Table 1 results on the 5,000-example KaliBench test split, in percent:
| Mode | Exact match | Tool accuracy | Optional F1 | Positional F1 | Total Score |
|---|---|---|---|---|---|
| Unrestricted | 32.2 | 77.9 | 56.1 | 76.9 | 70.3 |
| Restricted | 37.4 | 95.0 | 57.9 | 80.9 | 77.9 |
| Hinted | 69.4 | 92.5 | 86.5 | 89.1 | 89.3 |
Average Total Score: 79.2%, up from 71.7% for RedSage-Ins (+7.5 percentage points). Gains are concentrated in unrestricted and restricted modes; hinted Total Score is slightly below the RedSage-Ins baseline.
Total Score averages tool accuracy, optional-argument F1, and positional-argument F1. Exact match uses canonicalization and alias-aware scoring. Evaluation uses vLLM in BF16, temperature 0.2, and a 16,384-token budget (8,192 input + 8,192 output), with thinking enabled. See the evaluation guide for the full protocol.
Usage
pip install transformers accelerate torch safetensors
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "RISys-Lab/RedSage-K-SFT-GRPO"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
).eval()
system_prompt = """You are a cybersecurity function-calling AI model.
You have access to the following tool:
<tools>
[{'type':'function','function':{'name':'run_terminal','description':'Execute a shell command in a Kali/Linux terminal and return stdout, stderr, and exit code.'}}]
</tools>
TASK:
Given a USER QUERY, generate the single most accurate shell command using any appropriate Kali/Linux tool(s) to solve the query.
REQUIREMENTS:
1. Use the correct command-line tool(s) appropriate for the task.
2. Include all required optional arguments (flags beginning with '-' or '--') necessary to accomplish the task.
3. Correctly pair option keys and values (e.g., `--port 80`, `-A INPUT`, or `--flag=value`).
4. Preserve and include any positional arguments (e.g., IPs, filenames, interfaces).
5. Do NOT invent flags/options that do not exist for real Kali/Linux tools.
6. Note: scoring will penalize missing optional arguments, incorrect option->value pairs, or omitted positional arguments.
OUTPUT FORMAT:
<think>
[your_reasoning]
</think>
<output>
[command]
</output>"""
user_template = """USER QUERY: "{query}"
Generate the single most accurate shell command for the query.
Your response must follow the required structure:
<think>
[your_reasoning]
</think>
<output>
[command]
</output>"""
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_template.format(
query="In list mode, display the privileges of user 'eve' as they would apply to the command 'cat /etc/shadow', using non-interactive mode."
)},
]
inputs = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs, max_new_tokens=8192, do_sample=True, use_cache=True,
temperature=0.2, top_p=1.0, top_k=0,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
)
completion = outputs[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(completion, skip_special_tokens=True))
This example uses the exact unrestricted thinking evaluation prompts from src/prompt.py. The bundled chat template supplies the opening <think> tag, so the decoded completion starts after it. Use --output-policy thinking with the released evaluator.
Precision: The example loads the stored BF16 weights on compatible hardware.
Intended use and limitations
Designed for cybersecurity research, education, and command assistance in authorized environments.
- Commands may contain incorrect tools, flags, or arguments. Review them before execution; behavior also depends on tool versions and the local environment.
- KaliBench measures single-command generation, not execution success or multi-step agent performance. Verification can accept environment-related runtime failures and timeouts.
- Synthetic labels may contain errors, and alias-aware scoring may miss valid alternatives. Training and test sets share tools.
- Results do not establish multilingual, long-context, general-chat, or misuse-resistance performance.
Citation
If you use RedSage-K-SFT-GRPO, please cite:
@inproceedings{li2026kalibench,
title={KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards},
author={Pengfei Li and Naufal Suryanto and Sicheng Zhang and Muzammal Naseer},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
year={2026},
url={https://github.com/RISys-Lab/KaliBench}
}
- Downloads last month
- 403
Model tree for RISys-Lab/RedSage-K-SFT-GRPO
Base model
Qwen/Qwen3-8B-BaseDataset used to train RISys-Lab/RedSage-K-SFT-GRPO
Collection including RISys-Lab/RedSage-K-SFT-GRPO
Paper for RISys-Lab/RedSage-K-SFT-GRPO
Evaluation results
- Unrestricted Exact command match (%) on KaliBenchtest set self-reported32.200
- Unrestricted Tool accuracy (%) on KaliBenchtest set self-reported77.900
- Unrestricted Optional-argument F1 (%) on KaliBenchtest set self-reported56.100
- Unrestricted Positional-argument F1 (%) on KaliBenchtest set self-reported76.900
- Unrestricted Total Score (%) on KaliBenchtest set self-reported70.300
- Restricted Exact command match (%) on KaliBenchtest set self-reported37.400
- Restricted Tool accuracy (%) on KaliBenchtest set self-reported95.000
- Restricted Optional-argument F1 (%) on KaliBenchtest set self-reported57.900