VirbiusGuard-4B

VirbiusAgent 安全分类器(Prompt L1 检测),基于 Qwen3Guard-Gen-4B 微调的 LoRA 模型。 输出严格 JSON:hit_rule 与 triggered_id。

同口径评测相对基座:漏检 15.4% 降到 0.8%(gold_1000 / V15),jailbreak 召回 57.1% 升到 100%。

0.6B 轻量版:https://www.modelscope.cn/models/i1see1you/VirbiusGuard

与 Qwen3Guard-Gen-4B 对比

基座用官方 Safety 模板(Unsafe / Controversial 视为拦截);VirbiusGuard-4B 用引擎 JSON 协议。评测集与口径相同。

gold_1000(主表)

评测集:data/eval/gold_1000.jsonl(615 unsafe / 385 safe)。误报 = FP / 385。

模型 acc recall 漏检 FP precision
Qwen3Guard-Gen-4B 87.9% 84.6% 15.4%(95/615) 6.8%(26/385) 95.2%
4B V13.3 93.6% 99.5% 0.5%(3/615) 15.8%(61/385) 90.9%
VirbiusGuard-4B V15 97.5% 99.2% 0.8%(5/615) 5.2%(20/385) 96.8%
4B V17 97.8% 98.0% 2.0%(12/615) 2.6%(10/385) 98.4%

基座漏掉的主要是越狱与 Agent 工具滥用。V13.3 召回拉满但误报过高;V15 起进入可用区。V17 误报最低,但召回/自伤回退。

关键类别召回(gold_1000)

类别 基座 V13.3 V15 V17
Jailbreak(98) 57.1% 100% 100% 98.0%
Agent Tool Misuse(84) 81.0% 100% 98.8% 100%
Suicide and Self-Harm(33) 93.9% 100% 93.9% 93.9%

版本取舍(仅 gold_1000)

  • 要最低误报:V17 (10/385)
  • 要高召回且可用:V15

模型简介

  • 架构:Qwen3ForCausalLM(4B),LoRA(rank 32 / alpha 64)
  • 基座:Qwen3Guard-Gen-4B
  • 版本:V17(当前默认)
  • 相对基座的补强:jailbreak 与 agent-behavior
  • V17 数据:与 0.6B V15 同口径,良性切片再平衡,含 oasst1、COIG 中文散文、OCR 风格文本

分类体系

输出 10 种 unsafe 类别(triggered_id)或 safe(hit_rule 为 false)。每条输入只输出一个主要类别:

Violent、Non-violent Illegal Acts、Unethical Acts、Suicide and Self-Harm、Jailbreak、PII、Copyright Violation、Politically Sensitive Topics、Sexual Content or Sexual Acts、Agent Tool Misuse。

训练数据按 A 口径(提及即违规)标注,Politically Sensitive 拦截较严。

下载

HuggingFace:https://huggingface.co/i1see1you/VirbiusGuard-4B

main 为最新 V17。

  • Transformers:model-00001-of-00005.safetensors 至 model-00005-of-00005.safetensors(fp16,约 7.5GB)
  • GGUF F16:gguf/virbiusguard-4b-v17-f16.gguf(约 7.5GB,Ollama / llama.cpp)
  • GGUF Q8_0:gguf/virbiusguard-4b-v17-q8_0.gguf(约 4.0GB)
  • GGUF Q4_K_M:gguf/virbiusguard-4b-v17-q4_k_m.gguf(约 2.3GB)

使用方式

Transformers:

from_pretrained("i1see1you/VirbiusGuard-4B"),用 tokenizer.apply_chat_template 组 prompt,max_new_tokens 至少 40。系统提示要求只输出 JSON 字段 hit_rule 与 triggered_id。

VirbiusAgent 引擎:替换环境变量 VIRBIUS_PROMPT_LLM_MODEL 即生效。

训练方法(概述)

  • 教师模型离线标注,知识蒸馏
  • mlx-lm LoRA 微调(rank 32 / alpha 64 / dropout 0.1 / lr 1.5e-4 / 2 epoch)
  • 训练集:与 0.6B V15 同份 virbius_v15_train,约 22793 条

许可证 / 归属

基于 Qwen3Guard-Gen-4B 微调,数据集由教师模型离线标注。

联系我们

Downloads last month
249
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for i1see1you/VirbiusGuard-4B

Finetuned
Qwen/Qwen3-4B
Quantized
(11)
this model
Quantizations
1 model