OneDecision-VisionGuard-9B-SFT

OneDecision-VisionGuard-9B-SFT

OneDecision-VisionGuard-9B-SFT is a dense 9-billion-parameter multimodal image classification model based on Qwen/Qwen3.5-9B and trained on the ImageShield-OneDecision-Classification content-safety guardrail dataset. The model is designed to classify visual content as Safe or NSFW, with a particular focus on detecting Not Safe for Work (NSFW) sensual content and other potentially sensitive visual content. OneDecision-VisionGuard-9B-SFT performs detailed visual analysis of dress codes, clothing exposure, poses, framing, and visual settings to support conservative content-safety classification.

This model is intended for research and content-safety classification only. Any outputs generated by this model are the sole responsibility of the user. The authors and hosting platform disclaim all liability for generated content. Users must ensure safe, ethical, and lawful usage.

Key Highlights

  • Qwen 3.5 Foundation: Built on Qwen/Qwen3.5-9B.
  • Dense 9B Architecture: A dense 9-billion-parameter multimodal model.
  • Multimodal Image Classification: Designed for visual content-safety classification and image analysis.
  • ImageShield-OneDecision-Classification Training: Fine-tuned on the prithivMLmods/ImageShield-OneDecision-Classification content-safety guardrail dataset.
  • Safe / NSFW Classification: Classifies visual content as either Safe or NSFW.
  • NSFW Sensual Content Detection: Focused on identifying potentially Not Safe for Work (NSFW) sensual content and related sensitive visual material.
  • Explicit & Sensitive Content Detection: Designed to identify explicit, revealing, provocative, and otherwise potentially sensitive visual content.
  • Dress-Code Analysis: Performs detailed analysis of clothing, coverage, exposure, and revealing garments.
  • Pose Analysis: Analyzes body poses, positioning, framing, and potentially suggestive presentation.
  • Context & Setting Analysis: Considers visual settings and contextual cues when determining content safety.
  • Zero-Tolerance Safety Classification: Uses a conservative classification policy that favors safety when content is ambiguous.
  • Structured Output: Produces a machine-readable JSON response containing a scene description, classification, and reasoning.

Standard Prompt

You are a strict content-safety guardrail classifier for images designed for parental control and minor protection.

Your task is to analyze the image thoroughly, provide a detailed visual description, and classify whether the image is safe or unsafe (NSFW).

Enforce a strict zero-tolerance threshold:
- Any form of nudity, exposed breasts, buttocks, or genitalia must be classified as NSFW.
- Partial nudity, revealing underwear, lingerie, suggestive swimwear, or see-through clothing must be classified as NSFW.
- Anime nudes, ecchi, hentai, 2D/3D adult illustrations, suggestive manga art, and sexualized poses must be classified as NSFW.
- Explicit sexual acts, fetish material, or overt sexual positioning must be classified as NSFW.
- When in doubt, classify as NSFW.

Output strictly valid JSON with no extra conversational text or markdown code blocks outside the JSON format. Use exactly these keys in this order:
{
  "scene_description": "<Detailed, objective description of the subject, clothing, exposure, and pose>",
  "is_nsfw": true | false,
  "classification_reason": "<Precise reason for the classification based on clothing, exposure, or pose>",
  "nsfw": 1 | 0,
  "safe": 1 | 0
}

Field constraints:
- "is_nsfw" must be a JSON boolean (true or false), never a string.
- "nsfw" and "safe" must be JSON integers (1 or 0) and complementary: if is_nsfw is true then nsfw = 1 and safe = 0; if is_nsfw is false then nsfw = 0 and safe = 1.

If you are using the transformers inference, the Standard Prompt is optional, you can use simple prompts like "Classify this image as safe or nsfw." to get the classification. The Standard Prompt is used in the GGUF Quantizations and others.

Quick Start with Transformers

pip install transformers
pip install accelerate
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
import torch

# default: Load the model on the available device(s)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "prithivMLmods/OneDecision-VisionGuard-9B-SFT",
    dtype="auto",
    device_map="auto"
)

# We recommend enabling flash_attention_2 for better acceleration and memory saving, especially in multi-image scenarios.
# model = Qwen3_5ForConditionalGeneration.from_pretrained(
#     "prithivMLmods/OneDecision-VisionGuard-9B-SFT",
#     dtype=torch.bfloat16,
#     attn_implementation="flash_attention_2",
#     device_map="auto",
# )

processor = AutoProcessor.from_pretrained("prithivMLmods/OneDecision-VisionGuard-9B-SFT")

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "https://path/to/your/image.jpg",
            },
            {
                "type": "text",
                "text": "Classify this image as safe or nsfw."
            },
        ],
    }
]

# Preparation for inference
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt"
)
inputs = inputs.to(model.device)

# Inference: Generation of the output
generated_ids = model.generate(**inputs, max_new_tokens=256)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text[0])

Sample Outputs

{"scene_description": "Two young Rottweiler puppies are playfully wrestling on a grassy lawn scattered with fallen leaves. The puppy on top has its mouth open, biting down gently on the other's head, while the second puppy lies on its back with paws raised in the air.", "is_nsfw": false, "classification_reason": "The image depicts two playful dogs engaging in natural canine play behavior. There is no evidence of nudity, revealing swimwear, suggestive poses, or any other elements that would indicate an NSFW classification.", "nsfw": 0, "safe": 1}
{"scene_description": "A middle-aged man with a beard stands in a kitchen, wearing a white t-shirt and grey pants. He is actively cooking, holding a frying pan over a gas stove while another pan sits on the burner containing food. The room has warm lighting coming from above, highlighting the cooking activity.", "is_nsfw": false, "classification_reason": "The image depicts a standard domestic scene focused on cooking. The subject is dressed in casual, everyday attire (t-shirt and trousers) suitable for a home environment. There is no evidence of nudity, revealing swimwear, suggestive poses, or any other content that would trigger the NSFW classification criteria.", "nsfw": 0, "safe": 1}
{"scene_description": "A blonde woman wearing sunglasses and a patterned skirt stands outdoors in a park setting, embracing another person whose upper body is largely exposed. The woman's attire includes pink bikini bottoms and a sheer, semi-transparent top that reveals her midriff and bare chest.", "is_nsfw": true, "classification_reason": "The image contains significant exposure and suggestive elements. The central figure is partially nude, showing bare skin on the torso and arms. She wears a sheer, see-through top that exposes her midriff and cleavage, along with pink bikini bottoms. The outdoor setting combined with the intimate embrace and lack of full coverage classifies this as NSFW content.", "nsfw": 1, "safe": 0}

Intended Use

  • Multimodal Image Classification: Classifying visual media as safe or NSFW.
  • Content Safety Classification: Supporting automated visual content-safety classification.
  • NSFW Sensual Content Detection: Supporting research into automated detection of potentially sensitive and sexually suggestive visual content.
  • Parental Controls: Building conservative visual content-safety filtering systems.
  • Content Moderation: Supporting automated safety classification pipelines.
  • Multimodal Safety Research: Evaluating content-safety behavior in multimodal language models.
  • Guardrail Development: Researching structured safety classification and filtering workflows.
  • Visual Safety Analysis: Analyzing clothing, dress codes, poses, framing, and visual settings for content-safety signals.

Limitations

  • Experimental Model: The model may produce incorrect or inconsistent classifications.
  • False Positives: Benign content may occasionally be classified as unsafe due to the conservative classification threshold.
  • False Negatives: Unsafe content may occasionally be missed.
  • Context Sensitivity: Classification performance depends on image quality, visual context, and the provided instruction.
  • Conservative Policy: The model intentionally uses a strict classification threshold and may flag content that would not be considered NSFW under less restrictive moderation policies.
  • Visual Ambiguity: Clothing, poses, artistic styles, and contextual cues may sometimes be difficult to interpret accurately.
  • Computational Requirements: As a dense 9-billion-parameter multimodal model, OneDecision-VisionGuard-9B-SFT requires notable computational resources for inference, though less than larger VisionGuardrail variants.
  • Automated Classification: The model should not be treated as a definitive legal or safety determination.

Model Variants

License and Attributions

  • Qwen/Qwen3.5-9B: Base multimodal model used for this project.

  • ImageShield-OneDecision-Classification: A multimodal classification dataset in single-turn format targeted at one-pass decision-making for unsafe, sensitive, and policy-violating imagery. Each entry pairs a visual input with a serialized JSON string in the Decision column, delivering scene narration, classification rationale, boolean flags, and numeric class indicators. Used to train OneDecision-VisionGuard-9B-SFT.

  • Initial training: strangeropshf/job-term-10_04-M

  • TRL – Transformers Reinforcement Learning: TRL is a full-stack library providing tools to train transformer language models with methods including Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), Direct Preference Optimization (DPO), Reward Modeling, and more.

  • Transformers: Transformers provides state-of-the-art machine learning models for text, computer vision, audio, video, and multimodal tasks, supporting both inference and training.

  • Developed by: prithivMLmods

  • License: This model follows the same Apache 2.0 license as the base model Qwen/Qwen3.5-9B.

Downloads last month
4
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/OneDecision-VisionGuard-9B-SFT

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(1027)
this model
Quantizations
2 models

Dataset used to train prithivMLmods/OneDecision-VisionGuard-9B-SFT

Collections including prithivMLmods/OneDecision-VisionGuard-9B-SFT