File size: 1,933 Bytes
a4cc862
 
 
e5e41ed
 
 
 
 
 
 
 
 
 
a4cc862
 
 
 
e5e41ed
 
a4cc862
 
 
 
 
 
 
 
 
 
 
9d86d36
a4cc862
9d86d36
a4cc862
 
 
 
 
 
 
 
 
 
e5e41ed
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
---
base_model: Qwen/Qwen2.5-VL-7B-Instruct
datasets:
- PolicyShiftBench/PolicyShiftBench
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- vision-language
- image-safety
- guardrails
- policy-conditioned
- qwen2.5-vl
---

# PolicyShiftGuard-7B

[📜 Paper](https://arxiv.org/abs/2607.05910) | [💻 Code](https://github.com/ssmisya/PolicyShiftGuard) | [🏠 Project Page](https://policyshiftguard.github.io/)

PolicyShiftGuard-7B is a policy-conditioned image guardrail model based on Qwen2.5-VL-7B. It is trained to follow a supplied policy bundle and produce structured image-safety decisions under changing application policies.

## Expected Output Format

```text
true | <two-digit risk category id> | <short reason>
false | <short reason>
```

## Training Data

This checkpoint is trained with the PolicyShiftBench public data release:

- Dataset: `PolicyShiftBench/PolicyShiftBench`
- Main evaluation splits: ID/adaptive branch and OOD/shift branch
- Training stages: randomized policy SFT followed by boundary-pair policy adaptation

## Intended Use

Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.

## Limitations

This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.

## Citation

If you use this model, please cite the paper:

```bibtex
@article{song2026policyshiftguard,
  title   = {PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails},
  author  = {Song, Mingyang and Xu, Luxin and Sun, Haoyu and Pan, Minzhou and Cheng, Yu and Li, Bo},
  journal = {arXiv preprint arXiv:2607.05910},
  year    = {2026}
}
```