PolicyShiftGuard-3B / README.md
hitsmy's picture
Add pipeline tag, library name, and paper/code links (#1)
c8ba92e
|
Raw
History Blame Contribute Delete
1.61 kB
metadata
base_model: Qwen/Qwen2.5-VL-3B-Instruct
datasets:
  - PolicyShiftBench/PolicyShiftBench
license: apache-2.0
pipeline_tag: image-text-to-text
library_name: transformers
tags:
  - vision-language
  - image-safety
  - guardrails
  - policy-conditioned
  - qwen2.5-vl

PolicyShiftGuard-3B

📚 Paper | 💻 GitHub | 🏠 Project Page

PolicyShiftGuard-3B is a policy-conditioned image guardrail model based on Qwen2.5-VL-3B. It is trained to decide whether an image violates a supplied policy bundle and to return a structured safe/unsafe decision with the violated risk category when applicable.

Expected Output Format

true | <two-digit risk category id> | <short reason>
false | <short reason>

Training Data

This checkpoint is trained with the PolicyShiftBench public data release:

  • Dataset: PolicyShiftBench/PolicyShiftBench
  • Main evaluation splits: ID/adaptive branch and OOD/shift branch
  • Training stages: randomized policy SFT followed by boundary-pair policy adaptation

Intended Use

Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.

Limitations

This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.