|
Download README.md from pp1618/Blue-Eye: direct link, hf CLI and curl.
- Browser
- Download file 8.35 kB
-
https://huggingface.co/pp1618/Blue-Eye/resolve/main/README.md
- Command line
-
hf download hf://pp1618/Blue-Eye/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/pp1618/Blue-Eye/resolve/main/README.md
8.35 kB
| license: other | |
| license_name: dinov3-license | |
| license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license | |
| pipeline_tag: image-classification | |
| library_name: pytorch | |
| base_model: facebook/dinov3-vitl16-pretrain-lvd1689m | |
| base_model_relation: finetune | |
| tags: | |
| - image-classification | |
| - content-moderation | |
| - nsfw-detection | |
| - dinov3 | |
| metrics: | |
| - accuracy | |
| extra_gated_prompt: >- | |
| Blue-Eye is a content-moderation model. By downloading it you agree to use it under the DINOv3 | |
| License and not to use it to monitor or make decisions about individual people. | |
| extra_gated_button_content: Agree and download | |
| <p align="center"> | |
| <img src="assets/logo.png" width="140" alt="Blue-Eye logo"> | |
| </p> | |
| # Model Card for Blue-Eye | |
| Blue-Eye is an image classifier for content moderation. Given an image, it predicts whether the content | |
| is **safe**, **suggestive** or **explicit**, together with a probability for each class. On a benchmark | |
| of 3,000 real photographs it reaches **88.9% accuracy**, ahead of AWS Rekognition, Gemini 3.1 Pro, | |
| Google Cloud Vision and the open-source NSFW detectors it was compared with. It handles both photographs | |
| and anime/illustration. | |
|  | |
| ## Model Details | |
| Blue-Eye is a DINOv3 ViT-L/16 vision transformer fine-tuned end to end for three-class content | |
| classification. The model takes a 512x512 RGB image and returns three class probabilities. | |
| ### Model Description | |
| - **Developed by:** Pranshu Patel | |
| - **Model type:** Vision Transformer image classifier | |
| - **Fine-tuned from:** [facebook/dinov3-vitl16-pretrain-lvd1689m](https://huggingface.co/facebook/dinov3-vitl16-pretrain-lvd1689m) | |
| - **Classes:** `safe` (0), `suggestive` (1), `explicit` (2) | |
| - **Parameters:** 303M | |
| - **License:** [DINOv3 License](https://ai.meta.com/resources/models-and-libraries/dinov3-license) | |
| ### Model Sources | |
| - **Repository:** https://github.com/prnshu-p/Blue-Eye | |
| ## Uses | |
| ### Direct Use | |
| Blue-Eye is built for moderating sexual content in images: | |
| - filtering explicit or suggestive images out of feeds, search results and timelines | |
| - blurring images or adding content warnings | |
| - age-gating content on platforms that allow adult material | |
| - prioritising images for human moderators | |
| - curating image datasets before training other models | |
| The class definitions follow a nudity and sexual-content rubric: | |
| - **safe:** no sexualised content, including swimwear, fitness, medical images, breastfeeding and | |
| non-sexual art | |
| - **suggestive:** sexualised but not explicit, such as posed lingerie shots or bare buttocks | |
| - **explicit:** exposed genitalia, sexual acts or full nudity | |
| ### Downstream Use | |
| The default prediction is the highest-probability class. Platforms with a stricter or looser policy can | |
| set their own threshold on `p(explicit)` or `p(suggestive) + p(explicit)`, tuned on their own data. The | |
| model can also be fine-tuned further on a platform's own labels. | |
| ### Out-of-Scope Use | |
| Blue-Eye covers sexual content only; violence, gore and other policy areas are out of scope. It does not | |
| estimate age and is not a tool for detecting child sexual abuse material, for which dedicated | |
| hash-matching services should be used. It should not be used to monitor or make decisions about | |
| individual people. | |
| ## Bias, Risks, and Limitations | |
| The boundary between suggestive and its neighbouring classes is the most subjective part of the task, | |
| for people and models alike, and it is where most disagreements occur. Performance across demographic | |
| groups was not part of this evaluation. | |
| ### Recommendations | |
| Validate the model on images representative of your own platform before deployment, and keep a human | |
| review step for actions that affect user accounts. | |
| ## How to Get Started with the Model | |
| ```bash | |
| pip install torch "transformers>=4.56" safetensors pillow numpy huggingface_hub | |
| ``` | |
| ```python | |
| import sys | |
| from huggingface_hub import snapshot_download | |
| path = snapshot_download("pp1618/Blue-Eye") | |
| sys.path.insert(0, path) | |
| from inference import classify | |
| for result in classify(["photo.jpg", "drawing.png"], model=path): | |
| print(result["label"], result["probabilities"]) | |
| ``` | |
| From the command line: | |
| ```bash | |
| python inference.py photo.jpg folder_of_images/ --model pp1618/Blue-Eye --device cuda | |
| ``` | |
| On GPUs with bfloat16 support, add `--precision bf16` (or `precision="bf16"` in Python) for faster | |
| inference with practically identical predictions. | |
| ## Training Details | |
| ### Training Data | |
| About 2.6 million web images covering real photographs and anime/illustration, labelled into the three | |
| classes using a commercial content-moderation service, source content ratings and model-assisted | |
| relabelling. Evaluation images were removed from the training data. | |
| ### Training Procedure | |
| Training ran in three progressive fine-tuning stages starting from the DINOv3 ViT-L/16 checkpoint: | |
| 1. about 1.0M real photographs | |
| 2. about 1.0M images combining anime/illustration with real photographs | |
| 3. about 650k class-balanced images | |
| **Training regime:** 2 epochs per stage at 512x512, AdamW with a one-cycle schedule, label smoothing | |
| 0.05, random resized crops, horizontal flips and colour jitter, float32 master weights with bfloat16 | |
| mixed precision. | |
| ## Evaluation | |
| ### Testing Data and Metrics | |
| The Blue-Eye benchmark contains 3,000 real photographs (1,480 safe, 513 suggestive, 1,007 explicit), | |
| including 525 non-sexual, skin-heavy images such as swimwear, fitness and medical photos. Every system | |
| below was evaluated on the same images with the same three-class labels; commercial services were | |
| queried in August 2026 and their outputs mapped to the three classes. Google Cloud Vision counts an | |
| image as explicit when `adult` is VERY_LIKELY and as suggestive when `racy` is VERY_LIKELY; AWS | |
| Rekognition uses its default 50% confidence. The metric is three-class accuracy. | |
| ### Results | |
| | System | Type | Accuracy | | |
| |---|---|---:| | |
| | **Blue-Eye** | open weights | **88.9%** | | |
| | AWS Rekognition | commercial API | 87.3% | | |
| | Gemini 3.1 Pro | commercial model | 86.4% | | |
| | Gemini 3.7 Flash | commercial model | 84.1% | | |
| | Google Cloud Vision SafeSearch | commercial API | 83.5% | | |
| | TostAI/nsfw-image-detection-large | open weights | 79.2% | | |
| | Marqo/nsfw-image-detection-384 | open weights | 73.8% | | |
| | Falconsai/nsfw_image_detection | open weights | 70.9% | | |
| | NudeNet | open weights | 70.4% | | |
| | Freepik/nsfw_image_detector | open weights | 67.6% | | |
| | AdamCodd/vit-base-nsfw-detector | open weights | 62.1% | | |
| Many open-source detectors are binary, so they were also compared on the two binary tasks: | |
| | System | Safe vs not safe | Explicit vs rest | | |
| |---|---:|---:| | |
| | **Blue-Eye** | **92.2%** | **95.1%** | | |
| | Marqo/nsfw-image-detection-384 | 86.4% | 78.3% | | |
| | TostAI/nsfw-image-detection-large | 85.7% | 90.1% | | |
| | Falconsai/nsfw_image_detection | 83.3% | 75.6% | | |
| | Freepik/nsfw_image_detector | 82.6% | 71.8% | | |
| | AdamCodd/vit-base-nsfw-detector | 78.6% | 62.7% | | |
| | NudeNet | 76.9% | 87.5% | | |
| Blue-Eye results across domains: | |
| | Evaluation set | Images | Accuracy | | |
| |---|---:|---:| | |
| | Real photographs | 4,000 | 91.4% | | |
| | Anime / illustration | 2,000 | 87.2% | | |
| | All | 6,000 | 90.0% | | |
| Per-class recall on the 3,000 real photographs: safe 89.5%, suggestive 75.4%, explicit 94.8%. | |
| ## Technical Specifications | |
| ### Model Architecture | |
| - **Backbone:** DINOv3 ViT-L/16: 24 layers, embedding dimension 1024, 16 heads, 4 register tokens, RoPE | |
| - **Pooling:** class token concatenated with the mean of the remaining output tokens (2048 features) | |
| - **Head:** LayerNorm followed by a linear layer to 3 classes | |
| - **Input:** RGB, shorter edge resized to 537 (bicubic), centre crop to 512x512, ImageNet normalisation | |
| - **Weights:** float32 safetensors; the head always runs in float32, including under bfloat16 inference | |
| ### Compute Infrastructure | |
| - **Hardware:** Google Cloud TPU v5e | |
| - **Software:** PyTorch, Hugging Face Transformers | |
| ## License | |
| Blue-Eye is a derivative of DINOv3 and is released under the | |
| [DINOv3 License](https://ai.meta.com/resources/models-and-libraries/dinov3-license). The full text is in | |
| `LICENSE`. Commercial use is permitted under its terms. | |
| ## Citation | |
| **BibTeX** | |
| ```bibtex | |
| @misc{patel2026blueeye, | |
| title = {Blue-Eye: a content-safety image classifier}, | |
| author = {Patel, Pranshu}, | |
| year = {2026}, | |
| url = {https://huggingface.co/pp1618/Blue-Eye} | |
| } | |
| ``` | |