Instructions to use BananaMind/BananaMind-2-Remove-Small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ultralytics
How to use BananaMind/BananaMind-2-Remove-Small with ultralytics:
from huggingface_hub import hf_hub_download from ultralytics import YOLO # pick the weights file from this repo's "Files and versions" tab weights = hf_hub_download("BananaMind/BananaMind-2-Remove-Small", "<weights>.pt") model = YOLO(weights) source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - Notebooks
- Google Colab
- Kaggle
BananaMind-2-Remove-Small
BananaMind-2-Remove-Small is the larger sibling of BananaMind-2-Remove-Nano: a single-class face detector for blurring or blacking out faces in video. It is YOLO11s fine-tuned from the COCO-pretrained weights on WIDER FACE and finds more of the small faces than the Nano model.
The model has 9,413,187 parameters (21.4 GFLOPs) and a 640x640 input. It ships as PyTorch, ONNX FP32 and ONNX FP16.
Model Details
| Field | Value |
|---|---|
| Parameters | 9,413,187 (fused) |
| Architecture | YOLO11s, anchor-free decoupled detection head |
| Layers (fused) | 100 |
| GFLOPs @ 640 | 21.4 |
| Starting weights | COCO-pretrained yolo11s.pt (fine-tuned, not trained from scratch) |
| Classes | 1 (face) |
| Input size | 640x640 |
| Weight formats | PyTorch .pt, ONNX FP32, ONNX FP16 |
| Library | Ultralytics 8.4.172 |
Training Data
| Split | Images | Faces |
|---|---|---|
| WIDER FACE train (converted) | 12,876 | 156,994 |
| WIDER FACE val (converted) | 3,222 | 39,112 |
Converted from CUHK-CSE/wider_face to one-class YOLO format. Boxes flagged invalid were skipped, small faces were kept, images whose only boxes were invalid were dropped
(8 images), and images were downscaled to 640 px on the long side.
Training Setup
| Field | Value |
|---|---|
| Hardware | 1x RTX 5070 Ti (16 GB) |
| Epochs | 40 (see note) |
| Image size | 640 |
| Batch size | 32 |
| Precision | AMP |
| Optimizer | AdamW (Ultralytics auto, lr 0.002) |
| Close mosaic | last 5 epochs |
| Augmentation | Ultralytics defaults + random motion blur (p 0.3, 3 to 15 px kernel) + JPEG compression (p 0.3, quality 20 to 90) to imitate video frames |
| Wall-clock time | about 30 minutes |
Note: the run was started as a 60-epoch schedule, stopped after epoch 34, and resumed from that checkpoint with the schedule shortened to end at epoch 40 (the learning-rate decay and the mosaic switch-off were recomputed for the new total). Epochs 1 to 34 therefore used the 60-epoch learning-rate curve.
Benchmarks
All models were run through the same evaluator on the same ground truth: the WIDER FACE validation split converted to YOLO format
(3,222 images, 39,112 labelled faces, images stored at 640 px or smaller, invalid boxes removed, every face size kept). mAP is COCO-style
(101-point AP via ultralytics.utils.metrics.ap_per_class, greedy score-ordered matching, IoU 0.5 and 0.5 to 0.95). "Recall" is the share of
labelled faces that receive a box with score >= 0.25 at IoU >= 0.5, which is the setting the blur script uses. Face height is measured in pixels
of the stored image. Speed is ONNX Runtime, batch 1, a random 640x640 input: "GPU" is an RTX 5070 Ti (CUDA provider, FP16 files for BananaMind models,
FP32 for the others), "CPU" is 4 threads (FP32 files for BananaMind FP32/FP16 rows, the INT8 file for the INT8 row). Timings exclude image decoding and NMS.
* SCRFD-10GF is det_10g.onnx from InsightFace's buffalo_l pack and YuNet is face_detection_yunet_2023mar.onnx from the OpenCV Zoo. Both were run with a
score floor of 0.02 and their standard decoding/NMS. They are different architectures trained on different data (not retrained here), so the comparison is
indicative, not a controlled study. SCRFD-10GF is the better detector for faces of 20 px and up; YuNet is far smaller and faster but much less accurate.
| Model | mAP50 | mAP50-95 | Recall (all) | Recall (>=20 px) | Recall (<10 px) | GPU ms | CPU ms | File MB |
|---|---|---|---|---|---|---|---|---|
| BananaMind-2-Remove-Small | 71.77% | 39.26% | 67.9% | 92.5% | 42.1% | 1.75 | 49.5 | 19.0 |
| BananaMind-2-Remove-Nano | 66.49% | 35.54% | 61.3% | 89.6% | 33.1% | 1.19 | 18.9 | 5.4 |
| BananaMind-2-Remove-Nano INT8 | 64.95% | 34.28% | 60.0% | 88.5% | 31.5% | 3.45 | 12.7 | 3.3 |
| SCRFD-10GF* | 68.66% | 37.18% | 66.8% | 94.1% | 36.3% | 2.48 | 50.5 | 16.9 |
| YuNet* | 47.48% | 22.93% | 50.0% | 86.9% | 13.3% | 0.77 | 2.1 | 0.2 |
Recall by face size
| Face height (px) | Faces | Remove-Nano | Remove-Nano INT8 | Remove-Small | SCRFD-10GF |
|---|---|---|---|---|---|
| under 10 | 16,305 | 33.1% | 31.5% | 42.1% | 36.3% |
| 10 to 20 | 10,521 | 72.1% | 70.7% | 79.0% | 82.2% |
| 20 and up | 12,286 | 89.6% | 88.5% | 92.5% | 94.1% |
| 40 and up | 4,808 | 94.3% | 93.6% | 95.9% | 96.5% |
| All faces | 39,112 | 61.3% | 60.0% | 67.9% | 66.8% |
Ultralytics val reference
Ultralytics' own validator matches predictions differently and reports slightly higher numbers for the same weights (it is not used in the table above).
| Format | mAP50 | mAP50-95 |
|---|---|---|
PyTorch .pt |
74.06% | 40.09% |
| ONNX FP16 | 73.88% | 39.92% |
Privacy notes
Missing a face in a single frame is enough to identify someone, and no detector finds every face: see the recall tables above, especially for tiny (under 20 px), turned-away, covered or motion-blurred faces. Check the output before sharing it.
We also tested our own pixelation (15% padding, blur_video.py logic) against a small learned attacker trained on WIDER FACE crops. At the default 8 blocks
the attacker recovered only coarse traits (head pose, hair and skin colour, sometimes glasses or a beard), not identity; with weaker pixelation (16 blocks)
plain face recognition on the pixelated image alone already matched 36.7% of faces among 1,500 candidates. The attacker was small and briefly trained, so
treat this as a floor on leakage. For strong anonymisation use --mode black.
Limitations
- Trained only on WIDER FACE. Performance on other kinds of footage (night, heavy compression, unusual cameras, non-human faces) is untested.
- Tiny and far-away faces are the weak spot: roughly a third to two fifths of faces under 10 px are found. Faces of 20 px and up are found about 90% of the time.
- Single-image detector. Video behaviour (tracking, hold, interpolation between detections) comes from the script, not the model.
- Not evaluated against adversarial inputs. Face blurring is not a guarantee of anonymity; voice, clothing, text and surroundings are untouched.
- Numbers above are self-reported on the WIDER validation split.
Usage
pip install -U ultralytics onnxruntime-gpu lap huggingface_hub
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
model_id = "BananaMind/BananaMind-2-Remove-Small"
path = hf_hub_download(model_id, "face_yolo11s.pt")
model = YOLO(path)
results = model.predict("photo.jpg", conf=0.25, imgsz=640)
for box in results[0].boxes:
print(box.xyxy[0].tolist(), float(box.conf[0]))
Blur or black out faces in a video
blur_video.py (included in this repo) detects faces, tracks them with ByteTrack, pixelates or blacks them out, and keeps the audio via ffmpeg.
# pixelate (default), detector on every frame
python blur_video.py in.mp4 out.mp4 --model face_yolo11s_fp16.onnx -n 1 --device 0
# solid black boxes
python blur_video.py in.mp4 out.mp4 --model face_yolo11s_fp16.onnx -n 1 --device 0 --mode black
Defaults: detector every 3 frames (-n), confidence 0.25, 15% box padding, 8 pixel blocks across the shorter side of each face (--blocks, lower is
stronger). Use -n 1 for the strictest coverage. With an ONNX model use --device 0; --device cpu makes Ultralytics try to pip-install onnxruntime.
The optional --keep mode (leave one person visible) needs face_id.py and InsightFace models that are not included here.
Files
| File | Notes |
|---|---|
face_yolo11s.pt |
PyTorch weights |
face_yolo11s_fp32.onnx, face_yolo11s_fp16.onnx |
ONNX, fixed 640x640, batch 1 |
blur_video.py |
detect, track and blur/black-out faces in a video, audio preserved |
Intended Use
Local face detection for blurring faces in video and images where the extra recall of a larger model is worth about 3.3x the compute of the Nano model. It is not intended as the sole safeguard where a missed face would cause serious harm.
License
AGPL-3.0, inherited from Ultralytics YOLO11 (the architecture and the COCO-pretrained starting weights). Using these weights or the included script in a product or network service means complying with AGPL-3.0 or obtaining an Ultralytics Enterprise licence.
The training data, WIDER FACE, is listed as CC BY-NC-ND 4.0 / non-commercial research on its Hugging Face card. Do not assume these weights are cleared for commercial use; check the dataset terms before relying on them commercially.
Citation
@inproceedings{yang2016wider,
author = {Yang, Shuo and Luo, Ping and Loy, Chen Change and Tang, Xiaoou},
title = {WIDER FACE: A Face Detection Benchmark},
booktitle = {IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2016}
}
@software{ultralytics_yolo11,
author = {Jocher, Glenn and Qiu, Jing},
title = {Ultralytics YOLO11},
year = {2024},
url = {https://github.com/ultralytics/ultralytics},
license = {AGPL-3.0}
}
🍌
- Downloads last month
- 136
Model tree for BananaMind/BananaMind-2-Remove-Small
Base model
Ultralytics/YOLO11
