YOLO Head Detection

A collection of seven YOLO checkpoints trained to detect human heads in images β€” useful for crowd counting, privacy blurring, and as a first stage before face recognition.

Trained on a dataset of 15,000 images.

Which one should you use?

  • Fastest on a small model β†’ v8n-head.pt (93 FPS, 6.2 MB).
  • Best balance β†’ v8s-head.pt (90 FPS, 22.5 MB, and zero false boxes on our negative test).
  • Highest recall β†’ 11m-head.pt / v8m-head.pt (more detections, ~50 FPS).

Note the pattern: every YOLOv8 checkpoint we shipped stays silent on a head-free image, while the YOLO11 family tends to emit low-confidence false boxes. If false positives cost you more than missed heads, start with v8s-head.pt.

πŸš€ Quick start

pip install ultralytics
from ultralytics import YOLO

# pick any checkpoint from the table above
model = YOLO("v8s-head.pt")

results = model("photo.jpg", conf=0.25, imgsz=640)

for box in results[0].boxes:
    x1, y1, x2, y2 = box.xyxy[0].tolist()
    conf = float(box.conf)
    print(f"head @ ({x1:.0f},{y1:.0f})-({x2:.0f},{y2:.0f})  conf={conf:.2f}")

results[0].save("annotated.jpg")

πŸ’‘ The checkpoints do not carry class names β€” class index 0 is the head class. To get readable labels on plotted boxes, map it yourself:

results[0].names = {0: "head"}   # visualisation only

πŸ” What it looks like

Street scene, 11m-head.pt β€” four people are visible, and two of the heads are boxed:

street scene

Box colour and the head <confidence> label are added by our own drawing code; the checkpoints themselves emit bare xyxy boxes with no class names.

⚠️ Failure cases β€” read this before deploying

1. Missed heads in clutter. In the street scene above, four people are visible but only two heads were boxed: the person seen from behind and the person clipped at the frame edge were missed entirely.

2. False positives on head-free images. Given a photo of two zebras β€” no humans at all β€” most checkpoints still emit boxes: 11s-relu-head.pt produced 4 false boxes, 11m-head.pt 3, while every YOLOv8 checkpoint stayed clean.

Low-confidence detections were also wrong on a tennis photo: two of four boxes landed on a player's shoulder and back instead of a head (confidences 0.78 and 0.49).

Practical advice: raise conf well above the default 0.25 for production β€” we used the low Ultralytics default deliberately, to expose these failures β€” and test against your own head-free negatives before trusting the model.

πŸ“Š How the numbers were measured

Hardware 1Γ— Tesla T4 (Kaggle), fp16, batch size 1
Latency cuda.synchronize() β†’ perf_counter β†’ cuda.synchronize(), 5 warm-up runs, median of 3 repetitions per image
Inference settings imgsz=640, conf=0.25, no TTA, no batching
Speed images 4 photos that genuinely contain human heads
Test images 2 Γ— Ultralytics sample images (bus.jpg, zidane.jpg) + 2 Γ— COCO val2017 images
Negative image 1 photo containing no humans (coco_1818.jpg, two zebras), used to expose false positives
Measured on 2026-10-04

⚠️ No mAP is reported. The test split used for training was not published alongside these checkpoints, so no ground truth is available to us β€” and we will not invent a number. Everything in this card is either a measured latency, a raw detection count, or a human-verified qualitative observation.

πŸŽ“ Training

  • Dataset size: ~15,000 images
  • Task: single-class object detection (human heads)
  • Label convention: class index 0; the checkpoints include no names mapping
  • Framework: Ultralytics β€” YOLOv8 and YOLO11

πŸ“„ License

Released under AGPL-3.0, inherited from the Ultralytics framework.

Citation

@misc{yolo-head-detection,
  title = {YOLO Head Detection},
  note  = {Seven YOLOv8 / YOLO11 head-detection checkpoints trained on ~15k images},
  year  = {2026}
}
Downloads last month
184
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support