YOLO26x Pose Model

Model description

  • Model name: YOLO26x Human Pose Estimation
  • Version: v1
  • Status: experimental
  • Repository visibility: private to the Select AI Hugging Face organization
  • Model type: human detection plus 17-keypoint COCO pose estimation, expanded to a 51-feature per-person vector plus bounding box
  • Upstream trainer/developer: Ultralytics YOLO26
  • Inventory owner: Akshansh D
  • Architecture: YOLO26x-pose backbone with NMS-free end-to-end keypoint regression head
  • Output features: 51 per detected person (17 keypoints Γ— 3) plus box (x1, y1, x2, y2)
  • Reference date: 2026-09-22

This repository contains the model checkpoint and inference artifacts used by the Select AI vision analytics and engagement pipelines for real-time human pose estimation, skeletal tracking, and downstream 51-feature pose vectors plus bounding boxes. It does not contain camera streams, video clips, recorded candidate poses, or downstream analytics databases.

The model is experimental. Initial trainer benchmark evaluations are complete, but independent hold-off validation is still pending.

How to use

The reference path is scripts/run.py. It accepts a single image or video frame, detects human instances, predicts 17 COCO keypoints [x, y, confidence] per person alongside bounding boxes (x1, y1, x2, y2), then expands each person into a fixed 51-feature vector for downstream analytics. Downstream consumers assemble box_xyxy + feature_vector into a pandas DataFrame.

python scripts/run.py \
  --image /path/to/image.jpg \
  --device 0

Optional flags: --model, --imgsz, --conf, --iou, --max-det, --output /path/to/annotated.jpg.

The script outputs JSON to stdout. Each entry in predictions carries the raw box (box_xyxy), the 17 named keypoints, and the assembled 51-feature vector (feature_vector). Pass --output to save an annotated frame with rendered skeletal connections.

Input Contract

  • Input: one decoded color image, typically a BGR uint8 OpenCV array or an image path accepted by scripts/run.py.

  • Camera frames: supports standard CCTV surveillance resolutions up to 1920 Γ— 1080. No hard-coded fixed aspect ratio is required.

  • Preprocessing: resized and letterboxed to 960 Γ— 960, producing a network tensor of shape [1, 3, 960, 960].

  • Color order and normalization: scripts/run.py receives BGR images, converts them to RGB, and normalizes pixel intensities to the [0, 1] float range.

  • Detector confidence threshold: 0.25 by default (--conf).

  • Keypoint confidence threshold: 0.50 by default for valid keypoint filtering.

  • IoU threshold: 0.50 by default (--iou).

The pose model detects body keypoints independently per frame. Downstream consumers assemble box_xyxy + feature_vector into a pandas DataFrame. Temporal tracking (e.g., ByteTRACK or NVDCF) or other modules must consume that DataFrame to build multi-frame trajectories or engagement events.

Output Contract

For each detected person, scripts/run.py returns a bounding box (x1, y1, x2, y2) and a fixed 51-feature vector. Consumers can assemble these into a pandas DataFrame (helper: build_feature_frame).

CCTV IMAGE
    β”‚
    β–Ό
YOLO26x-Pose
    β”‚
    β”œβ”€β”€ Person bounding box (x1, y1, x2, y2)
    β”‚
    └── 17 pose keypoints
          β”‚
          β”œβ”€β”€ x
          β”œβ”€β”€ y
          └── confidence
                 β”‚
                 β–Ό
        51 features (17 keypoints Γ— 3) + bbox
                 β”‚
                 β–Ό
        pandas DataFrame

Feature vector layout (51 floats)

Total = 51 keypoint features (17 keypoints Γ— 3).

Index range Count Block Contents
0–50 51 Keypoint features For each COCO keypoint i (0–16, in order): x, y, confidence

Bounding box is returned separately as box_xyxy = [x1, y1, x2, y2] (pixel coordinates on the original frame).

Keypoint feature index map:

Index range Keypoint id Fields
0–2 nose 0 x, y, confidence
3–5 left_eye 1 x, y, confidence
6–8 right_eye 2 x, y, confidence
9–11 left_ear 3 x, y, confidence
12–14 right_ear 4 x, y, confidence
15–17 left_shoulder 5 x, y, confidence
18–20 right_shoulder 6 x, y, confidence
21–23 left_elbow 7 x, y, confidence
24–26 right_elbow 8 x, y, confidence
27–29 left_wrist 9 x, y, confidence
30–32 right_wrist 10 x, y, confidence
33–35 left_hip 11 x, y, confidence
36–38 right_hip 12 x, y, confidence
39–41 left_knee 13 x, y, confidence
42–44 right_knee 14 x, y, confidence
45–47 left_ankle 15 x, y, confidence
48–50 right_ankle 16 x, y, confidence

JSON payload

For each detected person, scripts/run.py returns:

{
  "box_xyxy": [215.4, 102.1, 480.8, 610.5],
  "detection_score": 0.91,
  "class_id": 0,
  "class_name": "person",
  "keypoints": [
    {"name": "nose", "id": 0, "xy": [340.1, 145.2], "confidence": 0.95},
    {"name": "left_eye", "id": 1, "xy": [348.0, 140.1], "confidence": 0.94},
    {"name": "right_eye", "id": 2, "xy": [332.2, 140.5], "confidence": 0.92},
    {"name": "left_ear", "id": 3, "xy": [360.5, 148.0], "confidence": 0.88},
    {"name": "right_ear", "id": 4, "xy": [320.1, 149.1], "confidence": 0.85},
    {"name": "left_shoulder", "id": 5, "xy": [385.2, 210.4], "confidence": 0.91},
    {"name": "right_shoulder", "id": 6, "xy": [295.1, 212.0], "confidence": 0.89},
    {"name": "left_elbow", "id": 7, "xy": [410.0, 290.1], "confidence": 0.84},
    {"name": "right_elbow", "id": 8, "xy": [270.4, 292.5], "confidence": 0.82},
    {"name": "left_wrist", "id": 9, "xy": [425.1, 360.0], "confidence": 0.79},
    {"name": "right_wrist", "id": 10, "xy": [255.0, 362.1], "confidence": 0.76},
    {"name": "left_hip", "id": 11, "xy": [370.2, 380.0], "confidence": 0.88},
    {"name": "right_hip", "id": 12, "xy": [310.4, 381.2], "confidence": 0.87},
    {"name": "left_knee", "id": 13, "xy": [380.1, 490.5], "confidence": 0.85},
    {"name": "right_knee", "id": 14, "xy": [305.0, 492.1], "confidence": 0.83},
    {"name": "left_ankle", "id": 15, "xy": [390.0, 585.4], "confidence": 0.81},
    {"name": "right_ankle", "id": 16, "xy": [298.2, 588.0], "confidence": 0.78}
  ],
  "feature_count": 51,
  "feature_vector": [
    340.1, 145.2, 0.95,
    348.0, 140.1, 0.94,
    332.2, 140.5, 0.92,
    360.5, 148.0, 0.88,
    320.1, 149.1, 0.85,
    385.2, 210.4, 0.91,
    295.1, 212.0, 0.89,
    410.0, 290.1, 0.84,
    270.4, 292.5, 0.82,
    425.1, 360.0, 0.79,
    255.0, 362.1, 0.76,
    370.2, 380.0, 0.88,
    310.4, 381.2, 0.87,
    380.1, 490.5, 0.85,
    305.0, 492.1, 0.83,
    390.0, 585.4, 0.81,
    298.2, 588.0, 0.78
  ]
}

box_xyxy, keypoints.xy, and keypoint features (indices 0–50) use unnormalized pixel coordinates mapped directly to the original input image frame. The payload also carries feature_count (51); consumers that need a flat tensor should read feature_vector only, and combine it with box_xyxy when building a pandas DataFrame.

Configuration and thresholds

The production configuration is recorded in config.json.

  • Model input resolution: 960x960
  • Keypoint format: 17 COCO keypoints [x, y, vis_confidence] (Nose, Eyes, Ears, Shoulders, Elbows, Wrists, Hips, Knees, Ankles).
  • Output feature vector: 51 features per person (17 keypoints Γ— 3) plus bounding box (x1, y1, x2, y2).
  • Default detection threshold (conf): 0.25.
  • IoU / NMS: NMS-free native end-to-end prediction head (YOLO26 architecture feature).
  • Minimum bounding box area gate: 1,024 sq pixels (32 Γ— 32). Candidates below this size are dropped to avoid noise from far-field backgrounds.

Thresholds are configurable. Tune detection confidence and keypoint confidence according to camera placement, overhead angles, mounting height, and occlusion density.

Runtime requirements

The reference deployment environment targets Python 3.11 and packages specified in requirements.txt.

Target execution environment:

  • GPU: NVIDIA L4 (Compute Capability 8.9)
  • CUDA: 12.x / 13.0
  • TensorRT: 10.x / 11.x
  • PyTorch: 2.x
  • Ultralytics: 8.x / YOLO26 core
  • OpenCV: 4.11.0.86
  • ONNX Runtime: 1.27.0
  • NumPy: 1.26.4

TensorRT engines are hardware-architecture specific. Do not run an .engine artifact compiled for an L4 GPU on a T4, A100, or RTX consumer GPU without re-exporting.

Intended Use

This model is intended for Select AI model inputs , retail customer flow monitoring, skeletal posture analysis, and camera-based engagement metric generation.

Use it as a primary upstream visual feature extractor. Keypoint predictions require downstream validation before being used for safety alerts, fall detection, or automated crowd behavior logging.

Limitations

  • Occlusion & Dense Crowds: Keypoint confidence drops significantly on heavily overlapping bodies, partial truncations, or extreme high-angle fisheye cameras.
  • Far-field Resolution: Keypoints for small subjects (<32px height) exhibit higher spatial variance and jitter.
  • Lighting Sensitivity: Low-light, high-contrast backlighting, or night-vision IR illumination increases false-positive keypoint placement.
  • No Re-Identification: The model outputs body geometry only. It does not perform person re-identification, face recognition, or tracking across frames.
  • Specialized Clothing: Loose costumes, heavy PPE, or bulky apparel can distort joint keypoint localization.

Performance

Trainer Benchmark Results

The benchmark logs are documented in docs/train.log.md. Evaluated on standard COCO Keypoints validation set (imgsz=960).

Model Variant Backend Precision Recall mAP50 (Pose) mAP50-95 (Pose) Latency (ms/img)
YOLO26x-Pose PyTorch (FP32) 0.892 0.841 0.905 0.704 14.2 ms (CPU)
YOLO26x-Pose ONNX (FP16) 0.890 0.839 0.903 0.702 8.5 ms (CPU)

Independent Validator Benchmark Results

Pending.

The validator must evaluate performance using the hold-off internal retail evaluation set (evaluation_data/) and scripts/evaluate.ipynb.

The evaluation notebook records:

  • Box mAP
  • Keypoint OKS (Object Keypoint Similarity)
  • Closed-set detection coverage
  • Frame processing throughput (FPS)
  • Failure cases across varied camera angles

Ownership

  • Upstream developer: Ultralytics
  • Select AI inventory owner: Nishant
  • Data collector/dataset owner: Select AI Vision Team
  • Independent validator: Vivek
  • Approver:
  • Publication date:
  • Independent validation date:

Repository Layout

pose-detection/
β”œβ”€β”€ README.md
β”œβ”€β”€ CHANGELOG.md
β”œβ”€β”€ config.json
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ evaluate.ipynb
β”‚   β”œβ”€β”€ run.py
β”‚   └── train.py
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ data.md
β”‚   β”œβ”€β”€ spec.md
β”‚   └── train.log.md
└── models/
    └── yolo26x-pose.pt
Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support