pose-detection / docs /spec.md
nishant2401's picture
Upload folder using huggingface_hub
9da727c verified
|
Raw History Blame Contribute Delete
3.29 kB

Model Specification

YOLO26x-Pose

| Property | Value | | Model | YOLO26x-Pose | | Task | Human Pose Estimation | | Framework | Ultralytics | | Input Resolution | 960 × 960 | | Dataset | COCO Keypoints | | Classes | 1 | | Class | person | | Keypoints | 17 | | Output features | 51 per person (17 keypoints × 3) + bbox (x1, y1, x2, y2) | | Checkpoint | models/yolo26x-pose.pt |

Model Configuration

  • Input format: RGB image
  • Input size: 960 × 960
  • Detection class: person
  • Pose keypoints: 17
  • Pose task: human keypoint estimation
  • Feature vector: 51 floats (indices 0–50 keypoint x/y/confidence); bounding box (x1, y1, x2, y2) returned separately as box_xyxy

Output Feature Vector (51) + bbox

Index range Count Block Contents
0–50 51 Keypoint features 17 COCO keypoints × (x, y, confidence)

Bounding box: box_xyxy = [x1, y1, x2, y2] (pixel coordinates, original frame). Downstream consumers assemble bbox + feature vector into a pandas DataFrame.

Keypoint feature order (each triplet x, y, confidence): nose, left_eye, right_eye, left_ear, right_ear, left_shoulder, right_shoulder, left_elbow, right_elbow, left_wrist, right_wrist, left_hip, right_hip, left_knee, right_knee, left_ankle, right_ankle.

Reference Benchmark

| Metric | Score | | mAP50-95 | 71.6% | | mAP50 | 91.6% |

These are reference benchmark values for the model and are not presented as an independently reproduced local evaluation.

Development specification

Scope

YOLO26x-Pose is a human pose-estimation model for detecting people and predicting 17 human body keypoints from an input image. Each detected person is expanded into a fixed 51-feature vector (17 keypoints × 3) plus bounding box (x1, y1, x2, y2) for downstream analytics.

The packaged v1 artifact contains the upstream pretrained YOLO26x-Pose checkpoint from Ultralytics. The model performs person detection and pose estimation in a single model pipeline. Downstream applications can use the predicted bounding boxes, confidence scores, 17 keypoints, and the 51-feature vector for pose analysis.

Architecture decisions

The model uses the Ultralytics YOLO26 pose architecture and is loaded through the Ultralytics framework.

The packaged checkpoint is configured for:

  • Task: Human Pose Estimation
  • Model: YOLO26x-Pose
  • Input resolution: 960 × 960
  • Number of classes: 1
  • Class: person
  • Number of keypoints: 17

The model accepts an RGB image and produces person detections together with human pose keypoints.

The packaged repository keeps the upstream model checkpoint and supporting inference, training-provenance, and evaluation files together. Application-specific tracking, quality filtering, identity association, or alert logic is outside the model itself.

Starting checkpoint and training

The starting and final checkpoint for v1 is the upstream pretrained YOLO26x-Pose artifact from Ultralytics.

Select AI did not train or fine-tune the checkpoint.

scripts/train.py records this provenance and intentionally does not launch a training job or download a training dataset.

The published checkpoint is:

models/yolo26x-pose.pt