Instructions to use select-ai/pose-detection with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ultralytics
How to use select-ai/pose-detection with ultralytics:
# Couldn't find a valid YOLO version tag. # Replace XX with the correct version. from ultralytics import YOLOvXX model = YOLOvXX.from_pretrained("select-ai/pose-detection") source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - TensorRT
How to use select-ai/pose-detection with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
YOLO26x Pose Model
Model description
- Model name: YOLO26x Human Pose Estimation
- Version:
v1 - Status:
experimental - Repository visibility: private to the Select AI Hugging Face organization
- Model type: human detection plus 17-keypoint COCO pose estimation, expanded to a 51-feature per-person vector plus bounding box
- Upstream trainer/developer: Ultralytics YOLO26
- Inventory owner: Akshansh D
- Architecture: YOLO26x-pose backbone with NMS-free end-to-end keypoint regression head
- Output features: 51 per detected person (17 keypoints Γ 3) plus box
(x1, y1, x2, y2) - Reference date: 2026-09-22
This repository contains the model checkpoint and inference artifacts used by the Select AI vision analytics and engagement pipelines for real-time human pose estimation, skeletal tracking, and downstream 51-feature pose vectors plus bounding boxes. It does not contain camera streams, video clips, recorded candidate poses, or downstream analytics databases.
The model is experimental. Initial trainer benchmark evaluations are complete, but independent hold-off validation is still pending.
How to use
The reference path is scripts/run.py. It accepts a single image or video frame, detects human instances, predicts 17 COCO keypoints [x, y, confidence] per person alongside bounding boxes (x1, y1, x2, y2), then expands each person into a fixed 51-feature vector for downstream analytics. Downstream consumers assemble box_xyxy + feature_vector into a pandas DataFrame.
python scripts/run.py \
--image /path/to/image.jpg \
--device 0
Optional flags: --model, --imgsz, --conf, --iou, --max-det, --output /path/to/annotated.jpg.
The script outputs JSON to stdout. Each entry in predictions carries the raw box (box_xyxy), the 17 named keypoints, and the assembled 51-feature vector (feature_vector). Pass --output to save an annotated frame with rendered skeletal connections.
Input Contract
Input: one decoded color image, typically a BGR uint8 OpenCV array or an image path accepted by
scripts/run.py.Camera frames: supports standard CCTV surveillance resolutions up to 1920 Γ 1080. No hard-coded fixed aspect ratio is required.
Preprocessing: resized and letterboxed to 960 Γ 960, producing a network tensor of shape
[1, 3, 960, 960].Color order and normalization:
scripts/run.pyreceives BGR images, converts them to RGB, and normalizes pixel intensities to the [0, 1] float range.Detector confidence threshold: 0.25 by default (
--conf).Keypoint confidence threshold: 0.50 by default for valid keypoint filtering.
IoU threshold: 0.50 by default (
--iou).
The pose model detects body keypoints independently per frame. Downstream consumers assemble box_xyxy + feature_vector into a pandas DataFrame. Temporal tracking (e.g., ByteTRACK or NVDCF) or other modules must consume that DataFrame to build multi-frame trajectories or engagement events.
Output Contract
For each detected person, scripts/run.py returns a bounding box (x1, y1, x2, y2) and a fixed 51-feature vector. Consumers can assemble these into a pandas DataFrame (helper: build_feature_frame).
CCTV IMAGE
β
βΌ
YOLO26x-Pose
β
βββ Person bounding box (x1, y1, x2, y2)
β
βββ 17 pose keypoints
β
βββ x
βββ y
βββ confidence
β
βΌ
51 features (17 keypoints Γ 3) + bbox
β
βΌ
pandas DataFrame
Feature vector layout (51 floats)
Total = 51 keypoint features (17 keypoints Γ 3).
| Index range | Count | Block | Contents |
|---|---|---|---|
| 0β50 | 51 | Keypoint features | For each COCO keypoint i (0β16, in order): x, y, confidence |
Bounding box is returned separately as box_xyxy = [x1, y1, x2, y2] (pixel coordinates on the original frame).
Keypoint feature index map:
| Index range | Keypoint | id | Fields |
|---|---|---|---|
| 0β2 | nose | 0 | x, y, confidence |
| 3β5 | left_eye | 1 | x, y, confidence |
| 6β8 | right_eye | 2 | x, y, confidence |
| 9β11 | left_ear | 3 | x, y, confidence |
| 12β14 | right_ear | 4 | x, y, confidence |
| 15β17 | left_shoulder | 5 | x, y, confidence |
| 18β20 | right_shoulder | 6 | x, y, confidence |
| 21β23 | left_elbow | 7 | x, y, confidence |
| 24β26 | right_elbow | 8 | x, y, confidence |
| 27β29 | left_wrist | 9 | x, y, confidence |
| 30β32 | right_wrist | 10 | x, y, confidence |
| 33β35 | left_hip | 11 | x, y, confidence |
| 36β38 | right_hip | 12 | x, y, confidence |
| 39β41 | left_knee | 13 | x, y, confidence |
| 42β44 | right_knee | 14 | x, y, confidence |
| 45β47 | left_ankle | 15 | x, y, confidence |
| 48β50 | right_ankle | 16 | x, y, confidence |
JSON payload
For each detected person, scripts/run.py returns:
{
"box_xyxy": [215.4, 102.1, 480.8, 610.5],
"detection_score": 0.91,
"class_id": 0,
"class_name": "person",
"keypoints": [
{"name": "nose", "id": 0, "xy": [340.1, 145.2], "confidence": 0.95},
{"name": "left_eye", "id": 1, "xy": [348.0, 140.1], "confidence": 0.94},
{"name": "right_eye", "id": 2, "xy": [332.2, 140.5], "confidence": 0.92},
{"name": "left_ear", "id": 3, "xy": [360.5, 148.0], "confidence": 0.88},
{"name": "right_ear", "id": 4, "xy": [320.1, 149.1], "confidence": 0.85},
{"name": "left_shoulder", "id": 5, "xy": [385.2, 210.4], "confidence": 0.91},
{"name": "right_shoulder", "id": 6, "xy": [295.1, 212.0], "confidence": 0.89},
{"name": "left_elbow", "id": 7, "xy": [410.0, 290.1], "confidence": 0.84},
{"name": "right_elbow", "id": 8, "xy": [270.4, 292.5], "confidence": 0.82},
{"name": "left_wrist", "id": 9, "xy": [425.1, 360.0], "confidence": 0.79},
{"name": "right_wrist", "id": 10, "xy": [255.0, 362.1], "confidence": 0.76},
{"name": "left_hip", "id": 11, "xy": [370.2, 380.0], "confidence": 0.88},
{"name": "right_hip", "id": 12, "xy": [310.4, 381.2], "confidence": 0.87},
{"name": "left_knee", "id": 13, "xy": [380.1, 490.5], "confidence": 0.85},
{"name": "right_knee", "id": 14, "xy": [305.0, 492.1], "confidence": 0.83},
{"name": "left_ankle", "id": 15, "xy": [390.0, 585.4], "confidence": 0.81},
{"name": "right_ankle", "id": 16, "xy": [298.2, 588.0], "confidence": 0.78}
],
"feature_count": 51,
"feature_vector": [
340.1, 145.2, 0.95,
348.0, 140.1, 0.94,
332.2, 140.5, 0.92,
360.5, 148.0, 0.88,
320.1, 149.1, 0.85,
385.2, 210.4, 0.91,
295.1, 212.0, 0.89,
410.0, 290.1, 0.84,
270.4, 292.5, 0.82,
425.1, 360.0, 0.79,
255.0, 362.1, 0.76,
370.2, 380.0, 0.88,
310.4, 381.2, 0.87,
380.1, 490.5, 0.85,
305.0, 492.1, 0.83,
390.0, 585.4, 0.81,
298.2, 588.0, 0.78
]
}
box_xyxy, keypoints.xy, and keypoint features (indices 0β50) use unnormalized pixel coordinates mapped directly to the original input image frame. The payload also carries feature_count (51); consumers that need a flat tensor should read feature_vector only, and combine it with box_xyxy when building a pandas DataFrame.
Configuration and thresholds
The production configuration is recorded in config.json.
- Model input resolution: 960x960
- Keypoint format: 17 COCO keypoints [x, y, vis_confidence] (Nose, Eyes, Ears, Shoulders, Elbows, Wrists, Hips, Knees, Ankles).
- Output feature vector: 51 features per person (17 keypoints Γ 3) plus bounding box
(x1, y1, x2, y2). - Default detection threshold (conf): 0.25.
- IoU / NMS: NMS-free native end-to-end prediction head (YOLO26 architecture feature).
- Minimum bounding box area gate: 1,024 sq pixels (32 Γ 32). Candidates below this size are dropped to avoid noise from far-field backgrounds.
Thresholds are configurable. Tune detection confidence and keypoint confidence according to camera placement, overhead angles, mounting height, and occlusion density.
Runtime requirements
The reference deployment environment targets Python 3.11 and packages specified in requirements.txt.
Target execution environment:
- GPU: NVIDIA L4 (Compute Capability 8.9)
- CUDA: 12.x / 13.0
- TensorRT: 10.x / 11.x
- PyTorch: 2.x
- Ultralytics: 8.x / YOLO26 core
- OpenCV: 4.11.0.86
- ONNX Runtime: 1.27.0
- NumPy: 1.26.4
TensorRT engines are hardware-architecture specific. Do not run an .engine artifact compiled for an L4 GPU on a T4, A100, or RTX consumer GPU without re-exporting.
Intended Use
This model is intended for Select AI model inputs , retail customer flow monitoring, skeletal posture analysis, and camera-based engagement metric generation.
Use it as a primary upstream visual feature extractor. Keypoint predictions require downstream validation before being used for safety alerts, fall detection, or automated crowd behavior logging.
Limitations
- Occlusion & Dense Crowds: Keypoint confidence drops significantly on heavily overlapping bodies, partial truncations, or extreme high-angle fisheye cameras.
- Far-field Resolution: Keypoints for small subjects (<32px height) exhibit higher spatial variance and jitter.
- Lighting Sensitivity: Low-light, high-contrast backlighting, or night-vision IR illumination increases false-positive keypoint placement.
- No Re-Identification: The model outputs body geometry only. It does not perform person re-identification, face recognition, or tracking across frames.
- Specialized Clothing: Loose costumes, heavy PPE, or bulky apparel can distort joint keypoint localization.
Performance
Trainer Benchmark Results
The benchmark logs are documented in docs/train.log.md. Evaluated on standard COCO Keypoints validation set (imgsz=960).
| Model Variant | Backend | Precision | Recall | mAP50 (Pose) | mAP50-95 (Pose) | Latency (ms/img) |
|---|---|---|---|---|---|---|
| YOLO26x-Pose | PyTorch (FP32) | 0.892 | 0.841 | 0.905 | 0.704 | 14.2 ms (CPU) |
| YOLO26x-Pose | ONNX (FP16) | 0.890 | 0.839 | 0.903 | 0.702 | 8.5 ms (CPU) |
Independent Validator Benchmark Results
Pending.
The validator must evaluate performance using the hold-off internal retail evaluation set (evaluation_data/) and scripts/evaluate.ipynb.
The evaluation notebook records:
- Box mAP
- Keypoint OKS (Object Keypoint Similarity)
- Closed-set detection coverage
- Frame processing throughput (FPS)
- Failure cases across varied camera angles
Ownership
- Upstream developer: Ultralytics
- Select AI inventory owner: Nishant
- Data collector/dataset owner: Select AI Vision Team
- Independent validator: Vivek
- Approver:
- Publication date:
- Independent validation date:
Repository Layout
pose-detection/
βββ README.md
βββ CHANGELOG.md
βββ config.json
βββ requirements.txt
βββ scripts/
β βββ evaluate.ipynb
β βββ run.py
β βββ train.py
βββ docs/
β βββ data.md
β βββ spec.md
β βββ train.log.md
βββ models/
βββ yolo26x-pose.pt
- Downloads last month
- 7
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js