{ "cells": [ { "cell_type": "markdown", "id": "873b8243", "metadata": {}, "source": [ "# Independent validation for YOLO26x-Pose v1\n", "\n", "This notebook runs the published YOLO26x-Pose pretrained model against\n", "evaluator-supplied human pose data using the shared inference helpers\n", "from `scripts/run.py` (`load_model`, `extract_predictions`,\n", "`build_feature_vector`, `build_feature_frame`).\n", "\n", "It does not retrain the model, modify the checkpoint, change model\n", "weights, or update the README automatically.\n", "\n", "## Feature pipeline\n", "\n", "```text\n", "CCTV IMAGE\n", " │\n", " ▼\n", "YOLO26x-Pose\n", " │\n", " ├── Person bounding box (x1, y1, x2, y2)\n", " │\n", " └── 17 pose keypoints\n", " │\n", " ├── x\n", " ├── y\n", " └── confidence\n", " │\n", " ▼\n", " 51 features (17 keypoints × 3) + bbox\n", " │\n", " ▼\n", " pandas DataFrame\n", "```\n", "\n", "Each detected person is expanded into a 51-feature vector\n", "(17 keypoints × 3) plus bounding box `(x1, y1, x2, y2)`,\n", "matching the `box_xyxy` and `feature_vector` fields emitted by\n", "`scripts/run.py`. Downstream, those fields are assembled into a\n", "pandas DataFrame; this notebook validates that stage.\n", "\n", "Run this notebook from the repository root after placing the hold-off\n", "evaluation data in the structure below.\n", "\n", "```text\n", "evaluation_data/\n", "├── images/\n", "│ ├── image_001.jpg\n", "│ ├── image_002.jpg\n", "│ └── ...\n", "├── labels/\n", "│ ├── image_001.txt\n", "│ ├── image_002.txt\n", "│ └── ...\n", "└── data.yaml" ] }, { "cell_type": "code", "execution_count": null, "id": "f9861682", "metadata": {}, "outputs": [], "source": [ "import json\n", "import math\n", "import sys\n", "import time\n", "from pathlib import Path\n", "\n", "import cv2\n", "import numpy as np\n", "import pandas as pd\n", "import matplotlib.pyplot as plt\n", "\n", "from PIL import Image\n", "\n", "from sklearn.metrics import (\n", " accuracy_score,\n", " confusion_matrix,\n", " f1_score,\n", " precision_score,\n", " recall_score,\n", ")\n", "\n", "def find_repo_root():\n", " candidates = [\n", " Path.cwd(),\n", " Path.cwd().parent,\n", " Path.cwd().parent.parent,\n", " ]\n", "\n", " for candidate in candidates:\n", " if (\n", " (candidate / 'models').is_dir()\n", " and (candidate / 'scripts' / 'run.py').exists()\n", " ):\n", " return candidate.resolve()\n", "\n", " raise FileNotFoundError(\n", " 'Run this notebook from the YOLO26x-Pose repository root.'\n", " )\n", "\n", "REPO_ROOT = find_repo_root()\n", "\n", "sys.path.insert(0, str(REPO_ROOT / 'scripts'))\n", "\n", "# Reuse the reference inference path from run.py so evaluation stays\n", "# aligned with the 51-feature + bbox output contract.\n", "import run as run_module\n", "from run import (\n", " KEYPOINT_NAMES,\n", " NUM_KEYPOINTS,\n", " KEYPOINT_FEATURE_COUNT,\n", " TOTAL_FEATURE_COUNT,\n", " BBOX_FIELDS,\n", " IMAGE_SUFFIXES,\n", " build_feature_frame,\n", " build_feature_vector,\n", " extract_predictions,\n", " load_model,\n", ")\n", "\n", "# Change this path if your hold-off evaluation data lives outside the repository.\n", "EVALUATION_ROOT = REPO_ROOT / 'evaluation_data'\n", "\n", "MODEL_PATH = REPO_ROOT / 'models' / 'yolo26x-pose.pt'\n", "\n", "IMAGE_ROOT = EVALUATION_ROOT / 'images'\n", "LABEL_ROOT = EVALUATION_ROOT / 'labels'\n", "\n", "# Defaults must match run.py\n", "IMAGE_SIZE = run_module.DEFAULT_IMAGE_SIZE\n", "DEVICE = run_module.DEFAULT_DEVICE\n", "CONF_THRESHOLD = run_module.DEFAULT_CONFIDENCE\n", "IOU_THRESHOLD = run_module.DEFAULT_IOU\n", "MAX_DET = 100\n", "\n", "CLASS_ID = 0\n", "CLASS_NAME = 'person'\n", "\n", "OUTPUT_ROOT = REPO_ROOT / 'results' / 'independent_validation'\n", "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", "\n", "OUTPUT_JSON = OUTPUT_ROOT / 'independent_validation_results.json'\n", "OUTPUT_CSV = OUTPUT_ROOT / 'independent_validation_results.csv'\n", "\n", "IMAGE_SUFFIXES = set(run_module.IMAGE_SUFFIXES)\n", "\n", "print('Repository:', REPO_ROOT)\n", "print('Evaluation data:', EVALUATION_ROOT)\n", "print('Model:', MODEL_PATH)\n", "print('Images:', IMAGE_ROOT)\n", "print('Labels:', LABEL_ROOT)\n", "print('Input size:', IMAGE_SIZE)\n", "print('Device:', DEVICE)\n", "print('Confidence threshold:', CONF_THRESHOLD)\n", "print('IoU threshold:', IOU_THRESHOLD)\n", "print('Task: pose')\n", "print('Class:', CLASS_NAME)\n", "print('Keypoints:', NUM_KEYPOINTS)\n", "print('Keypoint features:', KEYPOINT_FEATURE_COUNT)\n", "print('Total features per person:', TOTAL_FEATURE_COUNT)\n", "print('Bbox fields:', BBOX_FIELDS)\n", "print('Max detections:', MAX_DET)\n", "print('Feature source: run.build_feature_vector (shared with scripts/run.py)')\n", "print('DataFrame helper: run.build_feature_frame (x1,y1,x2,y2 + 51 features)')" ] }, { "cell_type": "code", "execution_count": null, "id": "859918a6", "metadata": {}, "outputs": [], "source": [ "def image_paths(root):\n", "\n", " return sorted(\n", " path\n", " for path in root.rglob('*')\n", " if path.is_file()\n", " and path.suffix.lower() in IMAGE_SUFFIXES\n", " )\n", "\n", "\n", "def label_for_image(image_path, image_root, label_root):\n", "\n", " relative = image_path.relative_to(image_root)\n", " label_path = label_root / relative.with_suffix('.txt')\n", "\n", " return label_path\n", "\n", "\n", "image_root = EVALUATION_ROOT / 'images'\n", "label_root = EVALUATION_ROOT / 'labels'\n", "\n", "if not image_root.is_dir() or not label_root.is_dir():\n", "\n", " raise FileNotFoundError(\n", " f'Create {image_root} and {label_root} using the folder structure in the first cell.'\n", " )\n", "\n", "\n", "evaluation_paths = image_paths(image_root)\n", "\n", "if not evaluation_paths:\n", "\n", " raise ValueError(\n", " f'No evaluation images found under {image_root}'\n", " )\n", "\n", "\n", "evaluation_manifest = pd.DataFrame(\n", " {\n", " 'path': [str(p.resolve()) for p in evaluation_paths]\n", " }\n", ")\n", "\n", "evaluation_manifest['label_path'] = [\n", " str(label_for_image(path, image_root, label_root).resolve())\n", " for path in evaluation_paths\n", "]\n", "\n", "evaluation_manifest['image_name'] = [\n", " path.name\n", " for path in evaluation_paths\n", "]\n", "\n", "evaluation_manifest['label_exists'] = [\n", " Path(path).exists()\n", " for path in evaluation_manifest['label_path']\n", "]\n", "\n", "missing_labels = evaluation_manifest[\n", " ~evaluation_manifest['label_exists']\n", "]\n", "\n", "print(\n", " f'Evaluation images: {len(evaluation_manifest)}'\n", ")\n", "\n", "print(\n", " f'Images with labels: '\n", " f'{evaluation_manifest[\"label_exists\"].sum()}'\n", ")\n", "\n", "print(\n", " f'Images without labels: '\n", " f'{len(missing_labels)}'\n", ")\n", "\n", "if not missing_labels.empty:\n", "\n", " print(\n", " 'First images without labels:',\n", " missing_labels['image_name'].head(10).tolist()\n", " )\n", "\n", "print(\n", " 'Image root:',\n", " image_root\n", ")\n", "\n", "print(\n", " 'Label root:',\n", " label_root\n", ")\n", "\n", "print(\n", " 'Evaluation dataset ready:',\n", " len(evaluation_manifest) > 0\n", ")" ] }, { "cell_type": "code", "execution_count": null, "id": "55440b8c", "metadata": {}, "outputs": [], "source": [ "from ultralytics import YOLO\n", "\n", "runtime = load_model(MODEL_PATH)\n", "\n", "print(f'Loaded YOLO26x-Pose model from {MODEL_PATH}')\n", "print(f'Task: {runtime.task}')\n", "print(f'Classes: {runtime.names}')\n", "print(f'Input size: {IMAGE_SIZE}')\n", "print(f'Confidence threshold: {CONF_THRESHOLD}')\n", "print(f'IoU threshold: {IOU_THRESHOLD}')\n", "print(f'Feature vector size: {TOTAL_FEATURE_COUNT}')" ] }, { "cell_type": "code", "execution_count": null, "id": "97c5d6ba", "metadata": {}, "outputs": [], "source": [ "def extract_records(manifest):\n", "\n", " records = []\n", "\n", " started = time.perf_counter()\n", "\n", " for row in manifest.itertuples(index=False):\n", "\n", " image_path = Path(row.path)\n", "\n", " image = cv2.imread(\n", " str(image_path),\n", " cv2.IMREAD_COLOR\n", " )\n", "\n", " if image is None:\n", "\n", " records.append({\n", " 'path': str(image_path),\n", " 'image_name': image_path.name,\n", " 'predicted_boxes': None,\n", " 'predicted_keypoints': None,\n", " 'predicted_keypoint_conf': None,\n", " 'prediction_scores': None,\n", " 'predicted_feature_vectors': None,\n", " 'predicted_persons': 0,\n", " 'image_width': None,\n", " 'image_height': None,\n", " 'inference_seconds': None,\n", " 'error': 'image could not be decoded',\n", " })\n", "\n", " continue\n", "\n", " height, width = image.shape[:2]\n", "\n", " inference_started = time.perf_counter()\n", "\n", " # Same predict call shape as run.run_pose\n", " results = runtime.predict(\n", " source=image,\n", " imgsz=IMAGE_SIZE,\n", " conf=CONF_THRESHOLD,\n", " iou=IOU_THRESHOLD,\n", " device=DEVICE,\n", " max_det=MAX_DET,\n", " verbose=False,\n", " save=False,\n", " )\n", "\n", " inference_elapsed = (\n", " time.perf_counter() - inference_started\n", " )\n", "\n", " result = results[0]\n", "\n", " # Shared extraction path with scripts/run.py (includes\n", " # 51-feature vectors and box_xyxy).\n", " predictions = extract_predictions(result)\n", "\n", " if predictions:\n", "\n", " predicted_boxes = np.asarray(\n", " [p['box_xyxy'] for p in predictions],\n", " dtype=np.float32,\n", " )\n", "\n", " prediction_scores = np.asarray(\n", " [p['confidence'] for p in predictions],\n", " dtype=np.float32,\n", " )\n", "\n", " predicted_keypoints = np.asarray(\n", " [\n", " [[kp['x'], kp['y']] for kp in p['keypoints']]\n", " for p in predictions\n", " ],\n", " dtype=np.float32,\n", " ) if predictions[0]['keypoints'] else np.empty(\n", " (0, NUM_KEYPOINTS, 2),\n", " dtype=np.float32,\n", " )\n", "\n", " predicted_keypoint_conf = np.asarray(\n", " [\n", " [\n", " kp['confidence'] if kp['confidence'] is not None else 0.0\n", " for kp in p['keypoints']\n", " ]\n", " for p in predictions\n", " ],\n", " dtype=np.float32,\n", " ) if predictions[0]['keypoints'] else np.empty(\n", " (0, NUM_KEYPOINTS),\n", " dtype=np.float32,\n", " )\n", "\n", " predicted_feature_vectors = np.asarray(\n", " [p['feature_vector'] for p in predictions],\n", " dtype=np.float32,\n", " )\n", "\n", " for p in predictions:\n", "\n", " if len(p['feature_vector']) != TOTAL_FEATURE_COUNT:\n", "\n", " raise RuntimeError(\n", " f\"{image_path.name}: expected \"\n", " f\"{TOTAL_FEATURE_COUNT} features, got \"\n", " f\"{len(p['feature_vector'])}.\"\n", " )\n", "\n", " else:\n", "\n", " predicted_boxes = np.empty((0, 4), dtype=np.float32)\n", " prediction_scores = np.empty((0,), dtype=np.float32)\n", " predicted_keypoints = np.empty(\n", " (0, NUM_KEYPOINTS, 2), dtype=np.float32\n", " )\n", " predicted_keypoint_conf = np.empty(\n", " (0, NUM_KEYPOINTS), dtype=np.float32\n", " )\n", " predicted_feature_vectors = np.empty(\n", " (0, TOTAL_FEATURE_COUNT), dtype=np.float32\n", " )\n", "\n", " records.append({\n", " 'path': str(image_path),\n", " 'image_name': image_path.name,\n", " 'predicted_boxes': predicted_boxes,\n", " 'predicted_keypoints': predicted_keypoints,\n", " 'predicted_keypoint_conf': predicted_keypoint_conf,\n", " 'prediction_scores': prediction_scores,\n", " 'predicted_feature_vectors': predicted_feature_vectors,\n", " 'predicted_persons': len(predicted_boxes),\n", " 'image_width': width,\n", " 'image_height': height,\n", " 'inference_seconds': inference_elapsed,\n", " 'error': None,\n", " })\n", "\n", " elapsed = time.perf_counter() - started\n", "\n", " usable = sum(\n", " record['error'] is None\n", " for record in records\n", " )\n", "\n", " predicted_persons = sum(\n", " record['predicted_persons']\n", " for record in records\n", " )\n", "\n", " feature_vectors = [\n", " vector\n", " for record in records\n", " if record.get('predicted_feature_vectors') is not None\n", " for vector in record['predicted_feature_vectors']\n", " ]\n", "\n", " feature_lengths = [\n", " len(vector)\n", " for vector in feature_vectors\n", " ]\n", "\n", " inference_times = [\n", " record['inference_seconds']\n", " for record in records\n", " if record['inference_seconds'] is not None\n", " ]\n", "\n", " stats = {\n", " 'images': len(records),\n", " 'usable_images': usable,\n", " 'prediction_success_rate': (\n", " usable / len(records)\n", " if records\n", " else 0.0\n", " ),\n", " 'predicted_persons': predicted_persons,\n", " 'feature_vectors_emitted': len(feature_vectors),\n", " 'feature_vector_size_expected': TOTAL_FEATURE_COUNT,\n", " 'feature_vector_size_min': (\n", " int(min(feature_lengths))\n", " if feature_lengths\n", " else 0\n", " ),\n", " 'feature_vector_size_max': (\n", " int(max(feature_lengths))\n", " if feature_lengths\n", " else 0\n", " ),\n", " 'feature_vector_size_valid': all(\n", " length == TOTAL_FEATURE_COUNT\n", " for length in feature_lengths\n", " ),\n", " 'keypoint_feature_count': KEYPOINT_FEATURE_COUNT,\n", " 'bbox_fields': list(BBOX_FIELDS),\n", " 'elapsed_seconds': elapsed,\n", " 'images_per_second': (\n", " len(records) / elapsed\n", " if elapsed\n", " else 0.0\n", " ),\n", " 'persons_per_second': (\n", " predicted_persons / elapsed\n", " if elapsed\n", " else 0.0\n", " ),\n", " 'mean_inference_ms': (\n", " float(np.mean(inference_times) * 1000)\n", " if inference_times\n", " else 0.0\n", " ),\n", " 'median_inference_ms': (\n", " float(np.median(inference_times) * 1000)\n", " if inference_times\n", " else 0.0\n", " ),\n", " }\n", "\n", " return records, stats\n", "\n", "\n", "evaluation_records, evaluation_stats = extract_records(\n", " evaluation_manifest\n", ")\n", "\n", "record_by_path = {\n", " record['path']: record\n", " for record in evaluation_records\n", "}\n", "\n", "print(\n", " json.dumps(\n", " evaluation_stats,\n", " indent=2\n", " )\n", ")\n", "\n", "# Assemble 51 features + bbox into a pandas DataFrame\n", "feature_rows = []\n", "\n", "for record in evaluation_records:\n", "\n", " if (\n", " record.get('predicted_boxes') is None\n", " or record.get('predicted_feature_vectors') is None\n", " ):\n", " continue\n", "\n", " for box, vector in zip(\n", " record['predicted_boxes'],\n", " record['predicted_feature_vectors'],\n", " ):\n", " feature_rows.append(\n", " [float(v) for v in box] + [float(v) for v in vector]\n", " )\n", "\n", "feature_columns = (\n", " list(BBOX_FIELDS) + run_module.feature_column_names()\n", ")\n", "\n", "feature_frame = pd.DataFrame(\n", " feature_rows,\n", " columns=feature_columns,\n", ")\n", "\n", "print('Feature DataFrame shape:', feature_frame.shape)\n", "print('Expected columns:', 4 + TOTAL_FEATURE_COUNT)\n", "print(feature_frame.head())" ] }, { "cell_type": "code", "execution_count": null, "id": "c9c4a812", "metadata": {}, "outputs": [], "source": [ "def load_ground_truth(image_path):\n", "\n", " image_path = Path(image_path)\n", "\n", " label_path = label_for_image(\n", " image_path,\n", " image_root,\n", " label_root,\n", " )\n", "\n", " if not label_path.exists():\n", "\n", " return {\n", " 'boxes': np.empty(\n", " (0, 4),\n", " dtype=np.float32\n", " ),\n", " 'keypoints': np.empty(\n", " (0, NUM_KEYPOINTS, 3),\n", " dtype=np.float32\n", " ),\n", " 'error': 'ground-truth label not found',\n", " }\n", "\n", " with Image.open(image_path) as image:\n", "\n", " width, height = image.size\n", "\n", " ground_truth_boxes = []\n", " ground_truth_keypoints = []\n", "\n", " for line_number, line in enumerate(\n", " label_path.read_text().splitlines(),\n", " start=1\n", " ):\n", "\n", " values = line.strip().split()\n", "\n", " if not values:\n", " continue\n", "\n", " expected_values = (\n", " 5 + NUM_KEYPOINTS * 3\n", " )\n", "\n", " if len(values) != expected_values:\n", "\n", " raise ValueError(\n", " f'{label_path}:{line_number} has '\n", " f'{len(values)} values; expected '\n", " f'{expected_values}.'\n", " )\n", "\n", " numbers = np.asarray(\n", " [float(value) for value in values],\n", " dtype=np.float32\n", " )\n", "\n", " class_id = int(numbers[0])\n", "\n", " if class_id != CLASS_ID:\n", " continue\n", "\n", " x_center, y_center, box_width, box_height = (\n", " numbers[1:5]\n", " )\n", "\n", " x1 = (\n", " x_center - box_width / 2\n", " ) * width\n", "\n", " y1 = (\n", " y_center - box_height / 2\n", " ) * height\n", "\n", " x2 = (\n", " x_center + box_width / 2\n", " ) * width\n", "\n", " y2 = (\n", " y_center + box_height / 2\n", " ) * height\n", "\n", " ground_truth_boxes.append(\n", " [x1, y1, x2, y2]\n", " )\n", "\n", " keypoints = numbers[5:].reshape(\n", " NUM_KEYPOINTS,\n", " 3\n", " )\n", "\n", " keypoints[:, 0] *= width\n", " keypoints[:, 1] *= height\n", "\n", " ground_truth_keypoints.append(\n", " keypoints\n", " )\n", "\n", " return {\n", " 'boxes': np.asarray(\n", " ground_truth_boxes,\n", " dtype=np.float32\n", " ).reshape(-1, 4),\n", "\n", " 'keypoints': np.asarray(\n", " ground_truth_keypoints,\n", " dtype=np.float32\n", " ).reshape(-1, NUM_KEYPOINTS, 3),\n", "\n", " 'error': None,\n", " }\n", "\n", "\n", "def box_iou(box_a, box_b):\n", "\n", " x1 = max(box_a[0], box_b[0])\n", " y1 = max(box_a[1], box_b[1])\n", " x2 = min(box_a[2], box_b[2])\n", " y2 = min(box_a[3], box_b[3])\n", "\n", " intersection_width = max(\n", " 0.0,\n", " x2 - x1\n", " )\n", "\n", " intersection_height = max(\n", " 0.0,\n", " y2 - y1\n", " )\n", "\n", " intersection = (\n", " intersection_width\n", " * intersection_height\n", " )\n", "\n", " area_a = (\n", " max(0.0, box_a[2] - box_a[0])\n", " * max(0.0, box_a[3] - box_a[1])\n", " )\n", "\n", " area_b = (\n", " max(0.0, box_b[2] - box_b[0])\n", " * max(0.0, box_b[3] - box_b[1])\n", " )\n", "\n", " union = area_a + area_b - intersection\n", "\n", " if union <= 0:\n", " return 0.0\n", "\n", " return intersection / union\n", "\n", "\n", "def match_persons(\n", " ground_truth_boxes,\n", " predicted_boxes,\n", " threshold=IOU_THRESHOLD,\n", "):\n", "\n", " candidates = []\n", "\n", " for gt_index, gt_box in enumerate(\n", " ground_truth_boxes\n", " ):\n", "\n", " for prediction_index, prediction_box in enumerate(\n", " predicted_boxes\n", " ):\n", "\n", " iou = box_iou(\n", " gt_box,\n", " prediction_box\n", " )\n", "\n", " candidates.append(\n", " (\n", " iou,\n", " gt_index,\n", " prediction_index,\n", " )\n", " )\n", "\n", " candidates.sort(\n", " key=lambda item: item[0],\n", " reverse=True\n", " )\n", "\n", " used_ground_truth = set()\n", " used_predictions = set()\n", " matches = []\n", "\n", " for iou, gt_index, prediction_index in candidates:\n", "\n", " if iou < threshold:\n", " break\n", "\n", " if gt_index in used_ground_truth:\n", " continue\n", "\n", " if prediction_index in used_predictions:\n", " continue\n", "\n", " used_ground_truth.add(gt_index)\n", " used_predictions.add(prediction_index)\n", "\n", " matches.append(\n", " (\n", " gt_index,\n", " prediction_index,\n", " float(iou),\n", " )\n", " )\n", "\n", " return matches\n", "\n", "\n", "matching_records = []\n", "\n", "total_ground_truth = 0\n", "total_predictions = 0\n", "total_matches = 0\n", "\n", "for record in evaluation_records:\n", "\n", " if record['error'] is not None:\n", " continue\n", "\n", " ground_truth = load_ground_truth(\n", " record['path']\n", " )\n", "\n", " gt_boxes = ground_truth['boxes']\n", " predicted_boxes = record['predicted_boxes']\n", "\n", " matches = match_persons(\n", " gt_boxes,\n", " predicted_boxes,\n", " IOU_THRESHOLD,\n", " )\n", "\n", " total_ground_truth += len(gt_boxes)\n", " total_predictions += len(predicted_boxes)\n", " total_matches += len(matches)\n", "\n", " for gt_index, prediction_index, iou in matches:\n", "\n", " matching_records.append({\n", " 'path': record['path'],\n", " 'gt_index': gt_index,\n", " 'prediction_index': prediction_index,\n", " 'bbox_iou': iou,\n", " })\n", "\n", "print(\n", " f'Ground-truth persons: {total_ground_truth}'\n", ")\n", "\n", "print(\n", " f'Predicted persons: {total_predictions}'\n", ")\n", "\n", "print(\n", " f'Matched persons: {total_matches}'\n", ")\n", "\n", "print(\n", " f'Missed persons: '\n", " f'{total_ground_truth - total_matches}'\n", ")\n", "\n", "print(\n", " f'False-positive predictions: '\n", " f'{total_predictions - total_matches}'\n", ")" ] }, { "cell_type": "code", "execution_count": null, "id": "d8ef64c0", "metadata": {}, "outputs": [], "source": [ "matched_persons = total_matches\n", "\n", "missed_persons = (\n", " total_ground_truth\n", " - matched_persons\n", ")\n", "\n", "false_positive_persons = (\n", " total_predictions\n", " - matched_persons\n", ")\n", "\n", "detection_precision = (\n", " matched_persons\n", " / (\n", " matched_persons\n", " + false_positive_persons\n", " )\n", " if (\n", " matched_persons\n", " + false_positive_persons\n", " )\n", " else 0.0\n", ")\n", "\n", "detection_recall = (\n", " matched_persons\n", " / (\n", " matched_persons\n", " + missed_persons\n", " )\n", " if (\n", " matched_persons\n", " + missed_persons\n", " )\n", " else 0.0\n", ")\n", "\n", "if (\n", " detection_precision\n", " + detection_recall\n", "):\n", "\n", " detection_f1 = (\n", " 2\n", " * detection_precision\n", " * detection_recall\n", " / (\n", " detection_precision\n", " + detection_recall\n", " )\n", " )\n", "\n", "else:\n", "\n", " detection_f1 = 0.0\n", "\n", "\n", "detection_results = {\n", " 'iou_threshold': float(IOU_THRESHOLD),\n", " 'ground_truth_persons': int(\n", " total_ground_truth\n", " ),\n", " 'predicted_persons': int(\n", " total_predictions\n", " ),\n", " 'matched_persons': int(\n", " matched_persons\n", " ),\n", " 'missed_persons': int(\n", " missed_persons\n", " ),\n", " 'false_positive_persons': int(\n", " false_positive_persons\n", " ),\n", " 'precision': float(\n", " detection_precision\n", " ),\n", " 'recall': float(\n", " detection_recall\n", " ),\n", " 'f1': float(\n", " detection_f1\n", " ),\n", "}\n", "\n", "print(\n", " json.dumps(\n", " detection_results,\n", " indent=2\n", " )\n", ")" ] }, { "cell_type": "code", "execution_count": null, "id": "cd644c2d", "metadata": {}, "outputs": [], "source": [ "def normalized_keypoint_distance(\n", " ground_truth_keypoints,\n", " predicted_keypoints,\n", " ground_truth_box,\n", "):\n", "\n", " box_width = (\n", " ground_truth_box[2]\n", " - ground_truth_box[0]\n", " )\n", "\n", " box_height = (\n", " ground_truth_box[3]\n", " - ground_truth_box[1]\n", " )\n", "\n", " normalization = max(\n", " box_width,\n", " box_height,\n", " 1.0,\n", " )\n", "\n", " distances = []\n", "\n", " for keypoint_index in range(\n", " NUM_KEYPOINTS\n", " ):\n", "\n", " visibility = (\n", " ground_truth_keypoints[\n", " keypoint_index,\n", " 2\n", " ]\n", " )\n", "\n", " if visibility <= 0:\n", " continue\n", "\n", " gt_xy = ground_truth_keypoints[\n", " keypoint_index,\n", " :2\n", " ]\n", "\n", " predicted_xy = predicted_keypoints[\n", " keypoint_index\n", " ]\n", "\n", " distance = (\n", " np.linalg.norm(\n", " gt_xy - predicted_xy\n", " )\n", " / normalization\n", " )\n", "\n", " distances.append(\n", " (\n", " keypoint_index,\n", " float(distance),\n", " )\n", " )\n", "\n", " return distances\n", "\n", "\n", "keypoint_distances = {\n", " name: []\n", " for name in KEYPOINT_NAMES\n", "}\n", "\n", "all_keypoint_distances = []\n", "\n", "for match in matching_records:\n", "\n", " record = record_by_path[\n", " match['path']\n", " ]\n", "\n", " ground_truth = load_ground_truth(\n", " match['path']\n", " )\n", "\n", " gt_index = match['gt_index']\n", " prediction_index = match[\n", " 'prediction_index'\n", " ]\n", "\n", " gt_keypoints = ground_truth[\n", " 'keypoints'\n", " ][gt_index]\n", "\n", " predicted_keypoints = (\n", " record['predicted_keypoints']\n", " [prediction_index]\n", " )\n", "\n", " gt_box = ground_truth[\n", " 'boxes'\n", " ][gt_index]\n", "\n", " distances = normalized_keypoint_distance(\n", " gt_keypoints,\n", " predicted_keypoints,\n", " gt_box,\n", " )\n", "\n", " for keypoint_index, distance in distances:\n", "\n", " keypoint_name = KEYPOINT_NAMES[\n", " keypoint_index\n", " ]\n", "\n", " keypoint_distances[\n", " keypoint_name\n", " ].append(distance)\n", "\n", " all_keypoint_distances.append(\n", " distance\n", " )\n", "\n", "\n", "mean_keypoint_error = (\n", " float(\n", " np.mean(\n", " all_keypoint_distances\n", " )\n", " )\n", " if all_keypoint_distances\n", " else None\n", ")\n", "\n", "print(\n", " 'Matched persons:',\n", " len(matching_records)\n", ")\n", "\n", "print(\n", " 'Visible keypoint measurements:',\n", " len(all_keypoint_distances)\n", ")\n", "\n", "print(\n", " 'Mean normalized keypoint error:',\n", " mean_keypoint_error\n", ")\n", "\n", "for name in KEYPOINT_NAMES:\n", "\n", " values = keypoint_distances[name]\n", "\n", " mean_value = (\n", " float(np.mean(values))\n", " if values\n", " else None\n", " )\n", "\n", " print(\n", " f'{name:16s}: '\n", " f'{mean_value}'\n", " )" ] }, { "cell_type": "code", "execution_count": null, "id": "9c7c58f5", "metadata": {}, "outputs": [], "source": [ "PCK_THRESHOLDS = [\n", " 0.05,\n", " 0.10,\n", " 0.20,\n", "]\n", "\n", "pck_results = {}\n", "\n", "for threshold in PCK_THRESHOLDS:\n", "\n", " if all_keypoint_distances:\n", "\n", " correct = sum(\n", " distance <= threshold\n", " for distance\n", " in all_keypoint_distances\n", " )\n", "\n", " pck = (\n", " correct\n", " / len(all_keypoint_distances)\n", " )\n", "\n", " else:\n", "\n", " pck = 0.0\n", "\n", " pck_results[\n", " str(threshold)\n", " ] = float(pck)\n", "\n", "\n", "per_keypoint_pck = {}\n", "\n", "for keypoint_name in KEYPOINT_NAMES:\n", "\n", " values = keypoint_distances[\n", " keypoint_name\n", " ]\n", "\n", " threshold_results = {}\n", "\n", " for threshold in PCK_THRESHOLDS:\n", "\n", " if values:\n", "\n", " threshold_results[\n", " str(threshold)\n", " ] = float(\n", " np.mean(\n", " np.asarray(values)\n", " <= threshold\n", " )\n", " )\n", "\n", " else:\n", "\n", " threshold_results[\n", " str(threshold)\n", " ] = None\n", "\n", " per_keypoint_pck[\n", " keypoint_name\n", " ] = threshold_results\n", "\n", "\n", "print('Overall PCK:')\n", "\n", "for threshold, value in pck_results.items():\n", "\n", " print(\n", " f'PCK@{threshold}: '\n", " f'{value:.4f}'\n", " )\n", "\n", "\n", "print('\\nPer-keypoint PCK:')\n", "\n", "for keypoint_name, values in (\n", " per_keypoint_pck.items()\n", "):\n", "\n", " print(\n", " f'{keypoint_name:16s}: '\n", " f'{values}'\n", " )" ] }, { "cell_type": "code", "execution_count": null, "id": "08976c79", "metadata": {}, "outputs": [], "source": [ "summary = {\n", " 'model_version': 'v1',\n", " 'model': 'YOLO26x-Pose',\n", " 'task': 'human pose estimation',\n", " 'checkpoint': str(\n", " MODEL_PATH\n", " ),\n", " 'reference_script': 'scripts/run.py',\n", " 'device': DEVICE,\n", " 'input_size': IMAGE_SIZE,\n", " 'class_id': CLASS_ID,\n", " 'class_name': CLASS_NAME,\n", " 'num_keypoints': NUM_KEYPOINTS,\n", " 'keypoint_names': list(KEYPOINT_NAMES),\n", " 'keypoint_feature_count': KEYPOINT_FEATURE_COUNT,\n", " 'bbox_fields': list(BBOX_FIELDS),\n", " 'total_feature_count': TOTAL_FEATURE_COUNT,\n", " 'feature_layout': 'indices 0-50: 17 keypoints x (x, y, confidence); box_xyxy = [x1, y1, x2, y2] returned separately',\n", " 'confidence_threshold': CONF_THRESHOLD,\n", " 'bbox_iou_threshold': IOU_THRESHOLD,\n", " 'max_det': MAX_DET,\n", " 'evaluation_root': str(\n", " EVALUATION_ROOT\n", " ),\n", " 'evaluation': evaluation_stats,\n", " 'detection': detection_results,\n", " 'feature_vectors': {\n", " 'expected_size': TOTAL_FEATURE_COUNT,\n", " 'emitted': evaluation_stats.get('feature_vectors_emitted', 0),\n", " 'all_valid': evaluation_stats.get('feature_vector_size_valid', False),\n", " 'keypoint_feature_count': KEYPOINT_FEATURE_COUNT,\n", " 'bbox_fields': list(BBOX_FIELDS),\n", " },\n", " 'keypoints': {\n", " 'mean_normalized_error':\n", " mean_keypoint_error,\n", " 'pck': pck_results,\n", " 'per_keypoint_pck':\n", " per_keypoint_pck,\n", " },\n", " 'official_reference_benchmark': {\n", " 'mAP50-95': 0.716,\n", " 'mAP50': 0.916,\n", " 'input_resolution': '960x960',\n", " 'dataset': 'COCO Keypoints',\n", " 'locally_reproduced': False,\n", " },\n", " 'ownership': {\n", " 'inventory_owner': 'Nishant',\n", " 'independent_validator': 'Vivek',\n", " },\n", "}\n", "\n", "OUTPUT_JSON.parent.mkdir(\n", " parents=True,\n", " exist_ok=True\n", ")\n", "\n", "OUTPUT_JSON.write_text(\n", " json.dumps(\n", " summary,\n", " indent=2\n", " ) + '\\n'\n", ")\n", "\n", "print(\n", " f'Wrote validation summary to '\n", " f'{OUTPUT_JSON}'\n", ")" ] }, { "cell_type": "code", "execution_count": null, "id": "8419e9be", "metadata": {}, "outputs": [], "source": [ "plot_path = (\n", " OUTPUT_ROOT\n", " / 'keypoint_error_by_joint.png'\n", ")\n", "\n", "plot_values = []\n", "\n", "for keypoint_name in KEYPOINT_NAMES:\n", "\n", " values = keypoint_distances[\n", " keypoint_name\n", " ]\n", "\n", " if values:\n", "\n", " plot_values.append(\n", " float(np.mean(values))\n", " )\n", "\n", " else:\n", "\n", " plot_values.append(\n", " np.nan\n", " )\n", "\n", "\n", "plt.figure(\n", " figsize=(12, 5)\n", ")\n", "\n", "plt.bar(\n", " KEYPOINT_NAMES,\n", " plot_values\n", ")\n", "\n", "plt.xlabel(\n", " 'Human pose keypoint'\n", ")\n", "\n", "plt.ylabel(\n", " 'Mean normalized keypoint error'\n", ")\n", "\n", "plt.title(\n", " 'YOLO26x-Pose keypoint error by joint'\n", ")\n", "\n", "plt.xticks(\n", " rotation=60,\n", " ha='right'\n", ")\n", "\n", "plt.grid(\n", " axis='y',\n", " alpha=0.25\n", ")\n", "\n", "plt.tight_layout()\n", "\n", "plt.savefig(\n", " plot_path,\n", " dpi=160,\n", " bbox_inches='tight'\n", ")\n", "\n", "plt.show()\n", "\n", "print(\n", " f'Saved keypoint error plot to {plot_path}'\n", ")\n", "\n", "print(\n", " 'Official reference mAP50-95: 71.6%'\n", ")\n", "\n", "print(\n", " 'Official reference mAP50: 91.6%'\n", ")\n", "\n", "print(\n", " 'Official reference values are not '\n", " 'presented as locally reproduced results.'\n", ")" ] }, { "cell_type": "markdown", "id": "939c4466", "metadata": {}, "source": [ "## Validator handoff\n", "\n", "Before updating the README, record the hold-off dataset source and version, capture period, evaluation image count, labeled person count, hardware, backend, input resolution, confidence threshold, bounding-box IoU threshold, detection coverage, throughput, inference latency, keypoint error, PCK results, and any failure cases.\n", "\n", "Record the validator name and validation date in the Independent Validator Benchmark Results section.\n", "\n", "Keep the model status `experimental` until the independent validation review is complete.\n", "\n", "The official Ultralytics reference benchmark for YOLO26x-Pose must remain clearly separated from independently measured validation results.\n", "\n", "Official reference benchmark:\n", "\n", "- mAP50-95: 71.6%\n", "- mAP50: 91.6%\n", "- Input resolution: 960 × 960\n", "- Dataset: COCO Keypoints\n", "\n", "These reference values must not be presented as locally reproduced measurements.\n", "\n", "For the independent validation, record:\n", "\n", "- Model: YOLO26x-Pose\n", "- Task: Human Pose Estimation\n", "- Owner: Nishant\n", "- Independent Validator: Vivek\n", "- Validation date: record on completion\n", "- Hold-off dataset source and version: record on completion\n", "- Capture period: record on completion\n", "- Evaluation image count: record on completion\n", "- Ground-truth person count: record on completion\n", "- Matched persons: record on completion\n", "- Missed persons: record on completion\n", "- False-positive detections: record on completion\n", "- Detection precision and recall: record on completion\n", "- Inference latency and throughput: record on completion\n", "- Mean normalized keypoint error: record on completion\n", "- PCK@0.05, PCK@0.10, and PCK@0.20: record on completion\n", "- Feature-vector output: confirm 51 features per person (17 keypoints × 3) plus bounding box `(x1, y1, x2, y2)` matching `scripts/run.py`\n", "- Failure cases: record on completion\n", "\n", "Keep evaluator-supplied images, labels, and other private hold-off data outside the Hugging Face repository." ] }, { "cell_type": "code", "execution_count": null, "id": "0d674201", "metadata": {}, "outputs": [], "source": [] } ], "metadata": { "language_info": { "name": "python" } }, "nbformat": 4, "nbformat_minor": 5 }