{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# Fall Detection — Evaluation (Input Contract: 51 Features + BBox)\n", "\n", "Evaluate the trained XGBoost classifier following the **input contract**.\n", "\n", "**Input contract (what goes IN — only this much):**\n", "```\n", "51 keypoint features: x_0,y_0,conf_0, ..., x_16,y_16,conf_16\n", "+ bbox: x1, y1, x2, y2\n", "= 55 input columns (pandas DataFrame)\n", "```\n", "The **extra 5 features are NOT input** — they are extracted inside `scripts/run.py`.\n", "\n", "**Flow:**\n", "```\n", "INPUT: upstream DataFrame (51 features + bbox) # input contract only\n", " ↓\n", "EXTRACT: run.py: build_feature_frame() → +5 features # done here, not received\n", " ↓\n", "MODEL: 56-column DataFrame (51 + 5) → XGBoost.predict_proba(DataFrame)\n", " ↓\n", "OUTPUT: Fall / Normal + fall_probability\n", "```\n", "\n", "Upstream (out of scope): CCTV image → YOLO26x-Pose → 17 keypoints → 51 + bbox." ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "The notebook has 8 cells:\n", "1. Markdown intro — input contract (this cell)\n", "2. Flow overview\n", "3. Imports (pulls input-contract constants from `run.py`)\n", "4. Model loading (XGBoost only — YOLO is upstream/out of scope)\n", "5. Load **input** DataFrame (51 features + bbox only)\n", "6. **Extract** the extra 5 in `run.py` → 56-column DataFrame\n", "7. Run inference (56-col DataFrame → XGBoost)\n", "8. Results summary" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from pathlib import Path\n", "import sys\n", "\n", "import numpy as np\n", "import pandas as pd\n", "import matplotlib.pyplot as plt\n", "\n", "# Reuse pipeline functions + input-contract constants from scripts/run.py\n", "SCRIPTS_DIR = Path.cwd().parent if Path.cwd().name == 'scripts' else Path.cwd() / 'scripts'\n", "if str(SCRIPTS_DIR) not in sys.path:\n", " sys.path.insert(0, str(SCRIPTS_DIR))\n", "\n", "from run import (\n", " load_model,\n", " build_feature_frame,\n", " predict,\n", " # Input contract (what goes IN)\n", " INPUT_COLUMNS,\n", " INPUT_FEATURE_COUNT,\n", " KEYPOINT_COLUMNS,\n", " BBOX_COLUMNS,\n", " # Extracted in run.py (not input)\n", " ENGINEERED_FEATURE_NAMES,\n", " ENGINEERED_COUNT,\n", " # Model input (after extraction)\n", " FEATURE_COLS,\n", " MODEL_FEATURE_COUNT,\n", " CONF_THRESH,\n", " FALL_PROB_THRESH,\n", ")\n", "\n", "print('Imports ready.')\n", "print(f'INPUT CONTRACT : {INPUT_FEATURE_COUNT} keypoint features + {len(BBOX_COLUMNS)} bbox '\n", " f'= {len(INPUT_COLUMNS)} columns (the extra {ENGINEERED_COUNT} are NOT input)')\n", "print(f'EXTRACTED HERE : {ENGINEERED_COUNT} features -> {ENGINEERED_FEATURE_NAMES}')\n", "print(f'MODEL INPUT : {MODEL_FEATURE_COUNT} columns (DataFrame = '\n", " f'{INPUT_FEATURE_COUNT} input + {ENGINEERED_COUNT} extracted)')\n", "print(f'Thresholds : CONF_THRESH={CONF_THRESH}, FALL_PROB_THRESH={FALL_PROB_THRESH}')" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# Load XGBoost classifier only (YOLO/Pose is upstream and out of scope)\n", "model = load_model()\n", "print('XGBoost ready — expects a 56-column DataFrame (51 input + 5 extracted in run.py).')" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# Load INPUT DataFrame — input contract only: 51 keypoint features + bbox\n", "# Required columns: x_0,y_0,conf_0, ..., x_16,y_16,conf_16, x1,y1,x2,y2\n", "# The extra 5 (aspect_ratio, ...) are NOT part of the input.\n", "import os\n", "\n", "UPSTREAM_CSV = os.environ.get(\n", " 'UPSTREAM_FEATURES_CSV',\n", " '../upstream_features.csv', # true upstream export: 51 + bbox\n", ")\n", "\n", "input_df = pd.read_csv(UPSTREAM_CSV)\n", "\n", "# Validate input contract: 51 keypoint features + bbox must be present\n", "missing = [c for c in INPUT_COLUMNS if c not in input_df.columns]\n", "if missing:\n", " raise ValueError(\n", " f'Input contract violated — missing {len(missing)} column(s), e.g. {missing[:6]}. '\n", " f'Expected {INPUT_FEATURE_COUNT} keypoint features (x_0..conf_16) + bbox (x1,y1,x2,y2) only. '\n", " 'The extra 5 features are extracted by run.py, not required as input.'\n", " )\n", "\n", "# If a training CSV (56 feats + split/label, no bbox) was provided, say so clearly\n", "if set(['split', 'label']).issubset(input_df.columns) and not set(BBOX_COLUMNS).issubset(input_df.columns):\n", " raise ValueError(\n", " 'This looks like the training CSV (has split/label, no bbox). '\n", " 'Evaluation follows the input contract: provide 51 keypoint features + bbox only '\n", " '(upstream export), so run.py can extract the extra 5 itself.'\n", " )\n", "\n", "print(f'INPUT DataFrame (contract): {input_df.shape[0]} rows × {input_df.shape[1]} columns')\n", "print(f' - Keypoint features: {INPUT_FEATURE_COUNT}')\n", "print(f' - Bbox columns: {BBOX_COLUMNS}')\n", "print(f' - Engineered features present in input: '\n", " f'{[c for c in ENGINEERED_FEATURE_NAMES if c in input_df.columns] or \"none (correct — extracted in run.py)\"}')\n", "input_df.head()" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# EXTRACT the extra 5 here (in run.py) → 56-column model DataFrame\n", "# Input stays 51 + bbox; the 5 are computed, never received.\n", "feature_df, skip_rows = build_feature_frame(input_df)\n", "\n", "assert isinstance(feature_df, pd.DataFrame)\n", "assert list(feature_df.columns) == FEATURE_COLS\n", "assert feature_df.shape[1] == MODEL_FEATURE_COUNT # 56 = 51 + 5\n", "assert feature_df.shape[1] == INPUT_FEATURE_COUNT + ENGINEERED_COUNT\n", "\n", "print(f'EXTRACTED in run.py ({ENGINEERED_COUNT}): {ENGINEERED_FEATURE_NAMES}')\n", "print(f'Model DataFrame: {feature_df.shape[0]} rows × {feature_df.shape[1]} columns '\n", " f'({INPUT_FEATURE_COUNT} input + {ENGINEERED_COUNT} extracted)')\n", "print(f'Skipped rows (<5 valid keypoints): {skip_rows}')\n", "feature_df[ENGINEERED_FEATURE_NAMES].head()" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# MODEL: 56-column DataFrame → XGBoost (in-memory; never CSV)\n", "pred_df = predict(model, feature_df, skip_rows)\n", "\n", "results = pd.concat([input_df, pred_df], axis=1)\n", "results[['prediction', 'fall_probability']].head(10)" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# Results summary\n", "scorable = results['fall_probability'].notna()\n", "fall_count = int((results['prediction'] == 'Fall').sum())\n", "normal_count = int((results['prediction'] == 'Normal').sum())\n", "\n", "print('=' * 50)\n", "print(' DETECTION SUMMARY')\n", "print('=' * 50)\n", "print('Input contract : 51 keypoint features + bbox only')\n", "print(f'Extracted here : {ENGINEERED_COUNT} features in run.py')\n", "print(f'Model input : pandas.DataFrame {feature_df.shape[0]} × {feature_df.shape[1]}')\n", "print(f'Total rows : {len(results)}')\n", "print(f' - Skipped : {len(skip_rows)}')\n", "print(f' - Fall : {fall_count}')\n", "print(f' - Normal : {normal_count}')\n", "\n", "if scorable.any():\n", " print(f' - Max P(Fall): {results.loc[scorable, \"fall_probability\"].max():.2%}')\n", "\n", "print(f'\\nCONF_THRESH={CONF_THRESH}, FALL_PROB_THRESH={FALL_PROB_THRESH}')\n", "print('=' * 50)\n", "\n", "# Probability distribution plot\n", "if scorable.any():\n", " fig, axes = plt.subplots(1, 2, figsize=(12, 4))\n", "\n", " axes[0].hist(results.loc[scorable, 'fall_probability'], bins=20, color='steelblue', edgecolor='black')\n", " axes[0].axvline(FALL_PROB_THRESH, color='red', linestyle='--', label=f'threshold={FALL_PROB_THRESH}')\n", " axes[0].set_xlabel('P(Fall)')\n", " axes[0].set_ylabel('Count')\n", " axes[0].set_title('Fall probability distribution')\n", " axes[0].legend()\n", "\n", " counts = [normal_count, fall_count]\n", " axes[1].bar(['Normal', 'Fall'], counts, color=['green', 'red'], edgecolor='black')\n", " axes[1].set_ylabel('Count')\n", " axes[1].set_title('Prediction counts')\n", "\n", " plt.tight_layout()\n", " plt.show()" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.10.0" } }, "nbformat": 4, "nbformat_minor": 4 }