File size: 9,700 Bytes
166743f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Fall Detection β€” Evaluation (Input Contract: 51 Features + BBox)\n",
    "\n",
    "Evaluate the trained XGBoost classifier following the **input contract**.\n",
    "\n",
    "**Input contract (what goes IN β€” only this much):**\n",
    "```\n",
    "51 keypoint features: x_0,y_0,conf_0, ..., x_16,y_16,conf_16\n",
    "+ bbox:               x1, y1, x2, y2\n",
    "= 55 input columns    (pandas DataFrame)\n",
    "```\n",
    "The **extra 5 features are NOT input** β€” they are extracted inside `scripts/run.py`.\n",
    "\n",
    "**Flow:**\n",
    "```\n",
    "INPUT:  upstream DataFrame (51 features + bbox)          # input contract only\n",
    "            ↓\n",
    "EXTRACT: run.py: build_feature_frame()  β†’ +5 features   # done here, not received\n",
    "            ↓\n",
    "MODEL:   56-column DataFrame (51 + 5) β†’ XGBoost.predict_proba(DataFrame)\n",
    "            ↓\n",
    "OUTPUT:  Fall / Normal + fall_probability\n",
    "```\n",
    "\n",
    "Upstream (out of scope): CCTV image β†’ YOLO26x-Pose β†’ 17 keypoints β†’ 51 + bbox."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "The notebook has 8 cells:\n",
    "1. Markdown intro β€” input contract (this cell)\n",
    "2. Flow overview\n",
    "3. Imports (pulls input-contract constants from `run.py`)\n",
    "4. Model loading (XGBoost only β€” YOLO is upstream/out of scope)\n",
    "5. Load **input** DataFrame (51 features + bbox only)\n",
    "6. **Extract** the extra 5 in `run.py` β†’ 56-column DataFrame\n",
    "7. Run inference (56-col DataFrame β†’ XGBoost)\n",
    "8. Results summary"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from pathlib import Path\n",
    "import sys\n",
    "\n",
    "import numpy as np\n",
    "import pandas as pd\n",
    "import matplotlib.pyplot as plt\n",
    "\n",
    "# Reuse pipeline functions + input-contract constants from scripts/run.py\n",
    "SCRIPTS_DIR = Path.cwd().parent if Path.cwd().name == 'scripts' else Path.cwd() / 'scripts'\n",
    "if str(SCRIPTS_DIR) not in sys.path:\n",
    "    sys.path.insert(0, str(SCRIPTS_DIR))\n",
    "\n",
    "from run import (\n",
    "    load_model,\n",
    "    build_feature_frame,\n",
    "    predict,\n",
    "    # Input contract (what goes IN)\n",
    "    INPUT_COLUMNS,\n",
    "    INPUT_FEATURE_COUNT,\n",
    "    KEYPOINT_COLUMNS,\n",
    "    BBOX_COLUMNS,\n",
    "    # Extracted in run.py (not input)\n",
    "    ENGINEERED_FEATURE_NAMES,\n",
    "    ENGINEERED_COUNT,\n",
    "    # Model input (after extraction)\n",
    "    FEATURE_COLS,\n",
    "    MODEL_FEATURE_COUNT,\n",
    "    CONF_THRESH,\n",
    "    FALL_PROB_THRESH,\n",
    ")\n",
    "\n",
    "print('Imports ready.')\n",
    "print(f'INPUT CONTRACT : {INPUT_FEATURE_COUNT} keypoint features + {len(BBOX_COLUMNS)} bbox '\n",
    "      f'= {len(INPUT_COLUMNS)} columns (the extra {ENGINEERED_COUNT} are NOT input)')\n",
    "print(f'EXTRACTED HERE : {ENGINEERED_COUNT} features -> {ENGINEERED_FEATURE_NAMES}')\n",
    "print(f'MODEL INPUT    : {MODEL_FEATURE_COUNT} columns (DataFrame = '\n",
    "      f'{INPUT_FEATURE_COUNT} input + {ENGINEERED_COUNT} extracted)')\n",
    "print(f'Thresholds     : CONF_THRESH={CONF_THRESH}, FALL_PROB_THRESH={FALL_PROB_THRESH}')"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Load XGBoost classifier only (YOLO/Pose is upstream and out of scope)\n",
    "model = load_model()\n",
    "print('XGBoost ready β€” expects a 56-column DataFrame (51 input + 5 extracted in run.py).')"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Load INPUT DataFrame β€” input contract only: 51 keypoint features + bbox\n",
    "# Required columns: x_0,y_0,conf_0, ..., x_16,y_16,conf_16, x1,y1,x2,y2\n",
    "# The extra 5 (aspect_ratio, ...) are NOT part of the input.\n",
    "import os\n",
    "\n",
    "UPSTREAM_CSV = os.environ.get(\n",
    "    'UPSTREAM_FEATURES_CSV',\n",
    "    '../upstream_features.csv',  # true upstream export: 51 + bbox\n",
    ")\n",
    "\n",
    "input_df = pd.read_csv(UPSTREAM_CSV)\n",
    "\n",
    "# Validate input contract: 51 keypoint features + bbox must be present\n",
    "missing = [c for c in INPUT_COLUMNS if c not in input_df.columns]\n",
    "if missing:\n",
    "    raise ValueError(\n",
    "        f'Input contract violated β€” missing {len(missing)} column(s), e.g. {missing[:6]}. '\n",
    "        f'Expected {INPUT_FEATURE_COUNT} keypoint features (x_0..conf_16) + bbox (x1,y1,x2,y2) only. '\n",
    "        'The extra 5 features are extracted by run.py, not required as input.'\n",
    "    )\n",
    "\n",
    "# If a training CSV (56 feats + split/label, no bbox) was provided, say so clearly\n",
    "if set(['split', 'label']).issubset(input_df.columns) and not set(BBOX_COLUMNS).issubset(input_df.columns):\n",
    "    raise ValueError(\n",
    "        'This looks like the training CSV (has split/label, no bbox). '\n",
    "        'Evaluation follows the input contract: provide 51 keypoint features + bbox only '\n",
    "        '(upstream export), so run.py can extract the extra 5 itself.'\n",
    "    )\n",
    "\n",
    "print(f'INPUT DataFrame (contract): {input_df.shape[0]} rows Γ— {input_df.shape[1]} columns')\n",
    "print(f'  - Keypoint features: {INPUT_FEATURE_COUNT}')\n",
    "print(f'  - Bbox columns: {BBOX_COLUMNS}')\n",
    "print(f'  - Engineered features present in input: '\n",
    "      f'{[c for c in ENGINEERED_FEATURE_NAMES if c in input_df.columns] or \"none (correct β€” extracted in run.py)\"}')\n",
    "input_df.head()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# EXTRACT the extra 5 here (in run.py) β†’ 56-column model DataFrame\n",
    "# Input stays 51 + bbox; the 5 are computed, never received.\n",
    "feature_df, skip_rows = build_feature_frame(input_df)\n",
    "\n",
    "assert isinstance(feature_df, pd.DataFrame)\n",
    "assert list(feature_df.columns) == FEATURE_COLS\n",
    "assert feature_df.shape[1] == MODEL_FEATURE_COUNT  # 56 = 51 + 5\n",
    "assert feature_df.shape[1] == INPUT_FEATURE_COUNT + ENGINEERED_COUNT\n",
    "\n",
    "print(f'EXTRACTED in run.py ({ENGINEERED_COUNT}): {ENGINEERED_FEATURE_NAMES}')\n",
    "print(f'Model DataFrame: {feature_df.shape[0]} rows Γ— {feature_df.shape[1]} columns '\n",
    "      f'({INPUT_FEATURE_COUNT} input + {ENGINEERED_COUNT} extracted)')\n",
    "print(f'Skipped rows (<5 valid keypoints): {skip_rows}')\n",
    "feature_df[ENGINEERED_FEATURE_NAMES].head()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# MODEL: 56-column DataFrame β†’ XGBoost (in-memory; never CSV)\n",
    "pred_df = predict(model, feature_df, skip_rows)\n",
    "\n",
    "results = pd.concat([input_df, pred_df], axis=1)\n",
    "results[['prediction', 'fall_probability']].head(10)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Results summary\n",
    "scorable = results['fall_probability'].notna()\n",
    "fall_count = int((results['prediction'] == 'Fall').sum())\n",
    "normal_count = int((results['prediction'] == 'Normal').sum())\n",
    "\n",
    "print('=' * 50)\n",
    "print('         DETECTION SUMMARY')\n",
    "print('=' * 50)\n",
    "print('Input contract : 51 keypoint features + bbox only')\n",
    "print(f'Extracted here : {ENGINEERED_COUNT} features in run.py')\n",
    "print(f'Model input    : pandas.DataFrame {feature_df.shape[0]} Γ— {feature_df.shape[1]}')\n",
    "print(f'Total rows     : {len(results)}')\n",
    "print(f'  - Skipped    : {len(skip_rows)}')\n",
    "print(f'  - Fall       : {fall_count}')\n",
    "print(f'  - Normal     : {normal_count}')\n",
    "\n",
    "if scorable.any():\n",
    "    print(f'  - Max P(Fall): {results.loc[scorable, \"fall_probability\"].max():.2%}')\n",
    "\n",
    "print(f'\\nCONF_THRESH={CONF_THRESH}, FALL_PROB_THRESH={FALL_PROB_THRESH}')\n",
    "print('=' * 50)\n",
    "\n",
    "# Probability distribution plot\n",
    "if scorable.any():\n",
    "    fig, axes = plt.subplots(1, 2, figsize=(12, 4))\n",
    "\n",
    "    axes[0].hist(results.loc[scorable, 'fall_probability'], bins=20, color='steelblue', edgecolor='black')\n",
    "    axes[0].axvline(FALL_PROB_THRESH, color='red', linestyle='--', label=f'threshold={FALL_PROB_THRESH}')\n",
    "    axes[0].set_xlabel('P(Fall)')\n",
    "    axes[0].set_ylabel('Count')\n",
    "    axes[0].set_title('Fall probability distribution')\n",
    "    axes[0].legend()\n",
    "\n",
    "    counts = [normal_count, fall_count]\n",
    "    axes[1].bar(['Normal', 'Fall'], counts, color=['green', 'red'], edgecolor='black')\n",
    "    axes[1].set_ylabel('Count')\n",
    "    axes[1].set_title('Prediction counts')\n",
    "\n",
    "    plt.tight_layout()\n",
    "    plt.show()"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.10.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}