# Development specification ## Scope Fall Detection v1 classifier is a binary XGBoost model that classifies people as `Fall` / `Normal` from a **56-column pandas DataFrame**. Upstream stages — CCTV image capture, YOLO26x-Pose (yolo26x-pose.pt) 17-keypoint extraction, and bounding-box normalization producing the **51 keypoint features + bbox** as a DataFrame — are **out of scope for this model**. The **5 Priority 1 features are computed inside the inference pipeline** (`scripts/run.py`) from that DataFrame, which is then expanded to 56 columns and passed to XGBoost (never as a CSV into the model). **Input contract (what goes IN to `run.py`):** | Columns | Count | Source | | --- | ---: | --- | | `x_i, y_i, conf_i` for i = 0…16 | **51** | Upstream (input contract) | | `x1, y1, x2, y2` | 4 | Upstream bbox (needed to compute the 5) | | | **55** | **Total input to `run.py`** | The **extra 5 features are extracted in `run.py`**, not received. Model input after extraction = **56 = 51 + 5**. ``` CCTV IMAGE → YOLO26x-Pose → 51 features + bbox (DataFrame) upstream (out of scope) ↓ scripts/run.py: +5 engineered → 56-col DataFrame → XGBoost → Fall / Normal ↑ computed in this repo ↑ direct input (DataFrame) ``` The system addresses a critical domain shift: raw-coordinate baselines reached 92.22% accuracy on standard splits but collapsed to 66.83% on real-world CCTV. The 56-feature representation (51 normalized keypoints + 5 Priority 1 CCTV-invariant features) resolves this, achieving 88.36% on out-of-distribution CCTV while maintaining 91% on standard test sets. ## Architecture decisions YOLO26x-Pose is used **only upstream to generate training and inference vectors**, not as part of the classifier. It provides 17-keypoint COCO topology with confidence scores, enabling confidence-based filtering. Raw keypoints are transformed into normalized, camera-invariant features before reaching the classifier. **Feature Engineering Strategy**: the Priority 1 set replaces camera-dependent absolute coordinates with relative geometric ratios: 1. **51 normalized keypoint features**: 17 joints × (`x_norm`, `y_norm`, `confidence`) where `norm_x = (x - x1) / w`, `norm_y = (y - y1) / h` relative to the person's bounding box. Makes features invariant to camera distance. 2. **5 Priority 1 CCTV-invariant features** (positions 52–56 of the 56-vector): - **aspect_ratio**: `w / h` — horizontal spread. Upright < 1.0; lying > 1.0. - **nose_relative_y**: `(Y_nose - y1) / h` — head vertical position. Standing 0.0–0.25; fallen 0.7–1.0. - **torso_angle**: angle HipMid→ShoulderMid vs. vertical Y-axis. Standing ≈0°; fallen 80–90°. - **norm_com_y**: `(Σ(Y_i × conf_i) / Σ(conf_i) - y1) / h` — confidence-weighted vertical center-of-mass, robust to occlusion. - **head_hip_v_dist**: `(Y_hip_mid - Y_nose) / h` — upper-body extension, near 0.0 when lying flat. Together they provide scale invariance, angle invariance, and occlusion robustness. **Classifier Selection**: XGBoost was chosen over Random Forest, SVM, and MLP because it natively handles `NaN` via `missing=np.nan`, eliminating imputation for occluded keypoints (confidence < 0.50 → `NaN` coordinates). 1:1 class balancing via undersampling prevents majority-class bias. ## Starting checkpoint and training No pose-estimator fine-tuning. The XGBoost classifier was trained **from scratch** on the engineered 56-feature dataset (`pose_benchmark_priority1_dataset.csv`): - **Training data**: 1:1 balanced train split from `pose_benchmark_priority1_dataset.csv` - **Hyperparameters**: - `n_estimators=150` - `max_depth=5` - `learning_rate=0.03` - `scale_pos_weight=1.0` (data pre-balanced) - `missing=np.nan` (native occlusion handling) - `eval_metric="logloss"` - `random_state=42` Test set was preserved intact without resampling. Upstream YOLO26x-Pose checkpoint (`yolo26x-pose.pt`) was used only to generate the CSV, not as a model checkpoint. Training achieves: - **Standard test accuracy**: 91% (Precision: 0.91, Recall: 0.91, F1: 0.91) - **Real-world accuracy**: 88.36% (Precision: 0.90, Recall: 0.88, F1: 0.87) — +21.53 pp over raw-keypoint baseline (66.83% → 88.36%) ## Preprocessing and post-processing ### Training Pipeline (to produce the classifier's training data) 1. **Keypoint Extraction (upstream, out of scope for classifier)**: Process image datasets (train/val/test with fall/normal subdirs) through YOLO26x-Pose with batch_size=32, detection confidence 0.20, keypoint confidence 0.25. This step is only to generate the dataset; the classifier does not run it. 2. **Feature Engineering (51 upstream; +5 computed in `scripts/run.py` at inference)**: For each detected person: - Extract bounding box `(x1, y1, x2, y2)`, compute `w`, `h` - Normalize all 17 keypoints: `x_norm = (x - x1) / w`, `y_norm = (y - y1) / h` → **51 features + bbox** (upstream deliverable) - Set keypoints with `confidence < 0.25` (training) / `< 0.50` (inference) to `NaN` for `x`/`y` (confidence preserved) - **`scripts/run.py` computes** the 5 Priority 1 features (`aspect_ratio`, `nose_relative_y`, `torso_angle`, `norm_com_y`, `head_hip_v_dist`) from the 51 + bbox - Concatenate into **56-column DataFrame (51 + 5)** in fixed order — this is the XGBoost input (CSV only used to transport training data into a DataFrame) Older docs described this as “58 features (52 + 5)”; correct count is **56 = 51 (17×3) + 5** (`config.json`, `scripts/train.py`). 3. **Data Balancing**: Undersample majority class to 1:1 Normal:Fall in training split only. 4. **Model Training**: Train XGBoost on 56 features with native NaN handling. Export to `models/xgboost_priority1_fall_model.pkl`. ### Classifier I/O - **Input**: `(n, 56)` float **pandas DataFrame** (one row per person); `NaN` allowed. CSV is never fed to the model — upstream DataFrame → `run.py` extracts 5 features → 56-column DataFrame → `predict_proba`. - **Output**: `Fall` / `Normal` via `model.predict` / `model.predict_proba`; decision rule `P(Fall) >= FALL_PROB_THRESH`. ## Approaches considered **Version 1 (Raw Keypoints)**: Direct XGBoost on 51 raw normalized coordinates (sometimes miscounted as 52). Achieved 92.22% on standard splits but collapsed to 66.83% on CCTV due to camera sensitivity and lack of posture ratios. **Alternative Feature Sets**: Priority 2 (joint velocities, inter-joint distances) and Priority 3 (full skeleton angles, convex hull) were considered but excluded — the minimal 5 Priority 1 features already resolved the domain shift. **Alternative Classifiers**: Random Forest, SVM (RBF), MLP (64-32) were benchmarked. XGBoost outperformed and avoided imputation pipelines. **Temporal Models**: LSTM/GRU and 10-frame buffers excluded as out of scope for the classifier; deployment-level smoothing is separate. ## Known design gaps The classifier system does not include: 1. **Multi-person tracking**: each 56-vector is classified independently; identity tracking would need deployment logic. 2. **Environmental context**: ignores scene context (bed vs. floor, stairs). 3. **Performance profiling**: per-sample latency measured; multi-camera throughput not benchmarked. 4. **Confidence calibration**: `FALL_PROB_THRESH = 0.70` chosen via grid search; per-environment calibration recommended. 5. **Adversarial validation**: not tested against yoga/exercise floor activities — future “non-fall floor activity” class needed. ## Dataset structure **Training/Validation/Test images (upstream only)**: `data/{train,val,test}/{fall,normal}/*.{jpg,jpeg,png,bmp}` — used only to generate the feature CSV. **Classifier training file**: `pose_benchmark_priority1_dataset.csv` — 58 columns total: - 51 pose features (17 joints × 3: x_norm, y_norm, confidence) → columns 1–51 - + 5 additional Priority 1 features (`aspect_ratio`, `nose_relative_y`, `torso_angle`, `norm_com_y`, `head_hip_v_dist`) → columns 52–56 (**56 classifier features**) - + 1 `split` (metadata: `train` / `val` / `test`) → column 57 - + 1 `label` (`0` = Normal, `1` = Fall) → column 58 - = **58 columns** In `pose_extract.py`: `row = norm_kpts + [aspect_ratio] + p1_features + [split, label]` where `norm_kpts` = 51 and `[aspect_ratio] + p1_features` = 5. **Very important:** XGBoost uses only **51 + 5 = 56 classifier features** (columns 1–56). Columns 57 → `split` and 58 → `label` are **metadata, not model input**. **Real-World Evaluation**: out-of-distribution CCTV-derived feature vectors at `/home/ctspl/model_training/fall/version3/data`, evaluated separately. ## Deployment configuration **Hardware**: CPU-only for classifier; no GPU required. Upstream YOLO stage (if run locally for data generation) benefits from CUDA, but classifier inference/training does not need it. **Thresholds** (tunable via `config.json`): - `CONF_THRESH = 0.50` — Minimum keypoint confidence; below → `NaN` coordinates in 56-vector - `FALL_PROB_THRESH = 0.70` — Minimum XGBoost `P(Fall)` for Fall classification (optimized from 0.95) `YOLO_CONF_THRESH` is an upstream data-generation parameter (0.25 optimized) and **not a classifier threshold**. ### Grid Search Optimization Results Grid search over classifier thresholds (YOLO stage excluded): - `CONF_THRESH` (keypoint): [0.30, 0.40, 0.50] - `FALL_PROB_THRESH`: [0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 0.95] **Top Parameter Combinations (By F1-Score)**: | Keypoint_Conf | Fall_Prob_Thresh | Accuracy | F1_Score | Precision | Recall | TN | FP | FN | TP | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 0.5 | 0.70 | 92.67 | 90.91 | 90.11 | 91.73 | 390 | 28 | 23 | 255 | | 0.5 | 0.80 | 92.67 | 90.68 | 92.19 | 89.21 | 397 | 21 | 30 | 248 | | 0.5 | 0.60 | 92.10 | 90.30 | 88.58 | 92.09 | 385 | 33 | 22 | 256 | | 0.5 | 0.90 | 92.53 | 90.04 | 96.31 | 84.53 | 409 | 9 | 43 | 235 | | 0.5 | 0.50 | 91.24 | 89.43 | 86.29 | 92.81 | 377 | 41 | 20 | 258 | **Best Configuration** (selected): - CONF_THRESH: 0.50 - FALL_PROB_THRESH: 0.70 Accuracy 92.67%, F1 90.91%, Precision 90.11%, Recall 91.73%.