fall-detection / docs /spec.md
nishant2401's picture
Upload folder using huggingface_hub
166743f verified
|
Raw
History Blame Contribute Delete
10.3 kB

Development specification

Scope

Fall Detection v1 classifier is a binary XGBoost model that classifies people as Fall / Normal from a 56-column pandas DataFrame. Upstream stages β€” CCTV image capture, YOLO26x-Pose (yolo26x-pose.pt) 17-keypoint extraction, and bounding-box normalization producing the 51 keypoint features + bbox as a DataFrame β€” are out of scope for this model. The 5 Priority 1 features are computed inside the inference pipeline (scripts/run.py) from that DataFrame, which is then expanded to 56 columns and passed to XGBoost (never as a CSV into the model).

Input contract (what goes IN to run.py):

Columns Count Source
x_i, y_i, conf_i for i = 0…16 51 Upstream (input contract)
x1, y1, x2, y2 4 Upstream bbox (needed to compute the 5)
55 Total input to run.py

The extra 5 features are extracted in run.py, not received. Model input after extraction = 56 = 51 + 5.

CCTV IMAGE β†’ YOLO26x-Pose β†’ 51 features + bbox (DataFrame)   upstream (out of scope)
                                      ↓
                    scripts/run.py: +5 engineered β†’ 56-col DataFrame β†’ XGBoost β†’ Fall / Normal
                                              ↑ computed in this repo      ↑ direct input (DataFrame)

The system addresses a critical domain shift: raw-coordinate baselines reached 92.22% accuracy on standard splits but collapsed to 66.83% on real-world CCTV. The 56-feature representation (51 normalized keypoints + 5 Priority 1 CCTV-invariant features) resolves this, achieving 88.36% on out-of-distribution CCTV while maintaining 91% on standard test sets.

Architecture decisions

YOLO26x-Pose is used only upstream to generate training and inference vectors, not as part of the classifier. It provides 17-keypoint COCO topology with confidence scores, enabling confidence-based filtering. Raw keypoints are transformed into normalized, camera-invariant features before reaching the classifier.

Feature Engineering Strategy: the Priority 1 set replaces camera-dependent absolute coordinates with relative geometric ratios:

  1. 51 normalized keypoint features: 17 joints Γ— (x_norm, y_norm, confidence) where norm_x = (x - x1) / w, norm_y = (y - y1) / h relative to the person's bounding box. Makes features invariant to camera distance.
  2. 5 Priority 1 CCTV-invariant features (positions 52–56 of the 56-vector):
    • aspect_ratio: w / h β€” horizontal spread. Upright < 1.0; lying > 1.0.
    • nose_relative_y: (Y_nose - y1) / h β€” head vertical position. Standing 0.0–0.25; fallen 0.7–1.0.
    • torso_angle: angle HipMidβ†’ShoulderMid vs. vertical Y-axis. Standing β‰ˆ0Β°; fallen 80–90Β°.
    • norm_com_y: (Ξ£(Y_i Γ— conf_i) / Ξ£(conf_i) - y1) / h β€” confidence-weighted vertical center-of-mass, robust to occlusion.
    • head_hip_v_dist: (Y_hip_mid - Y_nose) / h β€” upper-body extension, near 0.0 when lying flat.

Together they provide scale invariance, angle invariance, and occlusion robustness.

Classifier Selection: XGBoost was chosen over Random Forest, SVM, and MLP because it natively handles NaN via missing=np.nan, eliminating imputation for occluded keypoints (confidence < 0.50 β†’ NaN coordinates). 1:1 class balancing via undersampling prevents majority-class bias.

Starting checkpoint and training

No pose-estimator fine-tuning. The XGBoost classifier was trained from scratch on the engineered 56-feature dataset (pose_benchmark_priority1_dataset.csv):

  • Training data: 1:1 balanced train split from pose_benchmark_priority1_dataset.csv
  • Hyperparameters:
    • n_estimators=150
    • max_depth=5
    • learning_rate=0.03
    • scale_pos_weight=1.0 (data pre-balanced)
    • missing=np.nan (native occlusion handling)
    • eval_metric="logloss"
    • random_state=42

Test set was preserved intact without resampling. Upstream YOLO26x-Pose checkpoint (yolo26x-pose.pt) was used only to generate the CSV, not as a model checkpoint.

Training achieves:

  • Standard test accuracy: 91% (Precision: 0.91, Recall: 0.91, F1: 0.91)
  • Real-world accuracy: 88.36% (Precision: 0.90, Recall: 0.88, F1: 0.87) β€” +21.53 pp over raw-keypoint baseline (66.83% β†’ 88.36%)

Preprocessing and post-processing

Training Pipeline (to produce the classifier's training data)

  1. Keypoint Extraction (upstream, out of scope for classifier): Process image datasets (train/val/test with fall/normal subdirs) through YOLO26x-Pose with batch_size=32, detection confidence 0.20, keypoint confidence 0.25. This step is only to generate the dataset; the classifier does not run it.

  2. Feature Engineering (51 upstream; +5 computed in scripts/run.py at inference): For each detected person:

    • Extract bounding box (x1, y1, x2, y2), compute w, h
    • Normalize all 17 keypoints: x_norm = (x - x1) / w, y_norm = (y - y1) / h β†’ 51 features + bbox (upstream deliverable)
    • Set keypoints with confidence < 0.25 (training) / < 0.50 (inference) to NaN for x/y (confidence preserved)
    • scripts/run.py computes the 5 Priority 1 features (aspect_ratio, nose_relative_y, torso_angle, norm_com_y, head_hip_v_dist) from the 51 + bbox
    • Concatenate into 56-column DataFrame (51 + 5) in fixed order β€” this is the XGBoost input (CSV only used to transport training data into a DataFrame)

    Older docs described this as β€œ58 features (52 + 5)”; correct count is 56 = 51 (17Γ—3) + 5 (config.json, scripts/train.py).

  3. Data Balancing: Undersample majority class to 1:1 Normal:Fall in training split only.

  4. Model Training: Train XGBoost on 56 features with native NaN handling. Export to models/xgboost_priority1_fall_model.pkl.

Classifier I/O

  • Input: (n, 56) float pandas DataFrame (one row per person); NaN allowed. CSV is never fed to the model β€” upstream DataFrame β†’ run.py extracts 5 features β†’ 56-column DataFrame β†’ predict_proba.
  • Output: Fall / Normal via model.predict / model.predict_proba; decision rule P(Fall) >= FALL_PROB_THRESH.

Approaches considered

Version 1 (Raw Keypoints): Direct XGBoost on 51 raw normalized coordinates (sometimes miscounted as 52). Achieved 92.22% on standard splits but collapsed to 66.83% on CCTV due to camera sensitivity and lack of posture ratios.

Alternative Feature Sets: Priority 2 (joint velocities, inter-joint distances) and Priority 3 (full skeleton angles, convex hull) were considered but excluded β€” the minimal 5 Priority 1 features already resolved the domain shift.

Alternative Classifiers: Random Forest, SVM (RBF), MLP (64-32) were benchmarked. XGBoost outperformed and avoided imputation pipelines.

Temporal Models: LSTM/GRU and 10-frame buffers excluded as out of scope for the classifier; deployment-level smoothing is separate.

Known design gaps

The classifier system does not include:

  1. Multi-person tracking: each 56-vector is classified independently; identity tracking would need deployment logic.
  2. Environmental context: ignores scene context (bed vs. floor, stairs).
  3. Performance profiling: per-sample latency measured; multi-camera throughput not benchmarked.
  4. Confidence calibration: FALL_PROB_THRESH = 0.70 chosen via grid search; per-environment calibration recommended.
  5. Adversarial validation: not tested against yoga/exercise floor activities β€” future β€œnon-fall floor activity” class needed.

Dataset structure

Training/Validation/Test images (upstream only): data/{train,val,test}/{fall,normal}/*.{jpg,jpeg,png,bmp} β€” used only to generate the feature CSV.

Classifier training file: pose_benchmark_priority1_dataset.csv β€” 58 columns total:

  • 51 pose features (17 joints Γ— 3: x_norm, y_norm, confidence) β†’ columns 1–51
    • 5 additional Priority 1 features (aspect_ratio, nose_relative_y, torso_angle, norm_com_y, head_hip_v_dist) β†’ columns 52–56 (56 classifier features)
    • 1 split (metadata: train / val / test) β†’ column 57
    • 1 label (0 = Normal, 1 = Fall) β†’ column 58
  • = 58 columns

In pose_extract.py: row = norm_kpts + [aspect_ratio] + p1_features + [split, label] where norm_kpts = 51 and [aspect_ratio] + p1_features = 5.

Very important: XGBoost uses only 51 + 5 = 56 classifier features (columns 1–56). Columns 57 β†’ split and 58 β†’ label are metadata, not model input.

Real-World Evaluation: out-of-distribution CCTV-derived feature vectors at /home/ctspl/model_training/fall/version3/data, evaluated separately.

Deployment configuration

Hardware: CPU-only for classifier; no GPU required. Upstream YOLO stage (if run locally for data generation) benefits from CUDA, but classifier inference/training does not need it.

Thresholds (tunable via config.json):

  • CONF_THRESH = 0.50 β€” Minimum keypoint confidence; below β†’ NaN coordinates in 56-vector
  • FALL_PROB_THRESH = 0.70 β€” Minimum XGBoost P(Fall) for Fall classification (optimized from 0.95)

YOLO_CONF_THRESH is an upstream data-generation parameter (0.25 optimized) and not a classifier threshold.

Grid Search Optimization Results

Grid search over classifier thresholds (YOLO stage excluded):

  • CONF_THRESH (keypoint): [0.30, 0.40, 0.50]
  • FALL_PROB_THRESH: [0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 0.95]

Top Parameter Combinations (By F1-Score):

Keypoint_Conf Fall_Prob_Thresh Accuracy F1_Score Precision Recall TN FP FN TP
0.5 0.70 92.67 90.91 90.11 91.73 390 28 23 255
0.5 0.80 92.67 90.68 92.19 89.21 397 21 30 248
0.5 0.60 92.10 90.30 88.58 92.09 385 33 22 256
0.5 0.90 92.53 90.04 96.31 84.53 409 9 43 235
0.5 0.50 91.24 89.43 86.29 92.81 377 41 20 258

Best Configuration (selected):

  • CONF_THRESH: 0.50
  • FALL_PROB_THRESH: 0.70

Accuracy 92.67%, F1 90.91%, Precision 90.11%, Recall 91.73%.