Training and evaluation record
v1 provenance
- Artifact type: trained XGBoost classifier on engineered pose features β input is a 56-column pandas DataFrame, not images and not a CSV file at inference
- Model:
xgboost.XGBClassifiertrained from scratch (no YOLO fine-tuning; YOLO26x-Pose used only upstream to generate the feature CSV for training) - Upstream pose detector (data generation only, out of scope for classifier): Ultralytics YOLO-Pose (yolo26x-pose.pt), pretrained, batch_size=32, det conf 0.20, kp conf 0.25, PyTorch AMP
- CTSPL training run: Fall Detection v1 (Experiment 01)
- CTSPL fine-tuning run: none
- Training dataset:
pose_benchmark_priority1_dataset.csvβ 6,607 samples after 1:1 class balancing (58 columns total = 51 pose features + 5 additional + 1 split + 1 label = 58; XGBoost uses only 51 + 5 = 56 classifier features, columns 1β56) - Raw image source for CSV generation (upstream):
/home/ctspl/model_training/fall/version3/data(train/val/test with fall/normal subdirs) - Training hardware: CPU-only XGBoost (no GPU required for classifier); upstream YOLO extraction used CUDA
- Training date: 2026-09-21
- Selected checkpoint:
models/xgboost_priority1_fall_model.pkl(input 56 features β output Fall/Normal)
TRAINING: CCTV IMAGE β YOLO26x-Pose β 51 features β +5 β CSV (58 cols) β DataFrame β XGBoost
INFERENCE: upstream DataFrame (51 + bbox = 55 cols) β run.py: +5 β 56-col DataFrame β XGBoost (no CSV into model)
Inference input contract (what goes IN to run.py):
| Columns | Count | Source |
|---|---|---|
x_i, y_i, conf_i for i = 0β¦16 |
51 | Upstream (input contract) |
x1, y1, x2, y2 |
4 | Upstream bbox (needed to compute the 5) |
| 55 | Total input to run.py |
The extra 5 features are extracted in run.py, not received. Model input after extraction = 56 = 51 + 5.
Upstream (data generation only, out of scope for classifier inference):
- Extract 17 keypoints per detected person using YOLO-Pose (each with
x,y,confidenceβ 51 features) - Training CSV has 58 columns total β 51 pose features + 5 additional Priority 1 features + 1
split+ 1label= 58 columns
At inference, upstream delivers a pandas DataFrame with the 51 keypoint columns + bounding box; scripts/run.py computes the 5 Priority 1 features as columns and passes the 56-column DataFrame to XGBoost.
In pose_extract.py: row = norm_kpts + [aspect_ratio] + p1_features + [split, label] where norm_kpts = 51 and [aspect_ratio] + p1_features = 5.
Very important: XGBoost classifier does not use split or label as input. It uses only 51 + 5 = 56 classifier features (columns 1β56). Columns 57 β split (train/val/test) and 58 β label (0 = Normal, 1 = Fall) are metadata, not model input.
The 2 metadata columns are:
split: which dataset split the row came from (train,val,test)label: ground-truth class (0= Normal,1= Fall)
Training configuration
Upstream pose extraction (data generation only, not classifier training):
- Model: yolo26x-pose.pt (Ultralytics)
- Batch size: 32
- Detection confidence threshold: 0.20
- Keypoint confidence threshold: 0.25 (for CSV generation; inference threshold for NaN masking is 0.50)
- Compute: PyTorch AMP mixed precision (CUDA) β only for extraction
Feature engineering (output is classifier input β 51 + 5 = 56, with 58 columns total in CSV):
- 51 pose features (cols 1β51): normalized keypoint features (17 COCO joints Γ 3: x_norm, y_norm, confidence) where
x_norm=(x-x1)/w,y_norm=(y-y1)/h - 5 additional Priority 1 features (cols 52β56):
aspect_ratio= w / h (col 52)nose_relative_y= (Y_nose - y1) / h (col 53)torso_angle= angle HipMidβShoulderMid vs. Y-axis (col 54)norm_com_y= confidence-weighted center of mass Y (col 55)head_hip_v_dist= (Y_hip_mid - Y_nose) / h (col 56)
- = 56 classifier features (XGBoost input, columns 1β56)
- 1
split(col 57,train/val/test) + 1label(col 58,0/1) = 58 columns total in CSV
- 1
In pose_extract.py: row = norm_kpts + [aspect_ratio] + p1_features + [split, label] β 51 + 5 + 1 + 1 = 58. XGBoost uses only columns 1β56.
XGBoost hyperparameters:
n_estimators: 150max_depth: 5learning_rate: 0.03scale_pos_weight: 1.0 (data pre-balanced)missing: np.naneval_metric: loglossrandom_state: 42
Classifier I/O:
- Input:
(n, 56)float pandas DataFrame (one row per person); CSV never enters the model at inference - Output:
Fall/Normalviapredict/predict_proba; ruleP(Fall) >= 0.70 β Fall
Training performance record
Standard test set evaluation (56-feature test vectors):
Training completed successfully. Model saved to models/xgboost_priority1_fall_model.pkl.
--- XGBoost Classification Report (56 features β Fall/Normal) ---
precision recall f1-score support
Normal (0) 0.90 0.91 0.91 3207
Fall (1) 0.92 0.91 0.91 3400
accuracy 0.91 6607
macro avg 0.91 0.91 0.91 6607
weighted avg 0.91 0.91 0.91 6607
Overall accuracy: 91.00%
Grid search (classifier thresholds only: CONF_THRESH 0.25/0.50, FALL_PROB_THRESH 0.20β0.95) achieved best: Accuracy 92.67%, F1 90.91% at CONF=0.50, Fall_Prob=0.70. See README.md / docs/spec.md for full table.