Scikit-learn
human-activity-recognition
wearable
wrist
time-series
cpu
scikit-learn
WISP / docs /EVALUATION.md
Zipeng365's picture
Add files using upload-large-folder tool
10ef792 verified
|
Raw History Blame Contribute Delete
2.93 kB

Evaluation and result authority

wisp evaluate loads a trusted fitted checkpoint and predicts the supplied labelled NPZ. It does not fit, search, tune a threshold, choose a fold/seed or select a model. It writes JSON metrics plus a confusion matrix whose rows and columns follow the complete global class vocabulary.

The four primary metrics follow the corrected frozen paper scorer:

  • Macro-F1: average F1 over classes observed in the ground truth or prediction; a class absent from both is not included in this average.
  • Accuracy: correctly classified windows divided by all evaluated windows.
  • Balanced accuracy: mean recall over classes present in the ground truth.
  • Worst-class F1: minimum F1 over the complete global vocabulary; an absent class has F1 zero.

Macro-F1 and worst-class F1 therefore do not necessarily average/minimize over the same class subset. Preserve the global class count even if a held-out partition omits one activity. Do not compare outputs from a differently defined scorer as though they were the same metric.

The inventory audit found 31 discrepancies between retained models' original adjacent metric metadata and the corrected frozen result records. These historical files are provenance and should not be silently edited. The corrected frozen four-metric catalogue in results/run_results.csv is the authority for reported paper values; wisp_release.scoring implements its definitions for new evaluation. A checkpoint's sibling metric JSON is not automatically the final leaderboard truth, especially for worst-class F1.

Checkpoint loading, prediction equivalence, input export and metric equivalence are different checks. Passing one does not establish the others. The root release audit states which files, datasets and folds were actually evaluated and whether their predictions/metrics match the frozen evidence. An inventory entry alone is not a successful inference test.

The recorded model gate loaded and smoke-tested all 360 retained checkpoints, and rescored their 360 immutable saved test-prediction records against the frozen four-metric catalogue. Full real-input inference was rerun for only six ADL-Wrist fold-0 checkpoints: Random/Evolution with seeds 42, 43 and 44. Each reproduced all 861 held-out predictions exactly both on the original frozen NPZ and on a fresh pinned benchmark export. The other 354 full test splits were not rerun in this gate.

ADL's original time coordinates are seconds; the fresh Hub export stores window ordinals. The numerical values therefore differ. Input signals, participant/label rows and within-participant stable chronological permutations matched; the decoder uses that ordering rather than elapsed-time gaps. Both input representations were evaluated explicitly and produced the same locked predictions. This does not authorize ignoring time metadata differences for another dataset or another decoder.