Instructions to use Zipeng365/WISP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use Zipeng365/WISP with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("Zipeng365/WISP", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
Download docs/EVALUATION.md from Zipeng365/WISP: direct link, hf CLI and curl.
- Browser
- Download file 2.93 kB
-
https://huggingface.co/Zipeng365/WISP/resolve/main/docs/EVALUATION.md
- Command line
-
hf download hf://Zipeng365/WISP/docs/EVALUATION.md
-
curl -L -o EVALUATION.md https://huggingface.co/Zipeng365/WISP/resolve/main/docs/EVALUATION.md
Evaluation and result authority
wisp evaluate loads a trusted fitted checkpoint and predicts the supplied labelled NPZ. It does not fit, search, tune a threshold, choose a fold/seed or select a model. It writes JSON metrics plus a confusion matrix whose rows and columns follow the complete global class vocabulary.
The four primary metrics follow the corrected frozen paper scorer:
- Macro-F1: average F1 over classes observed in the ground truth or prediction; a class absent from both is not included in this average.
- Accuracy: correctly classified windows divided by all evaluated windows.
- Balanced accuracy: mean recall over classes present in the ground truth.
- Worst-class F1: minimum F1 over the complete global vocabulary; an absent class has F1 zero.
Macro-F1 and worst-class F1 therefore do not necessarily average/minimize over the same class subset. Preserve the global class count even if a held-out partition omits one activity. Do not compare outputs from a differently defined scorer as though they were the same metric.
The inventory audit found 31 discrepancies between retained models' original adjacent metric metadata and the corrected frozen result records. These historical files are provenance and should not be silently edited. The corrected frozen four-metric catalogue in results/run_results.csv is the authority for reported paper values; wisp_release.scoring implements its definitions for new evaluation. A checkpoint's sibling metric JSON is not automatically the final leaderboard truth, especially for worst-class F1.
Checkpoint loading, prediction equivalence, input export and metric equivalence are different checks. Passing one does not establish the others. The root release audit states which files, datasets and folds were actually evaluated and whether their predictions/metrics match the frozen evidence. An inventory entry alone is not a successful inference test.
The recorded model gate loaded and smoke-tested all 360 retained checkpoints, and rescored their 360 immutable saved test-prediction records against the frozen four-metric catalogue. Full real-input inference was rerun for only six ADL-Wrist fold-0 checkpoints: Random/Evolution with seeds 42, 43 and 44. Each reproduced all 861 held-out predictions exactly both on the original frozen NPZ and on a fresh pinned benchmark export. The other 354 full test splits were not rerun in this gate.
ADL's original time coordinates are seconds; the fresh Hub export stores window ordinals. The numerical values therefore differ. Input signals, participant/label rows and within-participant stable chronological permutations matched; the decoder uses that ordering rather than elapsed-time gaps. Both input representations were evaluated explicitly and produced the same locked predictions. This does not authorize ignoring time metadata differences for another dataset or another decoder.