|
Download DOCS.md from omayib/badak: direct link, hf CLI and curl.
- Browser
- Download file 11.1 kB
-
https://huggingface.co/omayib/badak/resolve/main/DOCS.md
- Command line
-
hf download hf://omayib/badak/DOCS.md
-
curl -L -o DOCS.md https://huggingface.co/omayib/badak/resolve/main/DOCS.md
11.1 kB
| # badminsens-analyzer | |
| Counts and classifies badminton strokes in Garmin (Forerunner 255) `.fit` files with models | |
| trained on the BADS smartwatch dataset (`BADS_CLL_OPENACCESS_V1`, BadminSense paper, `paper.pdf`). | |
| The recordings come from a custom Garmin watch app that logs raw IMU data: | |
| [github.com/omayib/badak](https://github.com/omayib/badak). Trained models and a Colab notebook are | |
| on Hugging Face: [omayib/badak](https://huggingface.co/omayib/badak). | |
| ## Quick start | |
| ```bash | |
| python3 -m venv .venv && .venv/bin/pip install -r requirements.txt | |
| .venv/bin/python -m badminsens.predict path/to/ACTIVITY.fit # default model | |
| .venv/bin/python -m badminsens.predict path/to/ACTIVITY.fit --model gru # any shipped model | |
| .venv/bin/python -m badminsens.predict path/to/ACTIVITY.fit --model all # compare all models | |
| .venv/bin/python -m badminsens.plot path/to/ACTIVITY.fit # 4 charts + desktop/mobile report pages | |
| .venv/bin/python -m badminsens.plot path/to/ACTIVITY.fit --charts speed-simple # just one | |
| ``` | |
| In Python or a notebook: `rows, summary = badminsens.predict.analyze("ACTIVITY.fit", model="svm")`, | |
| then `badminsens.plot.plot_session("ACTIVITY.fit", rows)`. The charts: `hr` heart rate + strokes, `speed` average racket | |
| speed (km/j) per 5 minutes, `speed-simple` average racket speed per stroke and game (km/j = km per hour, | |
| estimated), times in the watch's local zone (WIB), `timeline` strokes per | |
| 2 minutes; `report-desktop` / `report-mobile` put all four on one page (a 2 x 2 grid, or one | |
| narrow column for phones). A gap in the IMU data of at least `--min-rest` minutes (default 2, i.e. the watch app paused) | |
| is drawn as rest and starts a new game. `badak_colab.ipynb` runs the same in | |
| Google Colab, straight from the Hugging Face repo `omayib/badak`. | |
| Output: counts and percentages per stroke in the terminal, `<ACTIVITY>_summary.json`, and | |
| `<ACTIVITY>_strokes.csv` (one row per swing: time, stroke, confidence, gate decision, class | |
| probabilities; with `--model all` every model's label). | |
| Options: `--min-peak-dps` (swing detection, default 800 °/s), `--min-conf` (below this a stroke is | |
| also flagged `Uncertain`, default 0.5), `--gate-quantile` (overhead-gate strictness, default 0.97), | |
| `--no-gate` (force every swing into a class), `--balance` (class-balanced session statistics | |
| iterations; default from the registry). | |
| ## What it can and cannot count | |
| BADS contains only **4 overhead strokes**, all right-handed with the watch on the racket wrist: | |
| | BADS label | stroke | side | | |
| |---|---|---| | |
| | `ForehandKill` | Smash | forehand | | |
| | `ForehandHigh` | Clear | forehand | | |
| | `ForehandLob` | Drop (吊球, "lob" is a mistranslation; it is **not** an underhand lift) | forehand | | |
| | `BackhandTransition` | Backhand clear | backhand | | |
| A match also contains lifts, net shots, drives, serves, defence and non-hitting swings. Earlier | |
| versions forced all of them into one of the 4 classes, which inflated "backhand" (308 of 484 | |
| swings). Now an **overhead gate** first checks whether a swing looks like a dataset stroke at all | |
| (forearm orientation before and after the swing, rotation axis at peak speed). The gate is | |
| Mahalanobis distance to the stroke classes, accepting 97% of training strokes and 91% of strokes | |
| of unseen players. Rejected swings are reported as `Other (not a dataset stroke)`. | |
| Result for `24469883647_ACTIVITY.fit` (match, right wrist, right-handed), default model: | |
| | | count | % of dataset strokes | % of all swings | | |
| |---|---|---|---| | |
| | Smash | 70 | 36.6 | 14.5 | | |
| | Clear | 7 | 3.7 | 1.4 | | |
| | Drop | 48 | 25.1 | 9.9 | | |
| | Backhand clear | 66 | 34.6 | 13.6 | | |
| | Other (not a dataset stroke) | 293 | | 60.5 | | |
| Forehand 65% / backhand 35% of the dataset strokes. The other models give 28–41% smash, 4–18% | |
| clear, 12–27% drop and 28–42% backhand clear (`--model all`). Treat that spread as the | |
| uncertainty: there are no labelled Garmin strokes, so accuracy on your watch cannot be measured. | |
| Clear is probably undercounted, because clear→smash is the main error on the dataset too (clear | |
| recall 0.3–0.66). | |
| ## Pipeline | |
| ```bash | |
| .venv/bin/python -m badminsens.dataset # 1. BADS JSON -> data/bads_windows.npz, player split | |
| .venv/bin/python -m badminsens.tune # 2. configuration search loop (CV on 9 dev players) | |
| .venv/bin/python -m badminsens.benchmark # 3. evaluate once on 3 test players, ship to models/ | |
| .venv/bin/python -m badminsens.select # 4. choose the default (leave-one-player-out, 12 players) | |
| ``` | |
| 1. **Pre-processing** (`dataset.py`, `io.py`, `features.py`): 848 strokes as 2 s windows at 100 Hz | |
| (acc m/s², gyro rad/s), aligned so the gyro peak is at sample 94 (like windows cut from a .fit). | |
| 2. **Split by player**: train 7 players (490 strokes) + val 2 (141) = 9 *dev* players used for | |
| tuning; **test 3 players (217 strokes: Play6, Play9, Play12)** used once, after tuning. | |
| 3. **Features**: the paper's 23 time/frequency descriptors per axis (trim 200 ms, 20 Hz low-pass, | |
| stats of signal and derivative, FFT energy/centroid/entropy, Welch PSD) + our 378 shape and | |
| orientation features; sequence nets get the 12/20 Hz low-passed signal at 25 Hz (8 channels). | |
| 4. **Session normalization**: each player/session is centred on its own mean (removes individual | |
| style and device offsets). At inference the Garmin session mean is computed over the gate-accepted | |
| swings, re-weighted so each predicted class counts equally (`balance_iters`), because a real | |
| session does not have the dataset's balanced stroke mix. (Z-scaling per session scored slightly | |
| higher in CV but broke on Garmin data, so it is excluded.) | |
| 5. **Augmentation**: session-level wrist rotation (watch worn differently), per-stroke jitter, | |
| time shift, amplitude scaling, noise. | |
| ## The tuning loop | |
| `badminsens.tune` searched 12 model families: 174 configurations in the final run, logged in | |
| `experiments/tuning_log.jsonl` (plus 80 in an earlier run before peak alignment, | |
| `tuning_log_v1_unaligned.jsonl`) (config + accuracy, balanced accuracy, macro P/R/F1, Cohen κ, MCC, | |
| log-loss, ROC-AUC, per-player accuracy, confusion matrix, skewed-session accuracy). Per family it | |
| tries the default, then random configurations, then mutations of the best one, and stops when | |
| leave-one-player-out CV accuracy reaches the 0.90 target or the budget runs out. Searched settings: | |
| - **SVM / LR / RF / ET / HGB**: feature set, normalization, C, γ, trees, max_features, leaf size, | |
| learning rate, leaf nodes, L2, number of augmented copies. | |
| - **Neural nets**: epochs (30/60/100), batch (16/32/64), optimizer (AdamW/Adam/SGD+Nesterov), | |
| learning rate, weight decay, scheduler (cosine/one-cycle/none), label smoothing, dropout, dense | |
| head (none, 64, 128, 128-64, 256-128, …), conv channels and kernel, RNN hidden size, layers and | |
| direction, transformer width, heads and layers, input low-pass and rate, augmentation strength. | |
| **No configuration reached 0.90 on unseen players.** The best was SVM at 0.892 CV. The biggest | |
| gains came from the data rather than from hyperparameters: session normalization (+7 points) and | |
| peak alignment (+1.5). | |
| ## Results | |
| Full tables (all metrics, per-class P/R/F1, configurations, confusion matrices, tuning summary): | |
| [`models/BENCHMARK.md`](models/BENCHMARK.md). | |
| | model | CV acc (9 dev) | **test acc (3 unseen)** | macro F1 | κ | skewed session | 12-player LOPO | | |
| |---|---|---|---|---|---|---| | |
| | **svm + cnn + transformer** (default) | | | | | 0.845 / 0.856* | **0.875** | | |
| | gru | 0.794 | **0.873 ± 0.021** | 0.870 | 0.829 | 0.861 | 0.821 | | |
| | rf | 0.870 | 0.842 | 0.825 | 0.788 | 0.722 | 0.861 | | |
| | transformer | 0.783 | 0.839 | 0.830 | 0.784 | 0.827 | 0.829 | | |
| | et | 0.872 | 0.836 | 0.819 | 0.780 | 0.721 | 0.853 | | |
| | svm | **0.892** | 0.814 | 0.783 | 0.751 | 0.716 | 0.871 | | |
| | hgb | 0.878 | 0.810 | 0.788 | 0.745 | 0.728 | 0.861 | | |
| | cnn | 0.792 | 0.810 | 0.811 | 0.745 | 0.805 | 0.807 | | |
| | mlp | 0.881 | 0.794 | 0.755 | 0.724 | 0.743 | 0.856 | | |
| | logreg | 0.853 | 0.779 | 0.738 | 0.703 | 0.684 | | | |
| | lstm | 0.778 | 0.768 | 0.769 | 0.689 | 0.767 | | | |
| | cnn_lstm | 0.759 | 0.754 | 0.743 | 0.672 | 0.757 | | | |
| | tcn | 0.799 | 0.736 | 0.741 | 0.647 | 0.735 | | | |
| *skewed session: test accuracy when the session statistics come from an unbalanced mix of the | |
| player's strokes (without / with class balancing); for the default it is the 12-player value.* | |
| How to read it: | |
| - The **CV and test rankings disagree**: SVM is best on the 9 dev players and GRU on the 3 test | |
| players. Accuracy varies more between players than between models. Every model scores about 0.73 | |
| on Play12 and 0.85–0.99 on Play9. With 3 test players, one standard error is about ±2.5 points. | |
| - The **default was therefore chosen after the test evaluation**, by leave-one-player-out over all | |
| 12 players (`badminsens.select`, 4× more unseen players). That makes the ensemble's number more | |
| reliable, but it is not an untouched test score. | |
| - The paper reports 91.43% (SVM, one random leave-3-users-out split). Reproducing its pipeline | |
| gives 78.9% ± 5.4 over 20 random 3-player splits (best split 88.1%). Split luck explains most of | |
| the gap. | |
| - Backhand clear is almost perfect, drop is good (F1 ≈ 0.9), and **clear vs smash** is the | |
| remaining error, as the paper also notes. | |
| Best configurations (all settings in `models/registry.json`): | |
| | model | key settings | | |
| |---|---| | |
| | svm | both feature sets, session-centred, RBF C=0.3, γ=3e-4, 4 augmented copies | | |
| | cnn | conv 64-128-128, kernel 9, dense 128, dropout 0.3, 30 epochs, batch 32, AdamW lr 5e-4, wd 1e-4, cosine, label smoothing 0.2, 20 Hz → 25 Hz input, session-centred | | |
| | transformer | d_model 32, 4 heads, 2 layers, no dense head, dropout 0.3, 30 epochs, batch 32, Adam lr 5e-4, wd 1e-3, cosine, session-centred | | |
| | gru | 2-layer unidirectional GRU, 64 hidden, dense 128, dropout 0.5, 60 epochs, batch 32, SGD+Nesterov lr 0.01, cosine, no session normalization | | |
| ## Requirements for your recordings | |
| - Raw accelerometer + gyroscope logging must be enabled (like `24469883647_ACTIVITY.fit`). The Garmin | |
| watch app at [github.com/omayib/badak](https://github.com/omayib/badak) records this kind of file. | |
| - Watch on the **racket-hand wrist**, right-handed (as in BADS). The Garmin→BADS axis mapping was | |
| verified from the data: the gyro/accel physics check passed, and the identity rotation fits the | |
| dataset's stroke orientations far better than the other 23 rotations (32% vs ≤10% of swings in | |
| distribution). | |
| ## Garmin ↔ BADS sensor mapping | |
| | | Garmin raw | Conversion | | |
| |---|---|---| | |
| | accel | 1 mg/LSB, ±8 g, **sign opposite to Android** | × −9.80665/1000 → m/s² | | |
| | gyro | 16.384 LSB/(°/s), ±2000 °/s | ÷ 16.384 → °/s → rad/s | | |
| ## Notes | |
| - Training uses the GPU when CUDA works and falls back to CPU otherwise. On this machine the | |
| driver was stuck (`CUDA unknown error`; fix with `sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm`), | |
| so everything was trained on CPU. | |
| - `experiments/`: tuning logs (v1 = before peak alignment, kept for reference), selection results | |
| and run logs. | |