FORGE β FOG Representation via Generative Encoding
Self-supervised spectral-temporal encoders for Freezing of Gait (FOG) detection from a single lower-back accelerometer. Pretrained by masked autoencoding on 11,724 h (~21M windows) of unlabeled at-home recordings from 65 participants, then trained for FOG detection on the 57-participant DeFOG cohort only. Evaluated with no target-cohort training on four external cohorts: FogAtHome-provoking, tDCS-FOG, Stanford and FogAtHome daily living.
Headline: the released nine-head frozen-encoder ensemble is evaluated against expert video annotation on an independent cohort β ICC(%TF) = 0.899 [0.700, 0.970], using one lower-back IMU and no target-cohort training.
External results (released detector)
Nine-head MC frozen-probe ensemble (3 participant folds x 3 seeds), evaluated with no target-cohort training and the unchanged DeFOG operating point of 0.35.
| Cohort (N) β shift | AUROC | AP | ICC(%TF) |
|---|---|---|---|
| FogAtHome-provoking (12) β cross-study | 0.887 [0.830, 0.922] | 0.804 [0.573, 0.902] | 0.899 [0.700, 0.970] |
| tDCS-FOG (71) β cross-protocol | 0.917 [0.863, 0.950] | 0.812 [0.550, 0.923] | 0.876 [0.810, 0.920] |
| Stanford (7) β site / device / med state | 0.734 [0.624, 0.846] | 0.400 | -0.119 [-0.920, 0.670] |
| FogAtHome daily living (11) β naturalistic* | 0.803 [0.737, 0.877] | 0.105 | 0.656 [-0.129, 0.872] |
* Daily living is scored inside a label-independent walking-and-standing domain
(58.18 h of 301.8 h, 2.92% FOG); it is gait-conditioned burden, not whole-recording
%TF. Stanford is negative evidence: discrimination survives the shift, the fixed
threshold does not (its oracle-rule threshold is 0.18). In-distribution reference:
window-level AP 0.730 on the DeFOG validation folds. Full definitions and confidence
intervals are in manifest.yaml under results:.
The tDCS-FOG AP of 0.812 [0.550, 0.923] above is retained as a historical
fold-safe analysis result, not the released nine-head reproduction target.
The existing public code manifest records AP 0.866 [0.661, 0.946] for the
all-71-participant nine-head assembly used by the reproduction workflow.
No result has been recomputed for this documentation update. Use the public
code's release/manifest.yaml for reproduction targets; this model repository's
older manifest.yaml also contains historical analysis entries.
Controlled comparison (a different model set)
The paper's headline effect is a separate, deliberately constrained experiment: two arms differing only in encoder initialization, under matched downstream training. Its numbers are lower than the released detector's on the same cohort because it is a different model set β not a worse estimate of the same thing.
| Cohort | Metric | Self-supervised | Supervised from scratch | Difference [95% CI] |
|---|---|---|---|---|
| FogAtHome-provoking | AUROC | 0.861 | 0.752 | +0.109 [0.029, 0.182] |
| FogAtHome-provoking | AP | 0.784 | 0.592 | +0.192 [0.064, 0.338] |
| tDCS-FOG | AUROC | 0.869 | 0.657 | +0.212 [0.126, 0.280] |
| tDCS-FOG | AP | 0.811 | 0.501 | +0.310 [0.133, 0.397] |
Each result in manifest.yaml names its model set and the manuscript table it comes
from, so the two sets stay distinguishable.
Released weights
Pretrained FORGE encoders (the backbones)
| Context | Window (frames) | File | Params |
|---|---|---|---|
| LC | 1000 | encoders/lc.safetensors |
14,147,072 |
| MC | 500 | encoders/mc.safetensors |
14,147,072 |
| SC | 200 | encoders/sc.safetensors |
12,918,272 |
Downstream classification weights (57-participant DeFOG, 3-fold participant-level CV)
| File | Context | Phase | Fold |
|---|---|---|---|
classification/lc_probe_fold0.safetensors |
lc | probe | 0 |
classification/lc_probe_fold1.safetensors |
lc | probe | 1 |
classification/lc_probe_fold2.safetensors |
lc | probe | 2 |
classification/mc_probe_fold0.safetensors |
mc | probe | 0 |
classification/mc_probe_fold1.safetensors |
mc | probe | 1 |
classification/mc_probe_fold2.safetensors |
mc | probe | 2 |
classification/sc_probe_fold0.safetensors |
sc | probe | 0 |
classification/sc_probe_fold1.safetensors |
sc | probe | 1 |
classification/sc_probe_fold2.safetensors |
sc | probe | 2 |
classification/lc_finetune_fold0.safetensors |
lc | finetune | 0 |
classification/lc_finetune_fold1.safetensors |
lc | finetune | 1 |
classification/lc_finetune_fold2.safetensors |
lc | finetune | 2 |
classification/mc_finetune_fold0.safetensors |
mc | finetune | 0 |
classification/mc_finetune_fold1.safetensors |
mc | finetune | 1 |
classification/mc_finetune_fold2.safetensors |
mc | finetune | 2 |
classification/sc_finetune_fold0.safetensors |
sc | finetune | 0 |
classification/sc_finetune_fold1.safetensors |
sc | finetune | 1 |
classification/sc_finetune_fold2.safetensors |
sc | finetune | 2 |
classification/lc_supervised_fold0.safetensors |
lc | supervised | 0 |
classification/lc_supervised_fold1.safetensors |
lc | supervised | 1 |
classification/lc_supervised_fold2.safetensors |
lc | supervised | 2 |
classification/mc_supervised_fold0.safetensors |
mc | supervised | 0 |
classification/mc_supervised_fold1.safetensors |
mc | supervised | 1 |
classification/mc_supervised_fold2.safetensors |
mc | supervised | 2 |
classification/sc_supervised_fold0.safetensors |
sc | supervised | 0 |
classification/sc_supervised_fold1.safetensors |
sc | supervised | 1 |
classification/sc_supervised_fold2.safetensors |
sc | supervised | 2 |
What this release contains
The released nine-head frozen-encoder ensemble contains all nine BiGRU heads:
three participant-level DeFOG folds (0, 1, 2) Γ seeds 42, 43, and 44, over one
shared pretrained frozen MC encoder (encoders/mc.safetensors). External
evaluation averages all nine heads; the DeFOG reference remains the seed-42
out-of-fold result. The 27 seed-42 classification files listed above are joined by:
| Seed | Fold 0 | Fold 1 | Fold 2 |
|---|---|---|---|
| 43 | classification/mc_probe_s43_fold0.safetensors |
classification/mc_probe_s43_fold1.safetensors |
classification/mc_probe_s43_fold2.safetensors |
| 44 | classification/mc_probe_s44_fold0.safetensors |
classification/mc_probe_s44_fold1.safetensors |
classification/mc_probe_s44_fold2.safetensors |
Usage
This release is weights only: each file is a .safetensors tensor set with small
string metadata (name, context, phase, fold, seed, and the experiment config that
rebuilds the model). No training configuration, optimizer state or local file path is
included, and loading executes no pickled code.
Rebuild a model from the companion repo, which composes the architecture from the
experiment config named in the file's metadata and in manifest.yaml:
from utils.released_weights import load_released_model
model, config = load_released_model(
"release/forge-fog/classification/mc_probe_fold0.safetensors",
)
Or read the tensors directly:
from safetensors.torch import load_file
from safetensors import safe_open
state_dict = load_file("classification/mc_probe_fold0.safetensors")
with safe_open("classification/mc_probe_fold0.safetensors", framework="pt") as f:
meta = f.metadata() # name / kind / context / phase / fold / seed / experiment / splits
Each classification file already contains its encoder, so encoders/*.safetensors are
needed only to train new heads.
Citation
Lior Nisimov, Amit Salomon, Eran Gazit, Talia Herman, Lior Rokach, Jeffrey M. Hausdorff, Nathaniel Shimoni. Self-supervised learning improves cross-cohort freezing-of-gait detection from a single lower-back accelerometer. Submitted to npj Digital Medicine, 2026.
Reproduce the released detector on the four external cohorts using the public
REPRODUCE.md
and ./reproduce.sh (or reproduce.bat on Windows). The public code's
release/manifest.yaml pins the weights and datasets. The workflow downloads
public safetensors and data, builds the evaluation inputs, and generates its own
predictions; private checkpoints, precomputed predictions, and private paths are
not required. This workflow does not reproduce every manuscript analysis.
Code: github.com/Lior-Nis/forge-public.
Data: Liornis/fog-dataset.
Intended use and limitations
For research use only. FORGE is not clinically validated and is not a medical device. It must not be used to diagnose, monitor, or make treatment decisions for an individual. Performance and burden calibration vary across cohorts, devices, and recording conditions.
License: MIT.