Du-IN — braindecode weights

braindecode-native re-host of the official Du-IN pretrained checkpoints (Zheng et al., NeurIPS 2024), converted so they load directly with braindecode.models.DuIN.

Du-IN decodes speech (61 Chinese words) from sEEG. Each subject's signal (10 selected bipolar channels, 1000 Hz) goes through a subject-specific spatial projection, a patch-wise convolutional tokenizer (one token per 100 ms) and an 8-layer Transformer encoder. The encoder is pretrained per subject by masked modelling of discrete codes learned by a VQ-VAE (mae); the VQ-VAE encoder itself (vqvae) is the other released stage.

Provenance

Original code https://github.com/liulab-repository/Du-IN (MIT)
Original weights https://huggingface.co/datasets/liulab-repository/Du-IN (pretrains/duin/<subject>/<stage>/model/checkpoint-399.pth, CC BY 4.0)
Paper Zheng et al. (2024), Du-IN: Discrete units-guided mask modeling for decoding speech from intracranial neural signals, NeurIPS 2024, arXiv:2405.11459

These weights are a format conversion only of the authors' checkpoints, with no re-training. For each of the 24 checkpoints:

  • the 183 encoder tensors are carried over unchanged (182 bit-for-bit, plus the subject layer W (160, 1), reshaped to spatial_projection.weight (16, 10));
  • the ported encoder reproduces the upstream duin_cls encoder output on the same input to < 1e-5 (max abs diff 7.2e-6 over all 24);
  • DuIN.from_pretrained on the uploaded file gives the same state dict and outputs as loading the original .pth.

The pretraining-only parts (mask_emb, VQ codebook, MAE token head) are dropped. The classification head (head_hidden, final_layer, 61 outputs) is randomly initialised and must be fine-tuned. Per-file source sha256 are in conversion_report.json.

Files

path content
config.json shared DuIN config (n_chans=10, n_times=3000, sfreq=1000, n_outputs=61)
subj-XXX/mae/model.safetensors MAE-pretrained encoder of subject XXX (001 to 012), the paper's Du-IN
subj-XXX/vqvae/model.safetensors VQ-VAE-stage encoder of subject XXX (the paper's "Du-IN (vqvae)")
model.safetensors default = subj-001/mae

Every checkpoint is subject-specific: it was trained on that subject's 10 channels and only fits that montage, in the channel order of the authors' dataset. They are not a cross-subject foundation model.

Usage

from braindecode.models import DuIN

# MAE-pretrained Du-IN of subject 002; the 61-word head is re-initialised.
model = DuIN.from_pretrained(
    "braindecode/duin-pretrained",
    filename="subj-002/mae/model.safetensors",
)
# Pass n_outputs=... for another task (the head is re-created).

# input: (batch, 10, 3000) — 3 s at 1000 Hz, z-scored per channel
logits = model(x)

The reference head ends with a sigmoid and is trained with F.cross_entropy(torch.sigmoid(logits), target); braindecode returns logits. Use that loss to match the reference training.

Replication (braindecode PR #1298)

Fine-tuning on the 61-word task with the paper's protocol (200 epochs, AdamW, per-word 80/10/10 split, 6 seeds × 12 subjects, mean ± SE over subjects):

paper (Table 2) braindecode, sigmoid loss braindecode, logits
Du-IN (mae) 62.70 ± 4.69 54.51 ± 4.77 50.07 ± 5.26
Du-IN (vqvae) 58.24 ± 4.83 49.33 ± 4.42 49.31 ± 5.04
scratch 56.29 ± 5.20 48.40 ± 4.98 47.32 ± 5.26

The pretraining gain over scratch is reproduced (+6.1 vs +6.4 in the paper, 12/12 subjects), but absolute accuracies are about 8 points lower. The authors' own run_cls.py on the same released data gives the same numbers as the port (subject 002: 47.8 % mae vs 83.6 % in the paper), so the gap comes from the released data and checkpoints, not from the conversion.

License

Weights: CC BY 4.0, as released by the authors; cite the original work. braindecode code: BSD-3-Clause; original Du-IN code: MIT.

Citation

@inproceedings{zheng2024duin,
  title={Du-IN: Discrete units-guided mask modeling for decoding speech from intracranial neural signals},
  author={Zheng, Hui and Wang, Hai-Teng and Jiang, Wei-Bang and Chen, Zhong-Tao and He, Li and Lin, Pei-Yang and Wei, Peng-Hu and Zhao, Guo-Guang and Liu, Yun-Zhe},
  booktitle={Advances in Neural Information Processing Systems},
  volume={37},
  year={2024}
}
Downloads last month
-
Safetensors
Model size
4.18M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for braindecode/duin-pretrained