MuVAP weights
Trained weights for the modules of MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild (Interspeech 2026).
Code: https://github.com/Haotian-Qi/MuVAP
A release is named {module}-{codebook}-{frontend}, with -mono- where the
model takes one channel plus VAD and a rate suffix where it does not run at
25 Hz. The name therefore states the configuration it was trained with. Each
release holds weights.pt, the config.yaml giving the architecture, and
provenance.json.
These are not Lightning checkpoints. They carry no optimizer state, and no copy of the frozen audio frontend, which is fetched from its own pretrained source when the model is built.
VAP releases
Eleven releases over three codebooks, two frontends and three frame rates. Every one is the epoch training ended on (4/4).
| Release | Audio in | Classes | Rate | Frontend | Params | Size | f1_macro | bacc | probe f1 |
|---|---|---|---|---|---|---|---|---|---|
vap-speaker-mimi-12hz |
2 ch | 256 | 12.5 | Mimi | 3.68M | 15 MB | 0.7776 | 0.7495 | 0.7714 |
vap-speaker-mimi |
2 ch | 256 | 25 | Mimi | 3.68M | 15 MB | 0.7755 | 0.7464 | 0.7716 |
vap-role-mimi |
1 ch | 136 | 25 | Mimi | 2.79M | 11 MB | 0.7589 | 0.7698 | 0.7623 |
vap-role-mimi-12hz |
1 ch | 136 | 12.5 | Mimi | 2.79M | 11 MB | 0.7565 | 0.7663 | 0.7570 |
vap-speaker-mono-mimi |
1 ch + VAD | 256 | 25 | Mimi | 2.83M | 11 MB | 0.7517 | 0.7216 | 0.7619 |
vap-speaker-mono-mimi-12hz |
1 ch + VAD | 256 | 12.5 | Mimi | 2.83M | 11 MB | 0.7484 | 0.7183 | 0.7585 |
vap-speaker-cpc-50hz |
2 ch | 256 | 50 | CPC | 3.75M | 15 MB | 0.7444 | 0.7129 | 0.7520 |
vap-role-cpc-50hz |
1 ch | 136 | 50 | CPC | 2.86M | 11 MB | 0.7358 | 0.7378 | 0.7382 |
vap-speaker-cpc |
2 ch | 256 | 25 | CPC | 3.94M | 16 MB | 0.7310 | 0.6983 | 0.7508 |
vap-role-cpc |
1 ch | 136 | 25 | CPC | 3.06M | 12 MB | 0.7289 | 0.7303 | 0.7304 |
vap-speaker-mono-cpc-50hz |
1 ch + VAD | 256 | 50 | CPC | 2.89M | 12 MB | 0.7148 | 0.6823 | 0.7398 |
The three codebooks differ in what a row of the projection window means. A speaker row is one audio channel, so the 2x4 future is 256 ordered states. A role row is the current speaker against the next, which makes the pair order-invariant and folds those 256 states onto 136 classes. A mono release uses the speaker codebook but hears one mixed channel, taking the causal voice activity behind each frame in place of the second channel - the original VAP arrangement, and the one release family that needs a VAD at inference.
f1_macro is zero-shot hold/shift on the Fisher turn-taking events, macro over
the two classes, decoded from the logits with no probe and no tuning. It is read
over p_future, the trailing 1.4 s of the projection window.
Three things to know before comparing rows.
The shift prior is a decision threshold. Scaling the shift slot by s and
renormalising is exactly thresholding at 1/(1+s). Each codebook is read at its
own default - 1.0 for speaker rows, which name channels, and 2.0 for role
rows, which name roles - so the speaker rows sit at 0.5 and the role rows at
0.333, and the bacc column is where that difference surfaces. Comparisons
within a codebook are sound; across codebooks they compare operating points as
much as models. probe f1, a logistic probe on the last-frame embedding, is
threshold-free and closer to like-for-like. Sweeping s on the test events
raises several rows, but that is measured on the scored split itself and is an
upper bound rather than a held-out result.
The frontend is the only large effect. Mimi over CPC is +0.045 on speaker and +0.030 on role at matched rate.
Frame rate barely matters. 25 -> 50 Hz on CPC gains 0.013 (speaker) and 0.007 (role). On Mimi, 25 -> 12.5 Hz costs nothing measurable: +0.002, -0.002, -0.003 across the three codebooks. Matching a 12.5 Hz downstream system is close to free. CPC halves a 100 Hz convolutional stack to the requested rate; Mimi taps either its 25 Hz encoder output or its 12.5 Hz latent, so 50 Hz is CPC-only.
ASD releases
| Release | What it is | Params | Size | Score |
|---|---|---|---|---|
asd-mimi |
audio-visual active speaker detection, frozen Mimi frontend | 18.62M | 75 MB | mAP_official 92.0429 |
asd-cpc |
as above, frozen CPC frontend | 18.88M | 76 MB | mAP_official 90.4983 |
mAP_official is the AVA-ActiveSpeaker mAP. Each -cpc / -mimi pair differs
solely in the frontend.
MuVAP releases
| Release | What it is | Runs with | Params | Size |
|---|---|---|---|---|
muvap-role-mimi |
multiparty fusion of the frozen VAP and ASD modules, role-relative codebook | vap-role-mimi + asd-mimi |
2.08M | 8 MB |
muvap-role-cpc |
as above, on the CPC pair | vap-role-cpc + asd-cpc |
2.08M | 8 MB |
A MuVAP release holds the fusion alone. It reads what its two frozen modules
produce, so it only works with the pair it was trained on; provenance.json
names that pair under pairs_with. Both are the epoch training ended on (4/4).
huggingface-cli download Haotian-Qi/MuVAP \
--include 'vap-role-mimi/*' 'asd-mimi/*' 'muvap-role-mimi/*' --local-dir weights
python -m demo.server --vap-weights weights/vap-role-mimi \
--asd-weights weights/asd-mimi --muvap-weights weights/muvap-role-mimi \
--certfile cert.pem --keyfile key.pem # live demo, see DEMO.md
Use
git clone https://github.com/Haotian-Qi/MuVAP && cd MuVAP
python -m pip install -e . # add '.[mimi]' for the Mimi releases
huggingface-cli download Haotian-Qi/MuVAP --include 'vap-role-mimi/*' --local-dir weights
python train_vap.py --config config/yaml/vap_role.yaml \
--test --weights weights/vap-role-mimi \
--set fisher_path=/path/to/fisher
That reproduces the score in the table exactly. The architecture comes from the
release's config.yaml; dataset roots and output directories come from the
--config you pass, so a release runs anywhere without reproducing the paths it
was trained with.
Or load one directly:
import yaml
from models.release import load_weights, resolve
from models.vap import build_vap
weights, config = resolve('weights/vap-role-mimi')
cfg = yaml.safe_load(open(config))['vap']
model = load_weights(build_vap(cfg), weights).eval()
logits = model(waveform) # [batch, frames, classes]
logits = model(waveform, vad) # the -mono- releases, VAD at the model rate
An ASD release loads the same way through models.asd.AudioVisualASD and the
asd key of its config.