MuVAP weights

Trained weights for the modules of MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild (Interspeech 2026).

Code: https://github.com/Haotian-Qi/MuVAP

A release is named {module}-{codebook}-{frontend}, with -mono- where the model takes one channel plus VAD and a rate suffix where it does not run at 25 Hz. The name therefore states the configuration it was trained with. Each release holds weights.pt, the config.yaml giving the architecture, and provenance.json.

These are not Lightning checkpoints. They carry no optimizer state, and no copy of the frozen audio frontend, which is fetched from its own pretrained source when the model is built.

VAP releases

Eleven releases over three codebooks, two frontends and three frame rates. Every one is the epoch training ended on (4/4).

Release Audio in Classes Rate Frontend Params Size f1_macro bacc probe f1
vap-speaker-mimi-12hz 2 ch 256 12.5 Mimi 3.68M 15 MB 0.7776 0.7495 0.7714
vap-speaker-mimi 2 ch 256 25 Mimi 3.68M 15 MB 0.7755 0.7464 0.7716
vap-role-mimi 1 ch 136 25 Mimi 2.79M 11 MB 0.7589 0.7698 0.7623
vap-role-mimi-12hz 1 ch 136 12.5 Mimi 2.79M 11 MB 0.7565 0.7663 0.7570
vap-speaker-mono-mimi 1 ch + VAD 256 25 Mimi 2.83M 11 MB 0.7517 0.7216 0.7619
vap-speaker-mono-mimi-12hz 1 ch + VAD 256 12.5 Mimi 2.83M 11 MB 0.7484 0.7183 0.7585
vap-speaker-cpc-50hz 2 ch 256 50 CPC 3.75M 15 MB 0.7444 0.7129 0.7520
vap-role-cpc-50hz 1 ch 136 50 CPC 2.86M 11 MB 0.7358 0.7378 0.7382
vap-speaker-cpc 2 ch 256 25 CPC 3.94M 16 MB 0.7310 0.6983 0.7508
vap-role-cpc 1 ch 136 25 CPC 3.06M 12 MB 0.7289 0.7303 0.7304
vap-speaker-mono-cpc-50hz 1 ch + VAD 256 50 CPC 2.89M 12 MB 0.7148 0.6823 0.7398

The three codebooks differ in what a row of the projection window means. A speaker row is one audio channel, so the 2x4 future is 256 ordered states. A role row is the current speaker against the next, which makes the pair order-invariant and folds those 256 states onto 136 classes. A mono release uses the speaker codebook but hears one mixed channel, taking the causal voice activity behind each frame in place of the second channel - the original VAP arrangement, and the one release family that needs a VAD at inference.

f1_macro is zero-shot hold/shift on the Fisher turn-taking events, macro over the two classes, decoded from the logits with no probe and no tuning. It is read over p_future, the trailing 1.4 s of the projection window.

Three things to know before comparing rows.

The shift prior is a decision threshold. Scaling the shift slot by s and renormalising is exactly thresholding at 1/(1+s). Each codebook is read at its own default - 1.0 for speaker rows, which name channels, and 2.0 for role rows, which name roles - so the speaker rows sit at 0.5 and the role rows at 0.333, and the bacc column is where that difference surfaces. Comparisons within a codebook are sound; across codebooks they compare operating points as much as models. probe f1, a logistic probe on the last-frame embedding, is threshold-free and closer to like-for-like. Sweeping s on the test events raises several rows, but that is measured on the scored split itself and is an upper bound rather than a held-out result.

The frontend is the only large effect. Mimi over CPC is +0.045 on speaker and +0.030 on role at matched rate.

Frame rate barely matters. 25 -> 50 Hz on CPC gains 0.013 (speaker) and 0.007 (role). On Mimi, 25 -> 12.5 Hz costs nothing measurable: +0.002, -0.002, -0.003 across the three codebooks. Matching a 12.5 Hz downstream system is close to free. CPC halves a 100 Hz convolutional stack to the requested rate; Mimi taps either its 25 Hz encoder output or its 12.5 Hz latent, so 50 Hz is CPC-only.

ASD releases

Release What it is Params Size Score
asd-mimi audio-visual active speaker detection, frozen Mimi frontend 18.62M 75 MB mAP_official 92.0429
asd-cpc as above, frozen CPC frontend 18.88M 76 MB mAP_official 90.4983

mAP_official is the AVA-ActiveSpeaker mAP. Each -cpc / -mimi pair differs solely in the frontend.

MuVAP releases

Release What it is Runs with Params Size
muvap-role-mimi multiparty fusion of the frozen VAP and ASD modules, role-relative codebook vap-role-mimi + asd-mimi 2.08M 8 MB
muvap-role-cpc as above, on the CPC pair vap-role-cpc + asd-cpc 2.08M 8 MB

A MuVAP release holds the fusion alone. It reads what its two frozen modules produce, so it only works with the pair it was trained on; provenance.json names that pair under pairs_with. Both are the epoch training ended on (4/4).

huggingface-cli download Haotian-Qi/MuVAP \
    --include 'vap-role-mimi/*' 'asd-mimi/*' 'muvap-role-mimi/*' --local-dir weights

python -m demo.server --vap-weights weights/vap-role-mimi \
    --asd-weights weights/asd-mimi --muvap-weights weights/muvap-role-mimi \
    --certfile cert.pem --keyfile key.pem          # live demo, see DEMO.md

Use

git clone https://github.com/Haotian-Qi/MuVAP && cd MuVAP
python -m pip install -e .            # add '.[mimi]' for the Mimi releases
huggingface-cli download Haotian-Qi/MuVAP --include 'vap-role-mimi/*' --local-dir weights

python train_vap.py --config config/yaml/vap_role.yaml \
    --test --weights weights/vap-role-mimi \
    --set fisher_path=/path/to/fisher

That reproduces the score in the table exactly. The architecture comes from the release's config.yaml; dataset roots and output directories come from the --config you pass, so a release runs anywhere without reproducing the paths it was trained with.

Or load one directly:

import yaml
from models.release import load_weights, resolve
from models.vap import build_vap

weights, config = resolve('weights/vap-role-mimi')
cfg = yaml.safe_load(open(config))['vap']
model = load_weights(build_vap(cfg), weights).eval()

logits = model(waveform)              # [batch, frames, classes]
logits = model(waveform, vad)         # the -mono- releases, VAD at the model rate

An ASD release loads the same way through models.asd.AudioVisualASD and the asd key of its config.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support