RysUpAlign learned alignment model v6
Frame embeddings for audio-to-audio vocal alignment: aligning a double, gang or backing vocal to a guide vocal without lyrics. Embeddings of the two takes are compared with dynamic time warping (DTW); the model is trained so that the same sung moment in two takes gets the same embedding even across singers, timbre and pitch (unison, harmony, octave).
It is the timing model of the RysUpAlign plug-in by Rys Up Audio. Benchmark, code and evaluation: https://github.com/rysupaudio-lab/rysupalign-bench
Files
| file | size | sha256 |
|---|---|---|
rysupalign_learned_v6.onnx |
305 MB, fp32 | db3bf5b88834bba926c147d8785e829845c7827248efc74e2966f0bec7ec8b32 |
rysupalign_learned_v6.int8.onnx |
112 MB, dynamic int8 | e48d692579af3d8ed710ff99e47a1c37d0c1e0205440d5a69d1521a69f505246 |
Both give the same accuracy on the benchmark (see below).
Input / output
- input
wav: float32[1, S], mono 16 kHz audio, normalised to zero mean and unit variance over the file - output
emb: float32[1, T, 128], unit-norm embeddings at 100 fps; frame k is centred at 10k + 12.5 ms
For long files, run 30 s chunks with 2 s overlap and drop 1 s on each side of every join.
import numpy as np, librosa, onnxruntime as ort
sess = ort.InferenceSession("rysupalign_learned_v6.onnx", providers=["CPUExecutionProvider"])
y, _ = librosa.load("take.wav", sr=16000, mono=True)
y = ((y - y.mean()) / (y.std() + 1e-7)).astype(np.float32)
emb = sess.run(None, {"wav": y[None]})[0][0] # [T, 128], 100 fps
# cost between two takes: 1 - emb_a @ emb_b.T, then DTW
The complete aligners used for the results (gating, log-mel block, DTW method M1) are
runners/learned_onnx.py and runners/m1_dtw.py in the GitHub repository.
Architecture
- Encoder: HuBERT-base (
facebook/hubert-base-ls960, revisionaf46f65f540dc3ca7aa59f46c6c3d5dbb4374fa8, Apache-2.0), frozen, first 9 transformer layers. - Head (1.66 M parameters): LayerNorm and a learned softmax-weighted sum of layers 3โ9, Linear 768โ256, ร2 transposed-conv upsampling to 100 fps, plus a log-mel branch (64 bands, 25 ms; 80 bands over 80โ2000 Hz, 64 ms) on the same grid, 3 residual dilated Conv1d blocks, Linear โ 128, L2 normalisation.
- The log-mel front end is part of the graph (STFT as fixed conv1d kernels), so the ONNX file takes raw audio.
Training
Symmetric frame-level InfoNCE (positive = the true fractional frame in the other take, negatives = all other frames within ยฑ1 s). Data: synthetic doubles with exact timing (WSOLA timing slop, pitch and formant shifts, EQ, drive, reverb, noise) made from Rys Up Audio's own studio vocal recordings and VocalSet 1.2 (CC BY 4.0), plus pseudo-labelled real double takes and VocalSet cross-singer pairs. No evaluation audio was used for training.
Results (RAB, mean / median / share within 20 ms of the timing error)
| system | R: real unison choir pairs (48) | S: exact-truth cases (120), mean | D: held-out Dagstuhl pairs (12), mean |
|---|---|---|---|
| no alignment | 52.0 / 34.8 ms / 33.7% | 39.5 ms | 65.1 ms |
| HuBERT-base L6 + log-mel DTW | 47.4 / 24.1 ms / 45.0% | 6.1 ms | 42.6 ms |
| MERT-v1-95M L4 + log-mel DTW | 44.8 / 20.0 ms / 49.9% | โ | 47.8 ms |
| MERT-v1-330M L10 + log-mel DTW | โ | โ | 45.2 ms |
| this model + log-mel, DTW | 43.8 / 20.3 ms / 49.5% | 3.9 ms | 42.0 ms |
| this model + log-mel, DTW method M1 | 39.8 / 14.8 ms / 58.4% | 3.2 ms | 43.0 ms (C++ engine) |
On the held-out suite D the model is statistically tied with HuBERT-base features (12 cases); the gain on R comes mostly from the DTW method and did not carry over to D. Suite R was used for model selection. Commercial alignment tools were not evaluated. Full tables, caveats (including an invalid annotation in the Choral Singing Dataset that affects 2 of the 48 R cases) and the scripts to reproduce everything are in the GitHub repository.
Licence
- Model weights: PolyForm Noncommercial License 1.0.0 (
LICENSE). Noncommercial use (research, personal, educational, by noncommercial organisations) is permitted under its terms. - The HuBERT-base encoder weights contained in the files remain under the Apache License 2.0; see
NOTICEand the appendix ofLICENSE. - Commercial licensing of the weights is available from Rys Up Audio: https://rysupaudio.com/pages/contact-us
Required Notice: Copyright 2026 Rys Up Audio LLC (https://rysupaudio.com)
Citation
Please cite the repository (https://github.com/rysupaudio-lab/rysupalign-bench, CITATION.cff) and
HuBERT (Hsu et al., IEEE/ACM TASLP 2021, arXiv:2106.07447) and VocalSet (Wilkins et al., ISMIR 2018).
Model tree for rysupaudio/rysupalign-learned-v6
Base model
facebook/hubert-base-ls960