SOTA-MOS: speech MOS at 16, 24 and 48 kHz
SOTA-MOS predicts the mean opinion score (1–5) of synthetic speech as listeners rate it in a test that mixes sampling rates. It is built for AudioMOS Challenge 2025 Track 3. Code, training and every experiment: Scicom-AI-Enterprise-Organization/SOTA-MOS.
It beats every Track 3 system, including the winner HighRateMOS, on 7 of 8 official metrics: utt MSE, utt LCC, utt SRCC, utt KTAU, sys LCC, sys SRCC, sys KTAU.
| eval (400 clips, 20 conditions) | utt MSE | utt LCC | utt SRCC | utt KTAU | sys MSE | sys LCC | sys SRCC | sys KTAU |
|---|---|---|---|---|---|---|---|---|
| SOTA-MOS (this model) | 0.208 | 0.888 | 0.803 | 0.623 | 0.101 | 0.989 | 0.968 | 0.884 |
| SOTA-MOS, train labels only | 0.252 | 0.846 | 0.718 | 0.534 | 0.114 | 0.970 | 0.910 | 0.779 |
| HighRateMOS (T17, challenge winner, paper) | 0.303 | 0.847 | 0.742 | 0.556 | 0.116 | 0.982 | 0.955 | 0.842 |
| HighRateMOS, our replication (3 seeds) | 0.264 | 0.838 | 0.686 | 0.492 | 0.130 | 0.960 | 0.959 | 0.842 |
| faster-UTMOSv2, off the shelf | 0.328 | 0.785 | 0.753 | 0.576 | 0.153 | 0.952 | 0.893 | 0.758 |
| best published Track 3 system, per metric | 0.238 | 0.847 | 0.742 | 0.589 | 0.056 | 0.982 | 0.955 | 0.842 |
Eval labels were never used to build or select the model. Members were chosen by greedy forward selection on out-of-fold predictions of the dev set, by a rule fixed before any of our systems was scored on eval. The eval set was scored once.
SOTA-MOS keeps the order of the 20 conditions across sampling rates. The replicated HighRateMOS puts every good condition near 3.6. Ours spreads them out, and it puts 16 kHz audio below the same system at 24 or 48 kHz, as the listeners did.
Use
git clone https://github.com/Scicom-AI-Enterprise-Organization/SOTA-MOS && cd SOTA-MOS
uv sync --extra serve
hf download Scicom-intl/HighRateMOS-VoiceMOS2025 --include "model/*" --local-dir .
uv run python -m sotamos.predict --system model/system.json clip_16k.wav clip_28k.wav clip_44k.wav clip_48k.wav clip_48k_world.wav
The same natural sentence at four rates, and a WORLD-vocoded copy at 48 kHz:
file,sampling_rate,mos
clip_16k.wav,16000,3.7021
clip_28k.wav,24000,3.9625
clip_44k.wav,48000,3.7948
clip_48k.wav,48000,3.8229
clip_48k_world.wav,48000,2.3365
From Python:
from sotamos.predict import SystemScorer, load_audio
scorer = SystemScorer("model/system.json", "cuda")
x, sr, x16 = load_audio("clip.wav") # any rate: resampled to the nearest of 16 / 24 / 48 kHz
print(scorer(x, sr, x16)) # MOS on the mixed-rate listening-test scale
Sampling rates
16, 24 and 48 kHz are scored as they are. Any other rate goes to the nearest of the three:
| input | scored at |
|---|---|
| 8 / 11.025 kHz | 16 kHz |
| 22.05 kHz | 24 kHz |
| 28 / 32 kHz | 24 kHz |
| 44.1 kHz | 48 kHz |
| 88.2 / 96 kHz | 48 kHz |
Serve
A FastAPI server with dynamic batching. It follows faster-UTMOSv2/serving, with batches built by clip length because SOTA-MOS scores the whole clip.
cd serving
SYSTEM=../model/system.json MAX_BATCH=16 PP_WORKERS=16 bash run_serve.sh
curl -X POST http://127.0.0.1:8000/predict -F file=@clip_44k.wav
# {"mos": 3.7948, "reps": 1, "gpu_ms": 141.17, "total_ms": 161.6, "batch_size": 1,
# "input_sampling_rate": 44100, "sampling_rate": 48000, "duration_s": 5.599}
One GPU, the 400 Track 3 eval clips (1,506 s of audio), same client for both servers, run back to back:
| concurrency | SOTA-MOS req/s | RTF | p50 | p95 | faster-UTMOSv2 req/s | RTF | p50 | p95 |
|---|---|---|---|---|---|---|---|---|
| 1 | 9.7 | 37× | 101 ms | 126 ms | 8.9 | 33× | 109 ms | 128 ms |
| 4 | 30.1 | 113× | 128 ms | 155 ms | 29.4 | 111× | 124 ms | 168 ms |
| 8 | 38.9 | 147× | 205 ms | 299 ms | 33.2 | 125× | 142 ms | 1400 ms |
| 16 | 44.9 | 169× | 339 ms | 469 ms | 27.1 | 102× | 316 ms | 1696 ms |
| 32 | 54.1 | 204× | 596 ms | 657 ms | 53.7 | 202× | 546 ms | 1235 ms |
| 64 | 63.3 | 238× | 866 ms | 1663 ms | 63.6 | 240× | 949 ms | 1375 ms |
SOTA-MOS serves as fast as faster-UTMOSv2 and holds up better at mid load. faster-UTMOSv2 scores one fold on a 3-second crop. SOTA-MOS scores the whole clip with 3 models × 5 folds. Served scores match the offline scores within 0.007 MOS on average (Pearson r ≥ 0.9998 at batch 1, 8 and 32).
The server is a drop-in replacement for faster-UTMOSv2's server. Same routes, same file field, same ?reps= and ?dataset= parameters
(accepted; SOTA-MOS is deterministic and has no data-domain), same response keys in the same order and types, same status codes.
It adds three keys after them: input_sampling_rate, sampling_rate (the trained rate the clip was scored at) and duration_s.
SOTA-MOS: {"mos": 3.7948, "reps": 1, "gpu_ms": 141.17, "total_ms": 161.6, "batch_size": 1, "input_sampling_rate": 44100, "sampling_rate": 48000, "duration_s": 5.599}
faster-UTMOSv2: {"mos": 3.21875, "reps": 1, "gpu_ms": 130.26, "total_ms": 150.69, "batch_size": 1}
serving/check_api_compat.py runs both servers side by side and checks every route: POST /predict (with and without reps/dataset),
invalid reps (422), a missing file (422), non-audio bytes (500), /health, /stats, /warmup and /. All pass.
faster-UTMOSv2's own benchmark.py runs against the SOTA-MOS server unchanged.
Model
Three frozen self-supervised speech models, each followed by a small trained head, 5 cross-validation folds each:
| backbone (frozen) | layers kept | folds |
|---|---|---|
| facebook/data2vec-audio-large | first 4 | 5 |
| facebook/wav2vec2-xls-r-300m | first 10 | 5 |
| facebook/wav2vec2-xls-r-1b | first 14 | 5 |
Each head combines the backbone's early layers (weighted sum over the kept layers), a multi-scale CNN on the clip's native-rate mel spectrogram on a 0–24 kHz axis, and learned embeddings for the sampling rate and the listening test. The mel branch is what lets a 16 kHz clip look different from a 48 kHz one. Training used the Track 3 train labels (single-rate tests) and the dev labels (mixed-rate test), with the listening test as an input.
The mixed-rate test drops 16 kHz audio by 0.35 MOS against its single-rate test. Train labels alone cannot teach that. The replication, the ablations and the frozen-feature probes are in the GitHub README.
Files
| path | contents |
|---|---|
model/ |
the SOTA-MOS ensemble: system.json, fold heads, model card |
track3_obf.tar.gz |
Track 3 train/dev audio and train labels, as distributed by the organisers |
audiomos2025-track3-eval-phase.zip |
Track 3 eval audio |
track3_post_eval_distro.tar.gz |
post-evaluation release: all 800 clips, train / dev / eval ratings per listener |
amc2025_track3_val_answer.zip |
dev system-level answers |
assets/ |
figures used on this page |
References
- W. Ren et al., HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment, 2025.
- W.-C. Huang et al., The AudioMOS Challenge 2025, 2025.
- K. Baba et al., UTMOSv2, 2024; faster-UTMOSv2.




