audio
speech
mos
speech-quality-assessment
audiomos
sampling-rate

SOTA-MOS: speech MOS at 16, 24 and 48 kHz

SOTA-MOS predicts the mean opinion score (1–5) of synthetic speech as listeners rate it in a test that mixes sampling rates. It is built for AudioMOS Challenge 2025 Track 3. Code, training and every experiment: Scicom-AI-Enterprise-Organization/SOTA-MOS.

It beats every Track 3 system, including the winner HighRateMOS, on 7 of 8 official metrics: utt MSE, utt LCC, utt SRCC, utt KTAU, sys LCC, sys SRCC, sys KTAU.

eval (400 clips, 20 conditions) utt MSE utt LCC utt SRCC utt KTAU sys MSE sys LCC sys SRCC sys KTAU
SOTA-MOS (this model) 0.208 0.888 0.803 0.623 0.101 0.989 0.968 0.884
SOTA-MOS, train labels only 0.252 0.846 0.718 0.534 0.114 0.970 0.910 0.779
HighRateMOS (T17, challenge winner, paper) 0.303 0.847 0.742 0.556 0.116 0.982 0.955 0.842
HighRateMOS, our replication (3 seeds) 0.264 0.838 0.686 0.492 0.130 0.960 0.959 0.842
faster-UTMOSv2, off the shelf 0.328 0.785 0.753 0.576 0.153 0.952 0.893 0.758
best published Track 3 system, per metric 0.238 0.847 0.742 0.589 0.056 0.982 0.955 0.842

Eval labels were never used to build or select the model. Members were chosen by greedy forward selection on out-of-fold predictions of the dev set, by a rule fixed before any of our systems was scored on eval. The eval set was scored once.

All eight metrics

Per-condition predictions

SOTA-MOS keeps the order of the 20 conditions across sampling rates. The replicated HighRateMOS puts every good condition near 3.6. Ours spreads them out, and it puts 16 kHz audio below the same system at 24 or 48 kHz, as the listeners did.

Use

git clone https://github.com/Scicom-AI-Enterprise-Organization/SOTA-MOS && cd SOTA-MOS
uv sync --extra serve
hf download Scicom-intl/HighRateMOS-VoiceMOS2025 --include "model/*" --local-dir .
uv run python -m sotamos.predict --system model/system.json clip_16k.wav clip_28k.wav clip_44k.wav clip_48k.wav clip_48k_world.wav

The same natural sentence at four rates, and a WORLD-vocoded copy at 48 kHz:

file,sampling_rate,mos
clip_16k.wav,16000,3.7021
clip_28k.wav,24000,3.9625
clip_44k.wav,48000,3.7948
clip_48k.wav,48000,3.8229
clip_48k_world.wav,48000,2.3365

From Python:

from sotamos.predict import SystemScorer, load_audio

scorer = SystemScorer("model/system.json", "cuda")
x, sr, x16 = load_audio("clip.wav")   # any rate: resampled to the nearest of 16 / 24 / 48 kHz
print(scorer(x, sr, x16))              # MOS on the mixed-rate listening-test scale

Sampling rates

16, 24 and 48 kHz are scored as they are. Any other rate goes to the nearest of the three:

input scored at
8 / 11.025 kHz 16 kHz
22.05 kHz 24 kHz
28 / 32 kHz 24 kHz
44.1 kHz 48 kHz
88.2 / 96 kHz 48 kHz

Serve

A FastAPI server with dynamic batching. It follows faster-UTMOSv2/serving, with batches built by clip length because SOTA-MOS scores the whole clip.

cd serving
SYSTEM=../model/system.json MAX_BATCH=16 PP_WORKERS=16 bash run_serve.sh
curl -X POST http://127.0.0.1:8000/predict -F file=@clip_44k.wav
# {"mos": 3.7948, "reps": 1, "gpu_ms": 141.17, "total_ms": 161.6, "batch_size": 1,
#  "input_sampling_rate": 44100, "sampling_rate": 48000, "duration_s": 5.599}

Serving

Serving latency

One GPU, the 400 Track 3 eval clips (1,506 s of audio), same client for both servers, run back to back:

concurrency SOTA-MOS req/s RTF p50 p95 faster-UTMOSv2 req/s RTF p50 p95
1 9.7 37× 101 ms 126 ms 8.9 33× 109 ms 128 ms
4 30.1 113× 128 ms 155 ms 29.4 111× 124 ms 168 ms
8 38.9 147× 205 ms 299 ms 33.2 125× 142 ms 1400 ms
16 44.9 169× 339 ms 469 ms 27.1 102× 316 ms 1696 ms
32 54.1 204× 596 ms 657 ms 53.7 202× 546 ms 1235 ms
64 63.3 238× 866 ms 1663 ms 63.6 240× 949 ms 1375 ms

SOTA-MOS serves as fast as faster-UTMOSv2 and holds up better at mid load. faster-UTMOSv2 scores one fold on a 3-second crop. SOTA-MOS scores the whole clip with 3 models × 5 folds. Served scores match the offline scores within 0.007 MOS on average (Pearson r ≥ 0.9998 at batch 1, 8 and 32).

The server is a drop-in replacement for faster-UTMOSv2's server. Same routes, same file field, same ?reps= and ?dataset= parameters (accepted; SOTA-MOS is deterministic and has no data-domain), same response keys in the same order and types, same status codes. It adds three keys after them: input_sampling_rate, sampling_rate (the trained rate the clip was scored at) and duration_s.

SOTA-MOS:       {"mos": 3.7948, "reps": 1, "gpu_ms": 141.17, "total_ms": 161.6, "batch_size": 1, "input_sampling_rate": 44100, "sampling_rate": 48000, "duration_s": 5.599}
faster-UTMOSv2: {"mos": 3.21875, "reps": 1, "gpu_ms": 130.26, "total_ms": 150.69, "batch_size": 1}

serving/check_api_compat.py runs both servers side by side and checks every route: POST /predict (with and without reps/dataset), invalid reps (422), a missing file (422), non-audio bytes (500), /health, /stats, /warmup and /. All pass. faster-UTMOSv2's own benchmark.py runs against the SOTA-MOS server unchanged.

Model

Three frozen self-supervised speech models, each followed by a small trained head, 5 cross-validation folds each:

backbone (frozen) layers kept folds
facebook/data2vec-audio-large first 4 5
facebook/wav2vec2-xls-r-300m first 10 5
facebook/wav2vec2-xls-r-1b first 14 5

Each head combines the backbone's early layers (weighted sum over the kept layers), a multi-scale CNN on the clip's native-rate mel spectrogram on a 0–24 kHz axis, and learned embeddings for the sampling rate and the listening test. The mel branch is what lets a 16 kHz clip look different from a 48 kHz one. Training used the Track 3 train labels (single-rate tests) and the dev labels (mixed-rate test), with the listening test as an input.

Where the data disagrees

The mixed-rate test drops 16 kHz audio by 0.35 MOS against its single-rate test. Train labels alone cannot teach that. The replication, the ablations and the frozen-feature probes are in the GitHub README.

Files

path contents
model/ the SOTA-MOS ensemble: system.json, fold heads, model card
track3_obf.tar.gz Track 3 train/dev audio and train labels, as distributed by the organisers
audiomos2025-track3-eval-phase.zip Track 3 eval audio
track3_post_eval_distro.tar.gz post-evaluation release: all 800 clips, train / dev / eval ratings per listener
amc2025_track3_val_answer.zip dev system-level answers
assets/ figures used on this page

References

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for Scicom-intl/HighRateMOS-VoiceMOS2025