Access SAWT V5

SAWT V5 is free for non-commercial use under the SAWT License 2.0. It is not available to for-profit organisations, and it may not be used to prepare training data for a commercial product. Two quick fields and one confirmation.

Log in or Sign Up to review the conditions and access this model content.

SAWT V5

SAWT V5

48 kHz speech restoration with a Whisper anchor

Noise, reverb, band limits, codecs and telephone lines in. Full-band 48 kHz speech in the same voice, with the same words, out.

Listening page · Quick start · How it works · Evaluation · Which SAWT to use · SAWT V4


SAWT V5 is the second public SAWT model from Quran-Lab. SAWT exists because a century of Quran recitation survives on tape, in echoing halls and in low-bitrate uploads, and restoring it must not change a syllable. That standard holds for every voice in every language, so SAWT restores general speech too.

V5 keeps what made V4 faithful: a deterministic anchor reads the damaged recording and tells a latent flow-matching generator what was said, by whom and at what pitch, and the generator only renders the texture. What changed is the reader. V5's anchor is Whisper large-v3 (encoder layers 1 to 24, with LoRA adapters), trained to emit from the damaged audio what the frozen encoder emits on the clean recording. It reads more languages and harder damage than V4's anchor did, and it brings a speaker vector with it. The generator was continued from V4 on new, cleaner targets.

Highlights

  • The best SAWT on general speech. On Google's sixteen Miipher-2 demo clips, the DistillMOS quality predictor rises from 3.83 (V4) to 4.18, with V5 ahead of V4 on 14 of the 16 clips. Miipher-2 averages 4.36, but V5 scores higher than it on English, Swahili and Urdu clips in that set, and it returns the full 48 kHz band, where Miipher-2's renders are 24 kHz audio with at most 12 kHz of bandwidth.

  • Faithful. On 200 LibriTTS-R utterances, word error against the reference text is 0.073 (V4 0.077; the unprocessed input reads 0.066), the recognised words of the input change less than with V4 (drift 0.036 vs 0.042), and speaker consistency is 0.990.

  • Recitation content at V4's level. On 120 real archive recitations, phoneme error against the canonical text is 0.074 with the default recipe and 0.071 with the input-start recipe, level with V4 (0.071; the recordings themselves read 0.062), and speaker consistency is higher than V4's (0.954 vs 0.941).

  • Honest about its limits. On archive recitation V5 leaves more background in the pauses than V4 does (median floor -51 vs -70 dBFS). For recitation archives, V4 remains our recommendation; see Which SAWT to use.

Hear it

Loudness-matched, headphones recommended. Every clip below, and six more, are on the listening page.

Urdu (FLEURS), one of Google's Miipher-2 demo clips. The input is a 16 kHz recording with a raised floor. SAWT V5 returns 21 kHz of bandwidth at a quality predictor of 4.53 (input 2.69, Miipher-2 4.42, SAWT V4 3.93), speaker similarity 0.98.

damaged SAWT V5 Miipher-2 (Google)

English, the clip SAWT V4 lost. On this demo clip V4 collapsed (quality predictor 1.59, speaker similarity 0.60). V5 restores it at 4.44 with speaker similarity 0.99.

damaged SAWT V5 SAWT V4

Quran recitation from the archive (surah 38). A damaged archive upload. Quality predictor 2.88 to 4.05 (V4 3.79), and every phoneme the recogniser reads matches what it read in the input.

damaged SAWT V5 SAWT V4

Danish over a telephone line, the opening clip of the V4 card. A 300 to 3700 Hz telephone band comes back as 22 kHz of speech. Quality predictor 3.06 to 4.68 (V4 4.60).

damaged SAWT V5 SAWT V4

Spectrograms: damaged input, SAWT V5 and a reference system for three clips

The lossless versions of all ten clips, with their measurements, are in samples/ (samples.json).

Quick start

pip install -r requirements.txt huggingface_hub
python - <<'EOF'
from huggingface_hub import snapshot_download
snapshot_download("Quran-Lab/sawt-v5", local_dir="sawt-v5", allow_patterns=[
    "sawt2/*", "config.json", "requirements.txt",                      # the restorer
    "generator.pt", "anchor.pt", "vae48_dacvae_d2.pt", "dacvae_stats.pt",  # the weights, 1.96 GB
])
EOF
cd sawt-v5
python sawt2/restore_v5.py --ckpt generator.pt --anchor anchor.pt --vae dacvae@vae48_dacvae_d2.pt \
    --stats dacvae_stats.pt --steps 64 --match-dynamics --in recordings/ --out restored/

That is the evaluated recipe: Euler with 64 steps from unit noise, 5.12 s windows with 1 s of overlap, one noise field per file, the file-level speaker vector, dynamics matched to the input, seed 0. The first run downloads the Whisper large-v3 encoder from openai/whisper-large-v3 (MIT). Inputs can be WAV or FLAC at any sample rate; outputs are 48 kHz mono WAV, each with a JSON sidecar holding the configuration, a fingerprint of the checkpoint, the measured noise floor and the reverb tail.

Useful flags:

  • --start 1 --guard-std 0.70 --guard-sem 0.85, the input-start recipe: the flow starts from the input's own latent plus noise, and a window is redrawn (up to three times) if its energy collapses or the anchor's reading of the restored window drifts from its reading of the input. Phoneme error on archive recitation 0.071 instead of 0.074, at a slightly lower quality predictor.
  • --steps 32 halves the time at a small cost in quality.
  • --no-metrics skips the sidecar's reverb measurements.

Speed. restore_v5.py is the evaluation reference and works one window at a time: about 0.3 times real time at 64 steps on an RTX 4090, a few GB of VRAM. --steps 32 roughly halves the time. It is built for faithfulness checks and archive batches overnight, not for live use; a batched fast path like V4's is not part of this release.

Smaller download. The repository is about 18 GB because it also carries half-precision weights, earlier checkpoints and the full training state. The list above is what restores audio. bf16/ holds the same three weight files in bfloat16 (0.99 GB instead of 1.96; the latent and feature statistics stay in float32), and the loader casts them back up on read. Measured against the float32 files on Google's sixteen demo clips: DistillMOS -0.006, Audiobox PQ +0.001, median waveform agreement 40.6 dB.

python sawt2/restore_v5.py --ckpt bf16/generator.pt --anchor bf16/anchor.pt --vae dacvae@bf16/vae48_dacvae_d2.pt     --stats dacvae_stats.pt --steps 64 --match-dynamics --in recordings/ --out restored/

How it works

SAWT V5 inference pipeline

  1. Representation. Meta's DAC-VAE, frozen, 48 kHz, 128 channels at 25 Hz, with SAWT V4's fine-tuned decoder head (no watermark branch). Restoration happens in this latent.
  2. Anchor. The encoder of Whisper large-v3 (layers 1 to 24), frozen, with rank-64 LoRA adapters on every attention and feed-forward projection and five heads. It reads the damaged recording at 16 kHz and is trained to emit what the frozen encoder emits on the clean recording: the content stream (layer 24), a broad stream (the mean of layers 13 to 24), pitch (log f0, voicing, periodicity), an estimate of the clean latent, and a speaker vector. It is deterministic: the same input always gives the same reading.
  3. Generator. A 372M diffusion transformer (24 layers, width 1024, 16 heads, adaLN, RoPE, qk-norm), started from V4. Inputs: the noised latent, the damaged latent and the anchor's latent estimate; per block the content, broad, pitch and damaged-latent streams; globally the speaker vector, the bandwidth of the input, a room token and a start flag. Flow matching with an x1-parametrised velocity loss. Two starts were trained together: from unit noise (the default) and from the input's own latent plus noise.
  4. Rendering. Per 5.12 s window the anchor reads the input and the generator integrates 64 Euler steps; windows overlap by 1 s and the latent is decoded once per file.

Training

Recipe. The anchor trained for 60k steps. The generator trained data parallel with ZeRO-1 AdamW and an EMA of 0.9999, on batches of 256 crops of 5.12 s. The damage is synthesised on the accelerator every step: reverb by partitioned overlap-save convolution with measured and simulated rooms and halls, noise, hum, band limits and the rest of V4's chain, plus MP3 and AAC versions rendered ahead of time. Clean targets are a fixed crop grid whose latents were encoded once. The released generator went from V4's weights through 50k steps of the V5b recipe and 100k steps of the V5c recipe.

Data. About 1,160 hours of speech and singing in more than 30 languages at 44.1 and 48 kHz (BibleTTS, AISHELL-3, VCTK, Hi-Fi TTS, GTSinger, the OpenSLR high-quality sets, M4Singer, Opencpop, DAPS, JSUT, Expresso, VocalSet, SIWIS, the Arabic Speech Corpus), plus FLEURS-R and LibriTTS-R at 24 kHz with a target-band flag, and Quran recitation: real archive crops whose background floor is below -60 dBFS and the best 100 hours of a cleaned recitation corpus. Targets that were not already clean went through a cleaning chain (dialogue isolation, dereverberation, denoising); for each 5.12 s crop the cleaned version replaced the original only when it was quieter, below -60 dBFS, and its phonemes still matched the original's. Mix: general speech 70 %, 24 kHz speech 10 %, recitation 20 %.

Evaluation

The same instruments and sets as SAWT V4, scored in one look of the final checkpoint (EMA, step 100k). Every quality number has a faithfulness number next to it. All files in eval/.

Google's sixteen Miipher-2 demo clips (FLEURS, MLS and LibriTTS inputs; loudness-matched)

SAWT V5 SAWT V4 Miipher-2 input
DistillMOS 4.18 3.83 4.36 3.28
Audiobox PQ 7.18 6.94 7.39 5.74
speaker consistency 0.977 0.957 0.974
transcript drift (Whisper) 0.446 0.385 0.375
clips where V5 scores higher 14 of 16 3 of 16

LibriTTS-R, 200 utterances, against Google's Miipher renders

SAWT V5 SAWT V4 Miipher input
DistillMOS 4.38 4.36 4.47 3.95
word error vs reference (CTC) 0.073 0.077 0.066
transcript drift (CTC) 0.036 0.042
speaker consistency 0.990 0.988
dropped words 0.7 % 0.7 %

Real archive recitation, 120 files (phoneme error against the canonical text, read by the bundled Quran recogniser)

SAWT V5 V5, input-start recipe SAWT V4 input
phoneme error vs canonical text 0.074 0.071 0.071 0.062
speaker consistency 0.954 0.962 0.941
background floor (median) -50.9 dBFS -45.9 dBFS -69.6 dBFS
speech to floor 28.8 dB 22.8 dB 47.2 dB
DistillMOS 3.48 3.36 3.83 2.77

Over twelve long archive files measured pause by pause, V5 is cleaner than V4 on three, level on two and noisier on seven.

What the numbers say. V5 is the better restorer of general and multilingual speech: higher quality than V4 on almost every clip, the lowest word error and the best speaker consistency of the three systems where they were all measured. Miipher-2 still scores higher on average on its own demo set, and its Whisper transcripts drift less there. On archive recitation V5 keeps the words as well as V4 but leaves more of the background. We traced part of this to the training targets: a screen that judged clips by their quietest frames counted the zero padding of short clips as silence, so about 160 hours of short, slightly noisy clips were used as clean targets. It is fixed for the next model.

Which SAWT to use

material use
general or multilingual speech, interviews, broadcasts, audiobooks, phone recordings SAWT V5
Quran recitation archives where silent pauses matter SAWT V4 (or V5 with the input-start recipe, then compare)
recitation with heavy damage, where words matter most either; check both against the text

Files

file what size
generator.pt 372M flow-matching DiT, EMA at step 100k (V5c) 1.49 GB
anchor.pt LoRA adapters and heads for the Whisper large-v3 encoder (the encoder downloads from Hugging Face) 189 MB
vae48_dacvae_d2.pt fine-tuned DAC-VAE decoder head (the base downloads with the dacvae package) 283 MB
dacvae_stats.pt latent statistics the generator was trained with 3 kB
bf16/ the three weight files in bfloat16, for half the download 0.99 GB
variants/ earlier checkpoints, see below 5.9 GB
sawt2/ restorer, anchor, generator (sawt2/v5/), evaluation instruments
scripts/ the evaluation script
quran_asr_v31/, quran_canon_tokens.json Quran phoneme recogniser (zipformer CTC, int8 ONNX) and canonical token cache, for the recitation measurements 76 MB
samples/ the ten showcase clips, lossless, with their measurements 15 MB
eval/ every metric file of the final look
training/ the state to resume or fine-tune, see below 8.4 GB

Variants. variants/v5c_step050000_ema.pt is V5c halfway: slightly worse on general speech and content, cleaner on recitation (archive floor -57.2 dBFS, DistillMOS 3.59). variants/v5b_step050000_ema.pt and variants/v5b_step100000_ema.pt are V5b, the same recipe before the cleaned speech targets. variants/v5a_step100000_ema.pt is the first V5 run, kept for the record: it learned to pass the input's noise through on recitation, which is what led to V5b.

Training from here

training/ holds the state the released weights came from, for fine-tuning or for continuing the run:

file what
training/generator_v5c/latest.pt the live generator weights at step 100k
training/generator_v5c/opt_rank{0..31}.pt ZeRO-1 optimizer shards (moments, EMA and master weights) for 32 ranks, 2M-element buckets
training/anchor/ the anchor at step 60k (latest.pt, used by the released generator), its EMA, both optimizer shards and the fitted statistics

The trainers themselves are not part of this release. The checkpoints are plain PyTorch state: the generator's generator_v5c/latest.pt loads into sawt2.v5.dit_v5.LatentDiTV5, and each optimizer shard holds the moments, EMA and master weights of a 1/32 slice of every 2M-element parameter bucket, in parameter order.

License

SAWT License 2.0, non-commercial, the same terms as SAWT V4. Run it, study it, build on it, host it, with attribution to Quran-Lab, for non-commercial purposes. No commercial use of any kind and no use by for-profit organisations; restoring or cleaning audio with SAWT to build a training set is use, and anything trained on such a set is a derivative. Open to individuals acting privately, charities and non-profits, educational and research institutions, religious institutions, public archives, libraries, museums and heritage bodies. It may not be used to make a recording say what was not said, to pass off altered audio as an original, or to deceive, impersonate, surveil or harm. Full text in LICENSE.

Upstream components, each under its own terms for its original portions: Whisper large-v3 (OpenAI, MIT), DAC-VAE (Meta, Apache 2.0), torchcrepe (MIT) and the SpeechBrain ECAPA speaker model (Apache 2.0) for the anchor's training targets. Several training corpora are licensed for non-commercial use only (Expresso, GTSinger, M4Singer, Opencpop, DAPS), which the non-commercial terms of this licence match.

Acknowledgements

OpenAI for Whisper, Meta for DAC-VAE, the creators of every corpus listed above, Google for publishing the Miipher demo clips and the LibriTTS-R renders that make paired comparison possible.

Citation

@misc{sawt_v5_2026,
  title  = {SAWT V5: Whisper-anchored speech restoration at 48 kHz},
  author = {Quran-Lab},
  year   = {2026},
  url    = {https://huggingface.co/Quran-Lab/sawt-v5}
}

Built by Quran-Lab, University of Copenhagen. SAWT means voice.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support