VyvoUp EfficientHBR v2.0
VyvoUp is a compact causal speech-bandwidth-extension model. It accepts mono
24 kHz speech and returns exactly aligned mono 48 kHz speech. This release is
the validation-promoted successor to the immutable v1.0-real-pcm tag. It
uses a conservatively retained 3% of a synthetic-speech adaptation update;
the other 97% of its trainable weights remain at the real-PCM v1 checkpoint.
The release includes the model, checksum sidecar, fixed-policy outputs, complete validation and test reports, failed adaptation audits, synthetic-data provenance, RTX 5070 Ti benchmarks, and all-utterance FlashSR/LSD and DNSMOS comparisons.
VoiceSR: Vyvo TTS ile 10 yeni karşılaştırma
Aşağıdaki 10 İngilizce giriş, sabitlenmiş
Vyvo/Vyvo-Multilingual-EN-FT-v0.1
TTS checkpoint'i ile üretildi. Her metin için dört TTS adayı oluşturuldu;
yalnızca 24 kHz girişin Whisper transkriptinde en az kelime hatası veren aday,
iki SR modeli çalıştırılmadan önce seçildi. SR çıktıları seçimi etkilemedi.
VyvoUp yayımlanan 24 kHz PCM16 girişi doğrudan alır. FlashSR yalnızca 16 kHz kabul ettiği için aynı konuşmanın deterministik 24→16 kHz sürümünü alır. Her iki SR çıktısı mono 48 kHz PCM16'dır. Oynat düğmeleri doğrudan bu model deposundaki dosyaları açar; bu örnekler Space sayfasına bağlı değildir.
| # | Metin | Vyvo TTS girişi, 24 kHz | VyvoUp, 48 kHz | FlashSR, 48 kHz |
|---|---|---|---|---|
| 01 | Clear speech should remain natural when the bandwidth is extended. | |||
| 02 | The quiet morning train arrived exactly at seven thirty. | |||
| 03 | Please place the bright blue folder beside the wooden chair. | |||
| 04 | A quick brown fox jumps over the lazy dog near the river. | |||
| 05 | Engineers compared every signal before choosing the final model. | |||
| 06 | Fresh coffee and warm bread filled the small kitchen. | |||
| 07 | Can you hear the soft rain tapping against the window? | |||
| 08 | Three curious children watched the silver rocket cross the sky. | |||
| 09 | Accurate audio restoration requires careful listening and fair tests. | |||
| 10 | Today we are testing two speech super resolution systems. |
Toplam benzersiz konuşma süresi 37,20 saniye. Sabitlenmiş Whisper-large-v3-turbo anlaşılabilirlik kontrolünde 97 hedef kelime için TTS girişi ve VyvoUp çıktısı 3, FlashSR çıktısı 5 kelime hatası verdi. Bu kontrol algısal ses kalitesini veya üretilen yüksek frekansların doğruluğunu ölçmez. 48 kHz gerçek hedef bulunmadığı için bu bölüm LSD/DNSMOS üstünlük testi değil, doğrudan dinleme karşılaştırmasıdır.
Hemen dinle: 5 giriş + 5 çıktı
Her satırda önce modelin aldığı mono 24 kHz giriş, sonra VyvoUp'ın ürettiği mono 48 kHz çıktı yer alır. Oynat düğmelerine basarak sayfadan ayrılmadan toplam 10 WAV dosyasını dinleyebilirsiniz. Kendi sesinizi yükleyip denemek için canlı demo Space'ini açın.
| Örnek | 24 kHz giriş | 48 kHz VyvoUp çıktısı | Süre |
|---|---|---|---|
| 1 — %10 yüksek-frekans enerjisi | 3.50 sn | ||
| 2 — %30 yüksek-frekans enerjisi | 2.97 sn | ||
| 3 — %50 yüksek-frekans enerjisi | 3.04 sn | ||
| 4 — %70 yüksek-frekans enerjisi | 3.16 sn | ||
| 5 — %90 yüksek-frekans enerjisi | 3.87 sn |
Dosyalar tarayıcı uyumluluğu için PCM16 WAV olarak yayımlandı. Model çıktıları,
yayımlanan giriş WAV'ları tekrar okunarak üretildi; dolayısıyla her çıktı tam
olarak yanındaki girişe karşılık gelir. Seçim politikası, kimlikler ve tüm
SHA-256 değerleri demo/metadata.json içindedir.
Model
| Property | Value |
|---|---|
| Architecture | EfficientHBR, bias-free causal gated temporal convolutions |
| Trainable parameters | 32,112 |
| Input / output | mono 24 kHz / mono 48 kHz |
| Receptive field | 2,059 input samples / 85.79 ms |
| Algorithmic delay | 308 output samples / 6.42 ms |
| Estimated compute | 801.288 million MAC/s |
| Streaming state (FP32) | 398,932 bytes |
| Selected checkpoint | 3% v2-adapted + 97% v1-real trainable parameters |
| Checkpoint SHA-256 | 133a1d019d05c9a146f3ea618b868021268a58842f09841f85f900e5b05774c7 |
The neural network predicts two interleaved high-band residual phases. A fixed complementary FIR path preserves the observed band and adds only a high-pass residual. Exact digital silence maps to exact silence.
Data
The real-speech source is the official CSTR VCTK Corpus 0.92 archive, microphone 1, licensed CC BY 4.0. The speaker-disjoint real split contains 10.0001 hours for training, 0.5001 hours for validation, and 0.5008 hours for the held-out minimum-phase test.
Synthetic targets were generated as exact mono 48 kHz audio. All model and
implementation revisions are pinned in data/synthetic-*-generation.json.
| Round | 48 kHz TTS sources | Nominal train | Unique usable train |
|---|---|---|---|
| 1 | dots.tts SOAR, Freya, Irodori v4.1 Small | 10.0057 h | 8.6692 h |
| 2 | dots.tts MF 2-step, Freya, Irodori v4.1 Small | 5.0063 h | 5.0063 h |
| Total | three independently implemented model families | 15.0120 h | 13.6755 h |
Round 1 contained 495 byte-identical repeated Freya targets totaling 1.3365 hours. The audit detected and removed them before combined training. Round 2 rejected duplicates globally during generation. The resulting combined train set has 23.6756 unique hours: 10.0001 real, 8.6692 round-1 synthetic, and 5.0063 round-2 synthetic.
Pinned round-2 sources, current at generation time (2026-08-13):
dots-studio/dots.tts-mf-2stepsrevision159b33d33de0f9610d9ea73725a0820d27261fd7, native 48 kHz.freyavoice/freya-ttsrevisiond124e07493615208f58bdd21d432736849ee4230, native 48 kHz.Aratako/Irodori-TTS-v4.1-Smallrevision2b28324dc263ed5e6638b3cf3dd94c82ead07b4b, 48 kHz output. Its SilentCipher watermark path internally uses 44.1 kHz and restores 48 kHz; this is disclosed because it is not an untouched native-48 kHz path.
Train/validation/test voice controls and target text are disjoint. Training
inputs use Kaiser-sinc and zero-phase Chebyshev-II degradations; tests use an
unseen minimum-phase degradation. Generator source snapshots and SHA-256
hashes are under data/source-snapshots/.
Improvement history
The initial real-only model trained for 80,000 steps on one RTX 5070 Ti. Two subsequent gradient-training attempts were kept as evidence but not promoted:
- Synthetic-heavy v2: 60,000 steps, 18m20.9s. It greatly improved synthetic validation but regressed unseen-real LSD, so its test splits stayed sealed.
- Real-replay v3: 5,000 steps with an explicit 80% real / 12% round-1 / 8% round-2 crop mix, 98.42s. It passed all safety gates but still exceeded the real-domain regression cap, so it was rejected without opening tests.
- Validation soup v4: a predeclared 1–10% interpolation grid between v1 and the strongest v2 validation checkpoint. The 3% candidate was the best one satisfying every hard gate and was promoted before test data was opened.
All rejected and promoted selection reports are in training/.
Validation and held-out tests
Lower 12–24 kHz log-spectral distance (LSD) is better. The locked validation score weights unseen-real VCTK 0.60, round-1 synthetic 0.25, and round-2 synthetic 0.15. Every validation set also required zero clipping and worst-case protected-band error at or below -115 dB.
| Validation set | v1 LSD | v2.0 LSD | Change |
|---|---|---|---|
| Unseen-real VCTK | 8.7003 | 8.7757 | +0.0754 dB |
| Round-1 synthetic | 41.7658 | 41.2764 | -0.4894 dB |
| Round-2/latest synthetic | 42.2015 | 41.7082 | -0.4934 dB |
| Weighted score | 21.9919 | 21.8408 | -0.1511 dB |
The synthetic tests were opened exactly once after checkpoint selection. The VCTK test had already been opened for v1 and is only a disclosed regression set for this release.
| Test set | v1 LSD | v2.0 LSD | Result |
|---|---|---|---|
| VCTK, unseen speakers + degradation (527 / 30.05 min) | 8.3674 | 8.3903 | +0.0229 dB |
| Round-1 synthetic (171 / 30.40 min) | 41.5747 | 41.0840 | 0.4907 dB better |
| Round-2/latest synthetic (65 / 15.57 min) | 42.8360 | 42.3354 | 0.5005 dB better |
All final test sets had zero clipping. Worst protected-band errors were
-126.25, -126.56, and -135.29 dB respectively. The complete per-item evidence
and the test-opening declaration are in metrics/.
Synthetic-domain LSD remains much higher than VCTK LSD and the model over-produces synthetic high-band energy by roughly 7–8 dB on average. This is an important limitation, not hidden by the relative improvement.
FlashSR comparison
The pinned official repository at ysharma3501/FlashSR, implementation
revision 2a69326250613c0a0f6c1c8d9f0c48cb779842b8, and checkpoint SHA-256
62c70874ac4efeb4dc9c8aa9dc0a611a951e1c36292abeb4c406d7fb91e0eefc
were evaluated on all 527 VCTK test utterances (30.05 minutes).
| Protocol | VyvoUp LSD | FlashSR LSD | VyvoUp margin |
|---|---|---|---|
| Native inputs: 24 kHz vs 16 kHz | 8.3320 | 11.7206 | 3.3886 dB |
| Same 16 kHz information | 9.8090 | 11.7206 | 1.9116 dB |
| Native, RMS-gain aligned diagnostic | 8.3300 | 12.9994 | 4.6695 dB |
| Same information, RMS-gain aligned | 9.7950 | 12.9994 | 3.2044 dB |
VyvoUp has 32,112 trainable parameters; the loaded FlashSR graph has 131,904
parameters (88,128 active in its forward path). Mean end-to-end RTF in this
comparison was 0.004044 for VyvoUp and 0.004042 for FlashSR. FlashSR's official
wrapper peak-normalizes every clip, so both raw wrapper output and a
gain-aligned diagnostic are shown. The linked FlashSR repository implements a
small HiFi-GAN/HierSpeech++-style upsampler, not the diffusion model from the
similarly named paper. Full caveats, per-item metrics, and four paired listening
sets are in comparison/.
DNSMOS P.835 diagnostic
The pinned local models from Microsoft DNS-Challenge were also run on the same 527 held-out utterances using the regular, non-personalized DNSMOS P.835 protocol. Higher scores are better. The main table below matches each output's RMS to its target with one scalar gain, preventing FlashSR's official per-clip peak normalization from dominating the result.
| Protocol / system | SIG | BAK | OVRL | P.808 |
|---|---|---|---|---|
| VyvoUp, native 24 kHz input | 3.5174 | 4.0397 | 3.2149 | 3.6400 |
| VyvoUp, same 16 kHz information | 3.5167 | 4.0396 | 3.2142 | 3.6383 |
| FlashSR, 16 kHz input | 3.4883 | 3.9483 | 3.1272 | 3.6124 |
Paired VyvoUp-minus-FlashSR differences and 95% bootstrap intervals:
| Protocol | SIG | BAK | OVRL | P.808 |
|---|---|---|---|---|
| Native | +0.0291 [0.0216, 0.0366] | +0.0915 [0.0820, 0.1011] | +0.0877 [0.0788, 0.0966] | +0.0276 [0.0220, 0.0332] |
| Same 16 kHz information | +0.0284 [0.0209, 0.0359] | +0.0914 [0.0819, 0.1010] | +0.0870 [0.0781, 0.0958] | +0.0259 [0.0205, 0.0313] |
The raw official-wrapper OVRL scores were 3.2151 for VyvoUp native, 3.2144 for
VyvoUp equal-information, and 2.9000 for FlashSR. DNSMOS downsamples every
output to 16 kHz, removing the synthesized 8–24 kHz band. It is therefore a
low-band speech/noise-quality diagnostic, not a direct high-band reconstruction
metric or a subjective listening result. The 12–24 kHz LSD comparison remains
primary. Full distributions, caveats, pinned Microsoft revision/model hashes,
and all 527 per-item records are in
comparison/dnsmos/.
Aynı 10 konuşma: üç çıkışı doğrudan dinle
DNSMOS ses üretmez; puanlama modelidir. Aşağıdaki üç sütun iki bağımsız modelin üç çalışma yolunu gösterir: VyvoUp doğal 24 kHz giriş, aynı VyvoUp checkpoint'i ile eş-bilgi 16 kHz yolu ve FlashSR 16 kHz yolu. Eş-bilgi VyvoUp ile FlashSR aynı 16 kHz WAV'dan başlar.
Yayımlanan PCM16 dosyalar üzerindeki ortalamalar:
| Çıkış | SIG ↑ | BAK ↑ | OVRL ↑ | P.808 ↑ | 12–24 kHz LSD ↓ |
|---|---|---|---|---|---|
| VyvoUp — doğal 24 kHz | 3.5624 | 4.0752 | 3.2724 | 3.5796 | 8.3094 |
| VyvoUp — aynı 16 kHz bilgi | 3.5622 | 4.0742 | 3.2717 | 3.5779 | 9.7898 |
| FlashSR — 16 kHz | 3.4984 | 4.0020 | 3.1655 | 3.5756 | 11.4514 |
Karşılaştırma demosunu açın veya tam girişleri, hedefleri, 60 oynatıcıyı ve dosya bazlı ölçümleri bu model deposunda açın.
| # | VyvoUp doğal | VyvoUp aynı bilgi | FlashSR |
|---|---|---|---|
| 01 | |||
| 02 | |||
| 03 | |||
| 04 | |||
| 05 | |||
| 06 | |||
| 07 | |||
| 08 | |||
| 09 | |||
| 10 |
Doğal/eş-bilgi VyvoUp, LSD'de sırasıyla 10/10 ve 9/10; RMS-hizalı DNSMOS
OVRL'de iki yol da 8/10 örnekte FlashSR'den daha iyi çıktı. 70 WAV'ın tamamı
mono PCM16, hedefle tam hizalı, benzersiz SHA-256'ya sahip ve clipping/saturation
içermiyor. metadata.json bütün
dosya hash'lerini, giriş sözleşmelerini ve örnek bazlı sonuçları içerir.
RTX 5070 Ti benchmark
One second of audio, BF16, batch one, 20 warmups and 200 measured iterations:
| Path | Mean | p95 | RTF | Peak allocated |
|---|---|---|---|---|
| Eager | 1.913 ms | 1.943 ms | 0.001913 | 26.0 MB |
torch.compile |
1.086 ms | 1.205 ms | 0.001086 | 14.5 MB |
Compilation took 2.57 seconds and is excluded from measured latency. See
benchmarks/ for environment and boundary details.
Example outputs
Examples follow a fixed policy: 10th, 50th, and 90th percentiles of target high-band energy, plus the hardest candidate LSD. Each directory contains the 24 kHz input, sinc baseline, model output, and 48 kHz target.
| Selection | Input | Baseline | VyvoUp v2.0 | Target |
|---|---|---|---|---|
| 10th percentile | input | sinc | output | target |
| Median | input | sinc | output | target |
| 90th percentile | input | sinc | output | target |
| Hardest LSD | input | sinc | output | target |
examples/metadata.json binds every WAV to its SHA-256 and evaluation report.
Inference
Install this VyvoUp source checkout and huggingface_hub, then download both
the checkpoint and its mandatory checksum sidecar:
import numpy as np
import soundfile as sf
import torch
from huggingface_hub import hf_hub_download
from audio_upscaler.models import load_efficient_hbr
repo_id = "kadirnar/VyvoUp"
checkpoint = hf_hub_download(repo_id, "model.pt")
hf_hub_download(repo_id, "model.json")
audio, sample_rate = sf.read("speech-24khz.wav", dtype="float32")
if sample_rate != 24_000 or audio.ndim != 1:
raise ValueError("Input must be mono 24 kHz audio")
model = load_efficient_hbr(checkpoint, device="cuda")
inputs = torch.from_numpy(np.asarray(audio)).view(1, 1, -1).cuda()
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
output = model.forward_aligned(inputs)
sf.write("speech-48khz.wav", output.float().cpu().numpy()[0, 0], 48_000)
forward_aligned compensates FIR delay and returns exactly 2 * input_samples.
For bounded-state deployment, use stream and flush.
CPU tabanlı demoda kullanılan dinamik-uzunluk TorchScript dosyası
model.torchscript.pt olarak ayrıca yayımlanır. Dosya kimliği ve giriş/çıkış
sözleşmesi model.torchscript.json içindedir.
Canlı demo, sunucu veya
GPU kullanmadan model.onnx dosyasını ziyaretçinin tarayıcı CPU'sunda çalıştırır.
ONNX grafiği dinamik ses uzunluğunu destekler; üç farklı uzunlukta ONNX Runtime
CPU doğrulaması, PyTorch'a karşı en fazla 2.24e-7 mutlak hata ve tam 2:1 çıktı
uzunluğu verdi. Ayrıntılar model.onnx.json içindedir.
Intended use, limits, and license
- Research use for 24→48 kHz mono speech bandwidth extension.
- English real-speech training is limited to VCTK. Other languages, music, environmental audio, stereo, telephony, and arbitrary rates are out of scope.
- Generated high frequencies are plausible estimates, not recovered evidence; do not use them for forensic claims.
- Objective metrics do not replace blinded listening or deployment-domain safety testing. No subjective-quality superiority claim is made.
- Irodori-generated data inherits upstream ethical-use restrictions; do not use this work for impersonation, deception, misinformation, or rights violations.
Weights and bundled VCTK-derived listening examples are released under CC BY 4.0 to preserve VCTK attribution. Attribute CSTR VCTK Corpus 0.92 and VyvoUp when redistributing them. The source-code repository did not declare a software license at release time; this weight license grants no additional source-code rights.
VCTK citation: Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald, CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit, version 0.92 (2019), DOI: 10.7488/ds/2645.
- Downloads last month
- -