You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

WavCochCausalV8192-vocoder

WavCoch is a causal waveform-to-cochleagram tokenizer by Greta Tuckute and Klemen Kotar.

Model Details

Parameter Value
Parameters ~24.42M
Window Size 1001
Hop Length 80
Encoder Dim 512
Vocabulary Size 8192
Includes Vocoder True

Usage

from transformers import AutoModel

wavcoch = AutoModel.from_pretrained(
    "TuKoResearch/WavCochCausalV8192-vocoder",
    trust_remote_code=True,
)

codes = wavcoch.quantize(waveform_tensor)
coch = wavcoch.decode(codes)
embeddings = wavcoch(
    input_values=waveform_tensor,
    output_hidden_states=True,
    sampling_rate=16000,
).hidden_states[0]

audio = wavcoch.decode_audio(codes)

Notes

This repo includes a bundled vocoder and supports decode_audio(...) for end-to-end waveform synthesis.

When called with output_hidden_states=True, WavCoch exposes a single hidden-state layer: the post-FSQ projected embedding sequence used for direct probing.

Loading audio and video files

load_audio() is an optional utility available on both the WavCoch class and loaded instances. It requires the FFmpeg executable on PATH, installed, for example, with conda install -c conda-forge ffmpeg. FFmpeg is only required when this helper is called.

wav = wavcoch.load_audio("random.mp3")  # or .wav, .flac, .m4a, .mp4, .mov, etc.
# wav is a 1D float32 CPU tensor, mono, sampled at 16,000 Hz.
codes = wavcoch.quantize(wav.to(wavcoch.device))

# Optional: decode/downmix/resample without changing the decoded RMS.
wav = wavcoch.load_audio("recording.mp4", normalize_rms=False)

# The same helper works without constructing another model:
wav = type(wavcoch).load_audio("random.mp3")  # WavCoch.load_audio(...) also works.

The helper decodes the first audio stream, downmixes to mono using FFmpeg, and resamples to 16 kHz. It supports the audio/video formats and codecs available in your FFmpeg build; video frames are ignored. Inputs are local file paths, and the complete audio track is loaded into memory. Missing or undecodable audio raises an error.

RMS is measured across the complete waveform after downmixing/resampling. The reference RMS is 0.11 and the allowed range is 0.0011 to 11.0 (0.01x to 100x the reference). Audio within that range is unchanged. Nonzero audio outside it is scaled to the nearest boundary, preserving its relative dynamics; exact silence stays silent. This is not peak normalization: samples are not clipped to [-1, 1]. Use normalize_rms=False to disable this adjustment.

Loading does not tokenize, move the waveform onto a GPU, pad or shift the audio, or change the model's causal/centered mode. Existing tensor-based calls remain unchanged and never invoke this helper automatically. On the centered-capable checkpoint, select centered encoding explicitly with wavcoch.quantize(wav.to(wavcoch.device), mode="centered").

The standalone function can also be imported from wavcoch_audio.py when using the runtime source directly. For a local repository checkout:

from prep_scripts.hf_upload.wavcoch_audio import load_audio

wav = load_audio("random.mp3")

To get this feature from Hugging Face, use the updated revision (or main). Older pinned revisions retain their original API.

Downloads last month
255
Safetensors
Model size
24.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TuKoResearch/WavCochCausalV8192-vocoder

Finetunes
1 model