notmax123's picture
Rename to encoder_supertonic3_decoder/ (descriptive names); update model card
b504d3a verified
|
Raw History Blame Contribute Delete
1.95 kB

BlueCodec encoder trained against the frozen official Supertonic-3 vocoder (Sept 2026)

File Content
encoder.safetensors Encoder weights only (encoder.* keys), for BlueCodec.from_pretrained(..., decoder="supertonic3")
ae_290000.pt Training checkpoint at step 290,000 without the decoder (encoder, MPD, MRD, optimizers, schedulers; decoder_source = "supertonic3"), resumable with train_autoencoder.py --encoder_only --decoder supertonic3 --resume
LICENSE.OpenRAIL-M Copy of the license of the Supertonic-3 model this encoder was trained against

The decoder is not in this repository. This encoder uses the official Supertonic-3 vocoder (onnx/vocoder.onnx from Supertone/supertonic-3, revision 3cadd1ee, Supertone Inc., BigScience OpenRAIL-M). The bluecodec package downloads it from the official repo at load time:

from bluecodec import BlueCodec
codec = BlueCodec.from_pretrained("notmax123/blue-codec", decoder="supertonic3")   # needs: pip install onnx
z = codec.encode(audio, edge_pad_chunks=2)   # edge-padded encoding, recommended
y = codec.decode(z)[..., :audio.shape[-1]]

Recipe: encoder initialised from the 1.5M-step BlueCodec encoder, decoder frozen, AdamW lr 8.5e-5 cosine to 1e-6 over 300k steps, 2x64 segments of 61,740 samples, full-band log-mel reconstruction (λ 45) + LSGAN + composite feature matching, 10k-step discriminator warm-up. Details, audit and the end-of-clip edge code (use edge_pad_chunks=2): https://github.com/maxmelichov/blue-codec

License note. This encoder contains none of Supertone's weights, but it was trained through their model, so it may be a "Derivative of the Model" under OpenRAIL-M §1(e). Its use is therefore subject to the use-based restrictions in Attachment A of LICENSE.OpenRAIL-M in addition to the MIT license of BlueCodec's own code and weights.