Dankker0900's picture
Upload BVFM image and speech model weights
3800c9d verified
|
Raw History Blame Contribute Delete
1.34 kB

BVFM speech/text weights

Files

  • bvfm_speech_step299999_inference.pt (947,317,708 bytes): deployment-only checkpoint extracted from the selected step-299,999 training checkpoint. It contains inference_modules, model config, normalization/tokenizer/speaker metadata, and no optimizer, scaler, RNG, training-module duplicate, or EMA training container.
  • merged_config.json: effective checkpoint configuration and training provenance. Runtime data/model locations can be overridden by the GitHub entry points.
  • semantic_vae_1000k/: released Semantic-VAE EMA decoder metadata and weights used to decode the 40 Hz, 64-D speech latent.

SHA-256

6c4f1974e0e29ce7c4d755f1c49875eda5a8e4a03663d4289f0cd97e52df3429  bvfm_speech_step299999_inference.pt
7c455aa8ab3f7d576b4834f8342558894aafaa61a371b84a9bfa4d10a100e516  semantic_vae_1000k/dac/ema_state_dict.pth

With the GitHub code:

export BVFM_WEIGHTS_ROOT=/path/to/model-repo
python scripts/infer_tts_one.py \
  --ckpt-dir /path/to/model-repo/speech \
  --checkpoint bvfm_speech_step299999_inference.pt \
  --text "A short synthesis example." \
  --ref-wav /path/to/reference.wav \
  --out-dir runs/demo

The large resumable latest.pt, final.pt, and periodic step snapshots are deliberately excluded from the Hugging Face package.