Download README.md from NYCU-MLLab/Bidirectional_ASR-TTS_Flow_Matching: direct link, hf CLI and curl.
- Browser
- Download file 1.43 kB
-
https://huggingface.co/NYCU-MLLab/Bidirectional_ASR-TTS_Flow_Matching/resolve/main/README.md
- Command line
-
hf download hf://NYCU-MLLab/Bidirectional_ASR-TTS_Flow_Matching/README.md
-
curl -L -o README.md https://huggingface.co/NYCU-MLLab/Bidirectional_ASR-TTS_Flow_Matching/resolve/main/README.md
library_name: pytorch
tags:
- bidirectional-flow-matching
- text-to-image
- image-to-text
- text-to-speech
- automatic-speech-recognition
BVFM model weights
This model repository contains the deployment weights for the image/text and speech/text branches of Bidirectional Variational Flow Matching (BVFM). Source code, training/inference entry points, and environment instructions live in the separate GitHub repository.
Layout
image/
bvfm_image_step40000.pt
FlowTiTok_512.bin
README.md
speech/
bvfm_speech_step299999_inference.pt
merged_config.json
semantic_vae_1000k/
config.json
metainfo.json
dac/ema_state_dict.pth
README.md
The repository intentionally excludes training optimizer states, periodic snapshots, logs, metrics, generated samples, datasets, and caches.
Clone with Git LFS or use huggingface-cli download. Point the GitHub code at
the downloaded root with:
export BVFM_WEIGHTS_ROOT=/path/to/this/model/repository
See the branch-specific READMEs for exact file roles and checksums.
Third-party components
image/FlowTiTok_512.bin is the released FlowTiTok image tokenizer required
by image encoding and decoding. speech/semantic_vae_1000k is the released
Semantic-VAE decoder required to map the 64-D speech latent to waveform. See
THIRD_PARTY_NOTICES.md and retain the corresponding upstream notices when
redistributing these files.