Dankker0900's picture
Upload BVFM image and speech model weights
3800c9d verified
|
Raw History Blame Contribute Delete
1.09 kB
# BVFM image/text weights
## Files
- `bvfm_image_step40000.pt` (3,230,123,258 bytes): selected bidirectional BVFM
checkpoint at step 40,000. Format
`flowtok_bvfm_shared_variational_v2`; contains the shared vector field,
pair posterior, text/image inference priors, and AR caption decoder. It has
no optimizer state.
- `FlowTiTok_512.bin` (1,560,616,638 bytes): released 512px FlowTiTok tokenizer
required by both T2I decoding and I2T image encoding.
Training-only initializers `FlowTok-XL.pth` and `decoder_init.pt` are not part
of this deployment package.
## SHA-256
```text
4fe414c46bb1b25781178e0f01f71047194fcc3af42eebeaf596f72386f0b6a3 bvfm_image_step40000.pt
9169d1e02b8e8fe1bff8a97796b3f1171415c44a1a125c050f7138805b13ebc8 FlowTiTok_512.bin
```
With the GitHub code, either set `BVFM_WEIGHTS_ROOT` to the downloaded model
repository root or pass the files explicitly:
```bash
export FLOWTITOK_CKPT=/path/to/model-repo/image/FlowTiTok_512.bin
python scripts/infer_t2i.py \
--checkpoint /path/to/model-repo/image/bvfm_image_step40000.pt \
--output-dir runs/demo
```