PhysWave / README.md
10wind's picture
Upload PhysWave checkpoints and model card
901de71 verified
|
Raw History Blame Contribute Delete
1.87 kB
metadata
license: apache-2.0
language:
  - en
pipeline_tag: text-to-audio
tags:
  - spatial-audio
  - ambisonics
  - diffusion
  - arxiv:2608.29549

PhysWave

Pretrained checkpoints for PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation (EMNLP 2026).

Code and instructions · Paper · Demos

PhysWave generates single-source spatial audio from an acoustic caption and waypoints, or a natural-language instruction parsed by an API. Output is approximately 10 seconds of 16 kHz first-order ambisonic audio in W, X, Y, Z channel order. Use an ambisonic decoder for spatial listening.

Checkpoints

Both files are required:

File Model
physwave.ckpt Text- and waypoint-conditioned diffusion model
vae.ckpt Four-channel audio VAE

Usage

Follow the installation instructions in the code repository. From its root, download the weights:

python -m pip install huggingface_hub
hf download 10wind/PhysWave physwave.ckpt vae.ckpt --local-dir checkpoints
python -m physwave --condition examples/telephone_static.json --output outputs/telephone.wav

See the code README for Gemini/OpenRouter configuration and more examples.

Citation

If you use PhysWave in your research, please cite:

@inproceedings{yao2026physwave,
  title={PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation},
  author={Yao, Lingfeng and Huang, Chenpei and Yang, Xingke and Geng, Ziye and Luo, Changqing and Wang, Hao and Liu, Jiang and Pan, Miao},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}