--- license: apache-2.0 language: - en pipeline_tag: text-to-audio tags: - spatial-audio - ambisonics - diffusion - arxiv:2608.29549 --- # PhysWave Pretrained checkpoints for **PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation** (EMNLP 2026). [Code and instructions](https://github.com/lingfengyao/PhysWave) · [Paper](https://arxiv.org/abs/2608.29549) · [Demos](https://lingfengyao.github.io/PhysWave/) PhysWave generates single-source spatial audio from an acoustic caption and waypoints, or a natural-language instruction parsed by an API. Output is approximately 10 seconds of 16 kHz first-order ambisonic audio in W, X, Y, Z channel order. Use an ambisonic decoder for spatial listening. ## Checkpoints Both files are required: | File | Model | | --- | --- | | `physwave.ckpt` | Text- and waypoint-conditioned diffusion model | | `vae.ckpt` | Four-channel audio VAE | ## Usage Follow the installation instructions in the code repository. From its root, download the weights: ```bash python -m pip install huggingface_hub hf download 10wind/PhysWave physwave.ckpt vae.ckpt --local-dir checkpoints python -m physwave --condition examples/telephone_static.json --output outputs/telephone.wav ``` See the [code README](https://github.com/lingfengyao/PhysWave#natural-language-input) for Gemini/OpenRouter configuration and more examples. ## Citation If you use PhysWave in your research, please cite: ```bibtex @inproceedings{yao2026physwave, title={PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation}, author={Yao, Lingfeng and Huang, Chenpei and Yang, Xingke and Geng, Ziye and Luo, Changqing and Wang, Hao and Liu, Jiang and Pan, Miao}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing}, year = {2026} } ```