PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation
Paper • 2608.29549 • Published
Pretrained checkpoints for PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation (EMNLP 2026).
Code and instructions · Paper · Demos
PhysWave generates single-source spatial audio from an acoustic caption and waypoints, or a natural-language instruction parsed by an API. Output is approximately 10 seconds of 16 kHz first-order ambisonic audio in W, X, Y, Z channel order. Use an ambisonic decoder for spatial listening.
Both files are required:
| File | Model |
|---|---|
physwave.ckpt |
Text- and waypoint-conditioned diffusion model |
vae.ckpt |
Four-channel audio VAE |
Follow the installation instructions in the code repository. From its root, download the weights:
python -m pip install huggingface_hub
hf download 10wind/PhysWave physwave.ckpt vae.ckpt --local-dir checkpoints
python -m physwave --condition examples/telephone_static.json --output outputs/telephone.wav
See the code README for Gemini/OpenRouter configuration and more examples.
If you use PhysWave in your research, please cite:
@inproceedings{yao2026physwave,
title={PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation},
author={Yao, Lingfeng and Huang, Chenpei and Yang, Xingke and Geng, Ziye and Luo, Changqing and Wang, Hao and Liu, Jiang and Pan, Miao},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}