PhysWave

Pretrained checkpoints for PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation (EMNLP 2026).

Code and instructions · Paper · Demos

PhysWave generates single-source spatial audio from an acoustic caption and waypoints, or a natural-language instruction parsed by an API. Output is approximately 10 seconds of 16 kHz first-order ambisonic audio in W, X, Y, Z channel order. Use an ambisonic decoder for spatial listening.

Checkpoints

Both files are required:

File Model
physwave.ckpt Text- and waypoint-conditioned diffusion model
vae.ckpt Four-channel audio VAE

Usage

Follow the installation instructions in the code repository. From its root, download the weights:

python -m pip install huggingface_hub
hf download 10wind/PhysWave physwave.ckpt vae.ckpt --local-dir checkpoints
python -m physwave --condition examples/telephone_static.json --output outputs/telephone.wav

See the code README for Gemini/OpenRouter configuration and more examples.

Citation

If you use PhysWave in your research, please cite:

@inproceedings{yao2026physwave,
  title={PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation},
  author={Yao, Lingfeng and Huang, Chenpei and Yang, Xingke and Geng, Ziye and Luo, Changqing and Wang, Hao and Liu, Jiang and Pan, Miao},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for 10wind/PhysWave