PhysWave / README.md
10wind's picture
Upload PhysWave checkpoints and model card
901de71 verified
|
Raw History Blame Contribute Delete
1.87 kB
---
license: apache-2.0
language:
- en
pipeline_tag: text-to-audio
tags:
- spatial-audio
- ambisonics
- diffusion
- arxiv:2608.29549
---
# PhysWave
Pretrained checkpoints for **PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation** (EMNLP 2026).
[Code and instructions](https://github.com/lingfengyao/PhysWave) 路 [Paper](https://arxiv.org/abs/2608.29549) 路 [Demos](https://lingfengyao.github.io/PhysWave/)
PhysWave generates single-source spatial audio from an acoustic caption and waypoints,
or a natural-language instruction parsed by an API. Output is approximately 10 seconds
of 16 kHz first-order ambisonic audio in W, X, Y, Z channel order.
Use an ambisonic decoder for spatial listening.
## Checkpoints
Both files are required:
| File | Model |
| --- | --- |
| `physwave.ckpt` | Text- and waypoint-conditioned diffusion model |
| `vae.ckpt` | Four-channel audio VAE |
## Usage
Follow the installation instructions in the code repository. From its root, download the weights:
```bash
python -m pip install huggingface_hub
hf download 10wind/PhysWave physwave.ckpt vae.ckpt --local-dir checkpoints
python -m physwave --condition examples/telephone_static.json --output outputs/telephone.wav
```
See the [code README](https://github.com/lingfengyao/PhysWave#natural-language-input)
for Gemini/OpenRouter configuration and more examples.
## Citation
If you use PhysWave in your research, please cite:
```bibtex
@inproceedings{yao2026physwave,
title={PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation},
author={Yao, Lingfeng and Huang, Chenpei and Yang, Xingke and Geng, Ziye and Luo, Changqing and Wang, Hao and Liu, Jiang and Pan, Miao},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}
```