Buckets:

OpenT2S/OpenDialog / README.md
heyunchao's picture
|
download
raw
1.36 kB
metadata
license: cc-by-nc-4.0
task_categories:
  - text-to-speech
language:
  - en
  - zh

OpenDialog

OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.

OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of:

  • English data: 5074 hours
  • Chinese data: 1759 hours

This dataset is also available at ModelScope (more friendly for users from China mainland).

Citation

@article{zhu2025zipvoicedialog,
      title={ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching},
      author={Zhu, Han and Kang, Wei and Guo, Liyong and Yao, Zengwei and Kuang, Fangjun and Zhuang, Weiji and Li, Zhaoqing and Han, Zhifeng and Zhang, Dong and Zhang, Xin and Song, Xingchen and Lin, Long and Povey, Daniel},
      journal={arXiv preprint arXiv:2507.09318},
      year={2025}
}

Xet Storage Details

Size:
1.36 kB
·
Xet hash:
cd67a54fc46a4f8d416e6bfd7da41294c3f1b26aaa474063fcb059769698a02b

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.