169 GB
358 files
Updated about 1 month ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| data | 354 items | ||
| .gitattributes | 2.31 kB xet | b6a9e0dd | |
| README.md | 3.45 kB xet | a85470ad | |
| viVoice-compare-spectrograms.png | 2.08 MB xet | 55387851 | |
| viVoice-total-duration.png | 36.8 kB xet | 0277b567 |
Important Note ⚠️
This dataset is only to be used for research purposes. Access requests must be made via your school, institution, or work email. Requests from common email services will be rejected. We apologize for any inconvenience.
viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis
For a comprehensive description, please visit https://github.com/thinhlpg/viVoice This dataset is licensed under CC-BY-NC-SA-4.0 and is intended for research purposes only.
Key Features and Statistic 📊
- All audio is cleaned from noise and music.
- Clean cuts are made at the beginning and end of sentences to eliminate any unnecessary silences or disruptions, while avoiding cutting in the middle of words.
- Sourced from 186 YouTube channels, with channel IDs included for transparency.
- Number of samples: 887,772
- Total duration: 1,016.97 hours
- Sampling rate: 24 kHz
- Number of splits: 1 (train only)
- Size: 169 GBs
- Gender distribution of speakers: 61.3% ± 3.02% male (manually estimated from a sample of 1,000 with a 95% confidence interval)
- Estimated transcription error rate: 1.8% ± 0.82% (manually estimated from a sample of 1,000 with a 95% confidence interval)
- This metric is for quick reference purposes only; users of this dataset should carefully inspect it to ensure it meets your requirements.
- The error rate only accounts for sentences with mistranscriptions (more or fewer words than expected).
- Other errors, such as missing punctuation or incorrect but phonetically similar words, are not counted.
- Total size
- 169 GB
- Files
- 358
- Last updated
- Sep 9
- Pre-warmed CDN
- US EU US EU