This is a collection of parallel text audio dataset that can be used for training text-to-speech or speech-to-text models.