HD-Emo: An Emotional Speech Codec for Preference Extraction
HD-Emo is an emotional speech codec introduced in HPRO. It extracts content and style preference representations from speech tokens, together with hierarchical emotion predictions and ASR transcription.
- Paper: HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
- GitHub Repository: XXH333/HD-EMO
- Project Page & Demo: HPRO Demo
β¨ Features
Given an English speech utterance, HD-Emo extracts:
- Frame-level speech tokens
- Frame-level content preference tokens
- Frame-level style preference tokens
- Word-level VAD predictions (Valence, Arousal, and Dominance)
- Sentence-level emotion predictions across 9 classes
- ASR transcription
HD-Emo uses dual preference extractors with FSQ bottlenecks. The content stream is supervised by ASR, while the style stream is supervised by hierarchical emotional objectives, including sentence-level emotion recognition and word-level VAD prediction.
Note: The released implementation supports feature extraction only. Audio reconstruction from the extracted tokens is not provided.
π Acknowledgements
HD-Emo builds upon the following open-source projects:
We sincerely thank their authors and contributors.
π Citation
If you find HD-Emo useful in your research, please cite:
@misc{nie2026hpro,
title={HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech},
author={Sihang Nie and Xiaofen Xing and Rui Xing and Haoming Li and Ruitong Xiao and Jingyuan Xing and Baiji Liu and Xiangmin Xu},
year={2026},
eprint={2606.28249},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2606.28249},
}