HD-Emo: An Emotional Speech Codec for Preference Extraction

HD-Emo is an emotional speech codec introduced in HPRO. It extracts content and style preference representations from speech tokens, together with hierarchical emotion predictions and ASR transcription.


✨ Features

Given an English speech utterance, HD-Emo extracts:

  • Frame-level speech tokens
  • Frame-level content preference tokens
  • Frame-level style preference tokens
  • Word-level VAD predictions (Valence, Arousal, and Dominance)
  • Sentence-level emotion predictions across 9 classes
  • ASR transcription

HD-Emo uses dual preference extractors with FSQ bottlenecks. The content stream is supervised by ASR, while the style stream is supervised by hierarchical emotional objectives, including sentence-level emotion recognition and word-level VAD prediction.

Note: The released implementation supports feature extraction only. Audio reconstruction from the extracted tokens is not provided.


πŸ™ Acknowledgements

HD-Emo builds upon the following open-source projects:

We sincerely thank their authors and contributors.


πŸ“ Citation

If you find HD-Emo useful in your research, please cite:

@misc{nie2026hpro,
      title={HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech}, 
      author={Sihang Nie and Xiaofen Xing and Rui Xing and Haoming Li and Ruitong Xiao and Jingyuan Xing and Baiji Liu and Xiangmin Xu},
      year={2026},
      eprint={2606.28249},
      archivePrefix={arXiv},
      primaryClass={eess.AS},
      url={https://arxiv.org/abs/2606.28249}, 
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for XXH333/HD-Emo