LoopTTS Refiner 1.5B
This repository contains the Refiner checkpoint only for LoopTTS (arXiv:2608.28970). Given an input utterance, its transcript, and a natural-language instruction, the Refiner generates corrected speech. The Filter and Judge stages are outside this release.
Checkpoint
| File | Description |
|---|---|
model.pt |
Refiner 1.5B checkpoint, epoch 31 / step 7190 (6,408,626,894 bytes) |
config.yaml |
Portable summary of the evaluated inference configuration |
SHA256SUMS |
Checkpoint checksum |
SHA-256 of model.pt: 995185563a1ae25c2fcc57e0e06b7e196252f8a5745b521902166093ab38580f.
The checkpoint was selected from the project's tested TTS_refiner_v1_1.5B_from_FTEmoVoice_GTGlobal_useRAWprompt run. The March 28, 2026 Refiner test script points to this epoch and step, and the adjacent inference log records the same checkpoint being loaded. It is fine-tuned from EmoVoice 1.5B.
Use
The LoopTTS standalone inference instructions provide one command that accepts an input WAV, transcript, and prosody instruction. The script implements the Refiner input packing and greedy decoding directly. It needs no EmoVoice checkout or patch; command-line paths are relative to the LoopTTS repository.
The runtime needs Qwen2.5-1.5B and CosyVoice-300M-SFT in addition to this checkpoint. These are separate dependencies from their original publishers; they are not included in this repository.
Example instruction: Speak with an angry emotion at a moderate speed and high pitch. Stress the word 'no'.
Intended use and limitations
This model is provided for noncommercial research on controllable speech and prosody correction. It was evaluated with greedy decoding, three CosyVoice code layers, a 22,050 Hz output rate, and raw_audio_position=end. The released input preparation supports WAV utterances up to 25 seconds. The model may change words, speaker characteristics, or intended emotion; review outputs before using them in a study.
model.pt is a PyTorch checkpoint read with torch.load; load it only from this repository or another source you trust. The checkpoint alone is not a self-contained inference package. The standalone script has not yet been run against the released checkpoint, so output equivalence remains unverified.
License and attribution
The Refiner checkpoint is released under CC BY-NC 4.0. The fine-tuning base is EmoVoice, whose authors state that their pretrained models are noncommercial. The LoopTTS Refiner inference code uses CC BY-NC 4.0; portions adapted from upstream EmoVoice retain their MIT attribution. Qwen and CosyVoice dependencies retain their own licenses.
Citation
@inproceedings{song2026loopt,
title = {Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction},
author = {Song, Zeyang and Liu, Tianchi and Wang, Tianrui and Xu, Chenglin and Guo, Yiwen and Li, Haizhou},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026},
eprint = {2608.28970},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.28970}
}
- Downloads last month
- 21
Model tree for Poookeman/LoopTTS-Refiner-1.5B
Base model
yhaha/EmoVoice