LoopTTS Refiner 1.5B

This repository contains the Refiner checkpoint only for LoopTTS (arXiv:2608.28970). Given an input utterance, its transcript, and a natural-language instruction, the Refiner generates corrected speech. The Filter and Judge stages are outside this release.

Checkpoint

File Description
model.pt Refiner 1.5B checkpoint, epoch 31 / step 7190 (6,408,626,894 bytes)
config.yaml Portable summary of the evaluated inference configuration
SHA256SUMS Checkpoint checksum

SHA-256 of model.pt: 995185563a1ae25c2fcc57e0e06b7e196252f8a5745b521902166093ab38580f.

The checkpoint was selected from the project's tested TTS_refiner_v1_1.5B_from_FTEmoVoice_GTGlobal_useRAWprompt run. The March 28, 2026 Refiner test script points to this epoch and step, and the adjacent inference log records the same checkpoint being loaded. It is fine-tuned from EmoVoice 1.5B.

Use

The LoopTTS standalone inference instructions provide one command that accepts an input WAV, transcript, and prosody instruction. The script implements the Refiner input packing and greedy decoding directly. It needs no EmoVoice checkout or patch; command-line paths are relative to the LoopTTS repository.

The runtime needs Qwen2.5-1.5B and CosyVoice-300M-SFT in addition to this checkpoint. These are separate dependencies from their original publishers; they are not included in this repository.

Example instruction: Speak with an angry emotion at a moderate speed and high pitch. Stress the word 'no'.

Intended use and limitations

This model is provided for noncommercial research on controllable speech and prosody correction. It was evaluated with greedy decoding, three CosyVoice code layers, a 22,050 Hz output rate, and raw_audio_position=end. The released input preparation supports WAV utterances up to 25 seconds. The model may change words, speaker characteristics, or intended emotion; review outputs before using them in a study.

model.pt is a PyTorch checkpoint read with torch.load; load it only from this repository or another source you trust. The checkpoint alone is not a self-contained inference package. The standalone script has not yet been run against the released checkpoint, so output equivalence remains unverified.

License and attribution

The Refiner checkpoint is released under CC BY-NC 4.0. The fine-tuning base is EmoVoice, whose authors state that their pretrained models are noncommercial. The LoopTTS Refiner inference code uses CC BY-NC 4.0; portions adapted from upstream EmoVoice retain their MIT attribution. Qwen and CosyVoice dependencies retain their own licenses.

Citation

@inproceedings{song2026loopt,
  title     = {Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction},
  author    = {Song, Zeyang and Liu, Tianchi and Wang, Tianrui and Xu, Chenglin and Guo, Yiwen and Li, Haizhou},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026},
  eprint    = {2608.28970},
  archivePrefix = {arXiv},
  url       = {https://arxiv.org/abs/2608.28970}
}
Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Poookeman/LoopTTS-Refiner-1.5B

Base model

yhaha/EmoVoice
Finetuned
(1)
this model

Paper for Poookeman/LoopTTS-Refiner-1.5B