ESPnet3 f5tts model

Packed model bundle generated from egs3/libritts/f5tts.

Model

  • Repository: NewGame/libritts_f5tts_training
  • Recipe: egs3/libritts/f5tts
  • Corpus: libritts
  • System: f5tts
  • Creator: thanapattrachu
  • Created: 2026-10-08T15:00:52
  • Branch: espnet3/recipe/f5tts_libritts
  • Git: 481290130d6 (clean)
  • Origin: https://github.com/NewGamezzz/espnet.git

Model summary

  • Class: F5TTS
  • Total parameters: 157,945,188
  • Learnable parameters: 157,945,188 (100.0%)
  • Non-trainable parameters: 0
  • Parameter size: 631.78 MB
  • Buffers: 4,194,338 (16.78 MB)
  • Modules: 418 total, 275 leaf
  • DType composition: torch.float32(102.7%), torch.int64(0.0%)

Usage

from espnet3.api.inference import load

model = load("NewGame/libritts_f5tts_training")
output = model(
    "The text to synthesize.",
    "reference.wav",  # the voice to clone: a path, or a (rate, samples) pair
    "The transcript of the reference recording.",
)
samples, sample_rate = output["wav"].array, output["wav"].rate

Packaging

  • Bundle: model_pack
  • Exp dir: ./exp/training
  • Strategy: copy experiment outputs; include extra recipe assets; apply exclude filters

Results

dataset fwhisper_cer fwhisper_cer_delete fwhisper_cer_equal fwhisper_cer_insert fwhisper_cer_replace fwhisper_wer fwhisper_wer_delete fwhisper_wer_equal fwhisper_wer_insert fwhisper_wer_replace spk_similarity utmos
librispeech_pc 0.7644 0.228 94.4667 0.354 0.1429 1.6404 0.0293 17.9494 0.0728 0.1961 0.7058 4.1267

Training config

expand
num_device: 4
num_nodes: 1
task: null
recipe_dir: .
data_dir: ./data
exp_tag: training
exp_dir: ./exp/training
stats_dir: ./exp/training/stats
inference_dir: ./exp/training/inference
create_dataset:
  recipe_dir: .
dataset:
  _target_: espnet3.components.data.data_organizer.DataOrganizer
  _recursive_: false
  recipe_dir: .
  train:
  - data_src_args:
      split: train
      manifest_path: ./data/manifests_filtered/train.tsv
      fs: 24000
  valid:
  - data_src_args:
      split: valid
      manifest_path: ./data/manifests_filtered/valid.tsv
      fs: 24000
  test: null
  preprocessor:
    _target_: espnet2.train.preprocessor.CommonPreprocessor
    token_type: char
    token_list: ./data/tokens/char_tokens.txt
    text_cleaner: tacotron
    g2p_type: null
    _convert_: all
  _convert_: all
model:
  _target_: espnet3.systems.f5tts.f5tts.F5TTS
  token_list: ./data/tokens/char_tokens.txt
  feats_extract_config:
    fs: 24000
    n_fft: 1024
    hop_length: 256
    win_length: 1024
    n_mels: 100
  hidden_size: 768
  depth: 18
  attention_heads: 12
  attention_head_size: 64
  feed_forward_multiplier: 2
  text_embedding_size: 512
  text_mask_padding: false
  rotary_attention_heads: 1
  convolution_layers: 4
  sigma: 0.0
  audio_drop_probability: 0.3
  condition_drop_probability: 0.2
  mask_fraction_range:
  - 0.7
  - 1.0
  ode_solver_method: euler
  _convert_: all
optimizer:
  _target_: torch.optim.AdamW
  lr: 7.5e-05
  betas:
  - 0.9
  - 0.999
  weight_decay: 0.01
  _convert_: all
scheduler:
  _target_: espnet3.components.schedulers.linear_warmup_decay.LinearWarmupDecayLR
  warmup_steps: 20000
  total_steps: 600000
  _convert_: all
scheduler_interval: step
scheduler_monitor: null
optimizers: null
schedulers: null
best_model_criterion:
- - valid/loss
  - 1
  - min
seed: null
init: null
parallel:
  env: local
  n_workers: 1
  allow_overwrite_lock: false
dataloader:
  collate_fn:
    _target_: espnet2.train.collate_fn.CommonCollateFn
    int_pad_value: 0
    float_pad_value: 0.0
    _convert_: all
  train:
    num_shards: 1
    num_workers: 0
    iter_factory:
      _target_: espnet2.iterators.sequence_iter_factory.SequenceIterFactory
      shuffle: true
      collate_fn:
        _target_: espnet2.train.collate_fn.CommonCollateFn
        int_pad_value: 0
        float_pad_value: 0.0
        _convert_: all
      batches:
        type: numel
        batch_size: 1
        min_batch_size: 4
        shape_files:
        - ./exp/training/stats/train/feats_shape
        batch_bins: 960000
      _convert_: all
  valid:
    num_shards: 1
    num_workers: 0
    iter_factory:
      _target_: espnet2.iterators.sequence_iter_factory.SequenceIterFactory
      shuffle: false
      collate_fn:
        _target_: espnet2.train.collate_fn.CommonCollateFn
        int_pad_value: 0
        float_pad_value: 0.0
        _convert_: all
      batches:
        type: numel
        batch_size: 1
        min_batch_size: 4
        shape_files:
        - ./exp/training/stats/valid/feats_shape
        batch_bins: 960000
      _convert_: all
trainer:
  accelerator: auto
  devices: 4
  num_nodes: 1
  strategy: auto
  accumulate_grad_batches: 8
  check_val_every_n_epoch: 1
  gradient_clip_val: 1.0
  log_every_n_steps: 50
  logger:
  - _target_: lightning.pytorch.loggers.TensorBoardLogger
    save_dir: ./exp/training/tensorboard
    name: tb_logger
    _convert_: all
  max_steps: 600000
  max_epochs: 1000
  callbacks:
  - _target_: espnet3.components.callbacks.ema.EMACallback
    decay: 0.9999
    _convert_: all
fit: {}
sample_rate: 24000
n_mel_channels: 100
hop_length: 256
win_length: 1024
n_fft: 1024
remove_long_short:
  min_wav_duration: 1.0
  max_wav_duration: 20.0
  splits:
  - train
  - valid
  manifest_paths:
    train: ./data/manifest/train.tsv
    valid: ./data/manifest/valid.tsv
  save_path: ./data/manifests_filtered
create_token_list:
  manifest_path: ./data/manifests_filtered/train.tsv
  save_path: ./data/tokens
  filename: char_tokens.txt
  token_type: char
  cleaner: tacotron
  g2p: null
  non_linguistic_symbols: null
  bpemodel: null
  space_symbol: <space>
  add_symbol:
  - <blank>:0
  - <unk>:1
  - <sos/eos>:-1
  vocabulary_size: -1
  cutoff: 0
token_list: ./data/tokens/char_tokens.txt

Citing F5-TTS

@inproceedings{chen-etal-2025-f5,
  author={Yushen Chen and Zhikang Niu and Ziyang Ma and Keqi Deng and Chunhui Wang
    and Jian Zhao and Kai Yu and Xie Chen},
  title={{F5-TTS}: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching},
  year={2025},
  booktitle={Proceedings of the 63rd Annual Meeting of the Association for
    Computational Linguistics (Volume 1: Long Papers)},
  pages={6255--6271},
  doi={10.18653/v1/2025.acl-long.313}
}

Citing ESPnet

@inproceedings{watanabe2018espnet,
  author={Shinji Watanabe and Takaaki Hori and Shigeki Karita and Tomoki Hayashi and
    Jiro Nishitoba and Yuya Unno and Nelson Yalta and Jahn Heymann and Matthew Wiesner
    and Nanxin Chen and Adithya Renduchintala and Tsubasa Ochiai},
  title={{ESPnet}: End-to-End Speech Processing Toolkit},
  year={2018},
  booktitle={Proceedings of Interspeech},
  pages={2207--2211},
  doi={10.21437/Interspeech.2018-1456}
}
Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support