MODUS TTS β€” Fallout 76 Voice Model

Piper TTS voice trained on MODUS dialogue from Fallout 76.


Models

File Epochs Training samples
modus_10000.onnx 10,000 421 voice lines
modus_10000_v2.onnx 10,000 682 voice lines

Both share the same base, settings and hardware. The voice character is near-identical; v2 is noticeably more robust and slurs less, especially on words outside the training vocabulary. Use v2.


Specs

Property Value
Base checkpoint en_US-lessac-high
Quality high
Sample rate 22,050 Hz
Format ONNX
Language English
Batch size 8
Precision 16-bit AMP
GPU NVIDIA RTX A2000 12GB
Training time ~4–5 days per run

Training vocabulary: 1,984 unique words across 11,117 tokens. Median line length 17 words.


Usage

pip install piper-tts
wget https://huggingface.co/petrusilius/modus-tts/resolve/main/modus_10000_v2.onnx
wget https://huggingface.co/petrusilius/modus-tts/resolve/main/modus_10000_v2.onnx.json

Both files must sit in the same folder.

echo "We have you now, General." | \
  piper --model modus_10000_v2.onnx --output_file output.wav

Useful flags:

Flag Effect
--length_scale 1.3 slower
--length_scale 0.8 faster
--sentence_silence 0.5 longer pause between sentences (default 0.2)

Input notes

Plain text only, no SSML.

  • Write numbers as words: forty two, not 42
  • Spell out abbreviations: General, not Gen.
  • Use ... for mid-sentence pauses
  • Keep sentences short for better prosody

Limitations

  • Mispronounces words absent from the training data. This is a vocabulary limit, not an epoch limit β€” more training does not fix it. Keep generated text close to the model's own register, or supply IPA via espeak-ng.
  • Long unpunctuated sentences sound rushed.
  • VITS samples at inference, so output varies slightly between runs on identical input.

Disclaimer

Non-commercial fan project. Fallout 76 and all related assets are property of Bethesda Softworks.

Downloads last month
77
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support