Humaneness Voice Large · Acting fine-tune (three-epoch checkpoint)
This repository contains the complete Large-model checkpoint after three full-parameter fine-tuning epochs on the English/German acting-challenges tuning split, plus a separate experimental two-epoch continuation. The starting weights were the earlier MOSS-local-transformer SFT-3 voice-acting model. This is not a LoRA. The original step 432 remains an immutable comparison point. New voice names are an experimental control surface, not a verified cloning guarantee.
The SFT-3 architecture is a MOSS-style semantic model and local Talker with about 4.13 billion trainable parameters in this run, using MOSS Audio Tokenizer v2 for 12 RVQ codes per 80-ms audio frame. The codec is a separate upstream dependency and is not bundled with these model weights.
What was trained
The pinned dataset revision is 2f914a5081f6c0d6eb348b509ea50cf1d56aa7e3. After excluding one anomalous empty-ASR clip, the training split contains 36,754 clips / 300.27 unique audio hours. Every clip is presented three times (110,262 presentations, 900.81 presentation-hours): GENERAL plus inline SCRIPT, GENERAL plus plain SCRIPT, and inline SCRIPT with no GENERAL. The three views are shuffled across epochs. Half of all presentations have a distinct same-speaker 12-second audio reference; half use one of 30 stable pseudonyms. A balanced 10% control-dropout arm removes the caption and inline controls while preserving the literal transcript. Held-out challenge groups are not training targets.
The MOSS Text field contains only spoken words. Acting directions, durations and vocal bursts belong in Instruction; the speaker alias belongs in - Reference(s):\nSpeaker: NAME. When reference-conditioned, the 12-second codec recording is passed through the actual audio-reference channel. For cloning, use a clean 5–15-second reference recording; do not reuse the target audio as its own reference.
Training used full parameters with 128 GH200 GPUs (32 JUPITER nodes), global batch 256, AdamW, BF16 compute with FP32 parameters, 3% linear warmup and one cosine decay to 10% of peak. Peak learning rates were 8e-6 for the semantic backbone and 2.4e-5 for Talker/bridge. Full optimizer/RNG states were saved after each epoch locally; this public repository contains final inference weights and the training implementation, not the much larger optimizer files.
Checkpoint and use
checkpoints/step-00000432/model.safetensors contains the trained model. Accelerate saved it under the base. wrapper prefix and omitted some tied LM heads. The tested loader in infer.py strips base., restores the weights and re-ties the shared heads. Do not point AutoModel.from_pretrained() directly at this safetensors file. Example:
An experimental continuation from the exact step-432 weights is also available at checkpoints/continuation-00000288/model.safetensors. It adds two complete Gemini acting-data passes (288 optimizer steps at global batch 256) with a fresh optimizer, 3% warm-up and cosine decay. Peaks are 8e-6 for the backbone and 2.4e-5 for the Talker. It has an independent run contract and locally retained optimizer/RNG checkpoints. This is not yet the default recommendation: matched reference-vs-alias listening and speaker/WER scoring must establish whether it improves identity without losing intelligibility. To try it, substitute this path in the command below.
python infer.py --checkpoint checkpoints/step-00000432/model.safetensors \
--text 'We finally made it home.' \
--prompt 'GENERAL: relieved and warm, natural close-mic speech.\nSCRIPT:\n(relieved, exhaling) [2.7 seconds duration] We finally made it home.' \
--alias Alice --frames 65 --output example.wav
Pass --reference-wav speaker.wav in place of --alias Alice for real-speaker conditioning. Run on a CUDA GPU with requirements.txt installed and the upstream codec available. The local generation path has been exercised on JUPITER GH200/PyTorch 2.9/BF16. training/ contains the packed-data reader, manifest/prompt compiler, training code and Slurm template. The training scripts use JUPITER-specific absolute paths; adapt them and run a canary for any new system or data revision.
The dataset and starting model contain more provenance. A local 70-prompt A/B listening grid exists, but it is a listening aid rather than a controlled quality benchmark. Formal reference-vs-alias speaker evaluation is ongoing. Alias identity, pronunciation, timing and emotional control can fail. This is a research release; disclose synthetic speech, avoid impersonation and obtain consent for real speaker references.
© LAION contributors. Model weights and newly authored release materials: CC BY 4.0. Attribute the base model and dataset; upstream codec and other components have their own terms.