Humaneness Voice Small
Humaneness Voice Small is an experimental English/German text-to-speech and voice-acting model. This repository contains ten stage-local checkpoints from one continuous S1→S10 training run, the training and inference source, run statistics, and a guide to the exact prompt surfaces. It is not a standard AutoModel.from_pretrained() checkpoint: it uses a Qwen3-0.6B semantic backbone, a MOSS-style local Talker, and the separately loaded MOSS Audio Tokenizer v2. Read the inference instructions before loading weights.
The semantic backbone was initialized from pretrained Qwen3-0.6B, not from random weights. The 112,764,928-parameter Talker and bridge were initialized fresh for this run. This corrects the loose description “trained from scratch” sometimes applied to earlier experiments. The Qwen hidden width is 1,024 and the local Talker width is 2,560. One semantic frame spans 80 ms; the Talker predicts 12 RVQ codebook indices per frame, or 150 codebook indices per second. It does not predict 32 codebooks per frame. The codec reconstructs 48-kHz audio. The full state is approximately 715 million parameters, including the frozen, unused score conditioner; the model-only exports are BF16 PyTorch state dictionaries.
Choose a checkpoint
The checkpoints/S*/model_bf16.pt files are the lowest independent held-out validation-loss checkpoint within each stage. The common validation set has 50 held-out prompts in each of two conditioning modes (100 rows/checkpoint), and all 49 complete candidates were swept. The acting benchmark did not select the checkpoints. A lower validation loss does not necessarily mean a more natural voice.
| Stage | Ladder tier | Approx. measured presentation-hours¹ | Best global step | Held-out loss | Shared-cell acting reward² |
|---|---|---|---|---|---|
| S1 | x100 | 188,354 | 16,265 | 4.8054 | 0.4292 |
| S2 | x50 | 116,838 | 26,284 | 4.7232 | 0.4489 |
| S3 | tier7 | 72,314 | 32,337 | 4.7045 | 0.4628 |
| S4 | tier6 | 57,419 | 37,182 | 4.6966 | 0.4581 |
| S5 | tier5 | 28,716 | 38,183 | 4.6987 | 0.4543 |
| S6 | tier4 | 14,361 | 40,935 | 4.7005 | 0.4539 |
| S7 | tier3 | 5,750 | 41,466 | 4.7056 | 0.4540 |
| S8 | tier2 | 2,877 | 41,739 | 4.7109 | 0.4586 |
| S9 | tier1 | 1,439 | 41,878 | 4.7198 | 0.4540 |
| S10 | tier0 | 578 | 41,936 | 4.7278 | 0.4572 |
¹ These are training presentations, not unique source-audio hours. Multi-format prompts repeat some recordings. The ladder tiers are nested and paired with sidecar material; do not add these hours to estimate unique audio. ² Target-aware automatic reward on 1,651 shared valid acting cells per stage, not a human MOS. The S3–S10 paired difference is +0.00557 with a 95% challenge-bootstrap CI of −0.00701 to +0.01899: the benchmark does not establish S3 as significantly better than S10. S3 is a reasonable provisional default; inspect the same prompts in S3/S10 before choosing for a production voice. S4 has the best held-out validation loss. All underlying values and intervals are in stats/acting_findings.json and stats/acting_analysis.json.gz.
The run used 32 JUPITER Booster nodes × 4 GPUs/node, batch 4,096, one 2,097-step linear warmup followed by one cosine decay across all 41,936 updates—no stage LR restarts. Peak LRs were 8e-5 for the semantic backbone and 2.4e-4 for the Talker, ending at 5% of peak. The source code/plan_caption_production_32nodes.json records the exact historical manifests and hashes; its absolute JUPITER paths are provenance, not portable defaults.
Prompt formats actually trained
The outer model template is <user_inst> with Reference(s), Instruction, Tokens, Quality, Sound Event, Ambient Sound, Language, and Text fields. The examples below show the Instruction field. Text separately repeats the SCRIPT after SCRIPT: for structured prompts, or otherwise follows the inherited packer behavior. Tokens is a codec-frame budget, not a literal transcript token count. Reference audio, when supplied, is codec-encoded through the dedicated reference channel. Do not paste a reference filename into the instruction.
Structured GENERAL/SCRIPT (C01). Use when a specific delivery per sentence, explicit durations, pauses or vocal bursts matter:
GENERAL: frightened, breath-held, intimate, increasingly panicked, adult, low-register
SCRIPT:
(quiet alarm, searching hesitation) [3.7 seconds duration] I locked the back door... didn't I?
(fear rising, clipped self-correction) [4.2 seconds duration] I heard the latch click, but now it is standing open.
(whispered panic, sharp inhale) [5.1 seconds duration] Wait—please don't move; there is someone breathing in the hallway.
The prompting guide and the historical training code distinguish GENERAL: global cues from the SCRIPT: line-by-line performance and quoted text of other formats. Sentence durations and pauses are optional training variations, not promises of exact timing. If a vocal burst is requested in the structured SCRIPT, put it at its intended position in the spoken sequence, outside the literal speech text.
Caption + exact transcript (C02–C10). The description outside quotes is English; the spoken text inside TRANSCRIPT may be English or German. The model saw BUD-E/Body Whisper v1.0 and v1.1, Timbre Whisper, Voice Tagging Whisper and one validated Gemma-4-E4B paraphrase per enriched UID. The six possible Gemma presentation modes were not all generated for each UID.
CAPTION: Use a close, low-register voice that begins in contained fear and tightens into whispered panic.
TRANSCRIPT: "I locked the back door... didn't I? I heard the latch click, but now it is standing open."
For Timbre Whisper and Voice Tagging Whisper, the training loader split comma-separated tags from prose and presented tags only, prose only, or both, each one-third of that family's examples. Tag order was shuffled deterministically. The combined form is literal:
CAPTION: Tags: increasingly_panicked, tense, close-mic, breath-held, low-register. Caption: Use a close, low-register voice that tightens into whispered panic.
TRANSCRIPT: "I locked the back door... didn't I?"
The model also saw transcript only with a distinct reference recording (C11):
TRANSCRIPT: "I locked the back door... didn't I?"
For ordinary intelligible speech, start with a concise CAPTION: and an exact quoted TRANSCRIPT:. For voice identity, supply a short, clean, distinct reference audio. The S3 benchmark's Timbre-prose surface (C05) had mean WER 0.0437 across its mixed reference assignments; S3 Gemma-style C10 had the highest descriptive caption-style reward (0.505), while S10's Voice-Tagging tags+prose C09 had WER 0.0473 and the highest caption-style emotion-band hit there. These are not causal prompt-format rankings: reference assignments and authored captions differ by condition. The structured C01 format is currently much less reliable for words (S3 mean WER ≈0.348; S10 ≈0.313 across reference modes), but it is the surface that explicitly expresses sentence timing and bursts. It has not shown reliable exact burst-onset control.
Evaluation: what the scores do and do not say
The public listening Space and direct site contain 1,872 takes per checkpoint × ten checkpoints = 18,720 takes: 50 English/German acting challenges, 11 prompt surfaces, three paired seeds, plus controlled reference and duration variations. Every stage has a dedicated listening page; replace S3 in that URL with S1–S10. Separate extreme S3 prompting and S10 reference/CFG exploration are exploratory, not the frozen 18,720-take benchmark. The audio/score dataset is here.
| Stage | Mean WER ↓ | Genuineness ↑ | Burst blend ↑ | Requested-emotion hit ↑ | VoiceNet target hit ↑ | Requested burst type+onset F1 ↑ | Sentence duration MAE, s ↓ |
|---|---|---|---|---|---|---|---|
| S1 | 0.152 | 1.368 | 4.042 | 0.192 | 0.323 | 0.017 | 0.870 |
| S2 | 0.128 | 1.533 | 3.981 | 0.199 | 0.338 | 0.053 | 0.850 |
| S3 | 0.121 | 1.636 | 4.470 | 0.193 | 0.329 | 0.011 | 0.884 |
| S4 | 0.120 | 1.615 | 4.354 | 0.191 | 0.333 | 0.030 | 0.934 |
| S5 | 0.113 | 1.571 | 4.139 | 0.198 | 0.336 | 0.022 | 0.949 |
| S6 | 0.122 | 1.609 | 4.577 | 0.201 | 0.326 | 0.031 | 0.924 |
| S7 | 0.127 | 1.588 | 4.606 | 0.195 | 0.323 | 0.035 | 1.022 |
| S8 | 0.115 | 1.591 | 4.593 | 0.192 | 0.326 | 0.015 | 1.007 |
| S9 | 0.124 | 1.565 | 4.561 | 0.196 | 0.324 | 0.022 | 1.024 |
| S10 | 0.121 | 1.561 | 4.556 | 0.207 | 0.320 | 0.030 | 0.966 |
WER comes from Parakeet v3. Genuineness, burst blend, emotion and VoiceNet are learned proxy scores, not human listener ratings. Burst realization and duration error are restricted to eligible prompts; duration MAE excludes failed alignments, so consult alignment coverage before comparing it. A per-stage optimum in one column is not an all-round win. The balanced caption-condition comparison suggests reference audio often reduces descriptive WER, but the paired C01 reference/no-reference WER differences for S3 and S10 have confidence intervals crossing zero. No universal causal reference benefit is claimed. The benchmark's C01 prompts did carry parenthesized sentence directions, but lacked explicit pause tags, requested 0.7-second bursts only at sentence onset, and used durations for only half the challenges. These are limitations, not successful fine-grained-control results.
Practical CFG / temperature recommendations and GPU measurements
A separate S3 pilot measured 660 synthetic takes across CFG, temperature, emotion and prompt recipes, including German, a second-reference sensitivity check and concise structured-cue follow-ups. The study Space hosts both English reports:
- Emotion, CFG and temperature study — full results and listening examples.
- On-device GPU throughput — offline, incremental PCM streaming and parallel batches.
Recommended starting point for the tested S3 demo: temperature 1.0, CFG 1.5, concise CAPTION: and exact quoted TRANSCRIPT:. CFG 2.0 at temperature 1.0 is a measured alternative to audition. Use a short, clean, distinct speech reference for voice identity; do not use the target recording as its own reference. These recommendations are provisional and do not establish an optimum for all voices, prompts or checkpoints.
| Setting / recipe | Recommendation from this pilot |
|---|---|
| Caption + exact transcript, CFG 1.5 / T 1.0 | Conservative demo default. Main English cohort: mean raw WER 0.9% across 10 takes. |
| Caption + exact transcript, CFG 2.0 / T 1.0 | Useful alternative. Equally weighted main/German/second-reference caption pool: raw WER 4.24% vs 4.55% for the default, 30 takes per setting; descriptive, not a proven universal improvement. |
| Caption, CFG 2.0 / T 1.2 | Emotion-first winner under the main-grid selection rule, but less robust across reference checks; not the global default. |
| Concise GENERAL/SCRIPT cues, CFG 3.0 / T 1.0 | Exploratory audition candidate only. Low absolute naturalness proxy and no German/second-reference validation for this follow-up. |
| Temperature 0 | True argmax for audio and stop/continue in the tested demo. Available, but not the recommended study default; poor performance on this fixed recipe does not mean it fails on every prompt. |
CFG here is audio-logit guidance, unconditional + g * (conditional - unconditional). The comparison prompt keeps the same words, reference and requested timing/bursts while neutralizing emotional delivery. CFG 1 means no guidance beyond the conditional branch; reference-free or differently constructed unconditional prompts may behave differently. At positive audio temperatures the stop temperature stays 1.0; the tested T=0 setting also changes stop sampling. Top-k 50, top-p 0.95 and repetition penalty 1.0 were held fixed. The standalone inference example below should not be assumed to implement this optimized demo's CFG branch and stop policy automatically.
Naturalness/genuineness and emotion were measured with learned model proxies, not human MOS. The first-grid emotion advantage over the default had wide intervals crossing zero. The prompt set is small; seed repeats at T=0 are not independent. Reference A is a positive-valence-labelled corpus speech clip; B is a hiss-labelled take containing speech, not a verified neutral control. The second-reference cohort is therefore a stress check, not a clean-reference replication. Verbose parenthesized script directions performed poorly, while shorter cues improved intelligibility; this does not imply that every structured prompt fails. The full report discloses the predeclared rule, follow-ups, reference provenance and all failures. WER uses Unicode-aware word Levenshtein (S+D+I)/N on untrimmed PCM; post-alignment/fade WER is reported separately. Do not directly equate these values with the separate historical acting benchmark above, whose protocol differs.
Throughput was measured on an RTX 3090, using S3 at revision 5de86032771c37a3008d14e78ee7cdb39368e2a8. Warm, fixed-eight-second single-stream RTF was approximately 0.20 offline and 0.30 incremental PCM after optimization. Ten parallel incremental streams achieved approximately 28.3 aggregate audio seconds per GPU second, with per-stream RTF approximately 0.35. Aggregate batch RTF divides GPU time by the sum of stream durations; it is not individual-stream latency. The measured path includes TTS and codec plus launch gaps, but excludes Luna, reference encoding, alignment, scoring, MP3 encoding and network. Browser playback in the tested demo still waits for alignment/scoring. These are workload-specific warm measurements; compilation, long-form memory and batch limits are documented in the report. Fused/compiled BF16 inference is not bit-identical to the original code, and numerical checks are not a human perceptual-equivalence test.
Files and inference
checkpoints/S1throughS10: selected BF16 model-only state plusmetadata.jsonwith SHA-256 and original checkpoint contract. Optimizer/RNG state is not uploaded. Full 8.5-GB-per-checkpoint resume states remain in the original JUPITER scratch run; the public weights alone do not resume the optimizer exactly.code/continuous_train.py,code/train.sbatch,code/plan_caption_production_32nodes.jsonand adjacent modules: the historical JUPITER source and training contract. The raw plan refers to local manifests and intentionally fails its hash checks if edited; regenerate plans and manifests for new data.code/infer.pyis the standalone example for the exported weights. The code bundles MOSS export modules, tokenizer configuration and score schema; the audio codec is fetched separately from its upstream Hugging Face repository.stats/training_steps.jsonl: every logged training-step scalar line across all four Slurm invocations.stats/invocation_summaries.json,stats/run_configuration.json,stats/stage_winners.json,stats/validation_loss_full.tar.gz,stats/acting_findings.json, andstats/acting_analysis.json.gzhold the remaining scalar summaries, all held-out checkpoint-sweep records, and the public acting aggregate.stats/CONTENTS.jsongives hashes.
Example (requires a CUDA environment with PyTorch, Transformers, NumPy, SoundFile, Torchaudio, Safetensors, and permission to download the separate codec):
python code/infer.py --stage S3 --language en --frames 75 \
--prompt $'CAPTION: Warm, softly amused, conversational.\nTRANSCRIPT: "I cannot believe we made it."' \
--text 'I cannot believe we made it.' --output example.wav
Add --reference-wav reference.wav for a distinct, clean recording under about 3 seconds. Do not use the target recording as its own reference. The example is a source-level starting point; the public listening Space is the already-audited playback path. This repository does not contain the large original SFT-3 model weights or original Qwen checkpoint separately, because the exported Humaneness Voice Small states include every trained model weight required for inference.
Training data, license and responsible use
The core nested scaling ladder and associated annotation sidecars are documented at laion/tts-scaling-ladder-de-en; inspect each linked dataset's license/provenance before a derivative training run. The Humaneness Voice DPO collection is a separate future-adaptation resource, not evidence that these base checkpoints were DPO-trained: real-speech DPO, voice-profile DPO, CFG DPO, and S5 pitch DPO. The pitch set's chosen/rejected quality gap is substantial; do not treat it as a clean voice-identity-only preference corpus.
Model weights and original release materials are provided under CC BY 4.0; attribute LAION / Humaneness Voice Small and preserve attribution to the source projects. Qwen3 and MOSS Audio Tokenizer v2 are separate upstream dependencies under their own stated licenses. Respect speakers' rights, consent, provenance and applicable law. Synthetic voice acting can be mistaken for real speech; disclose generated audio and avoid impersonation or deceptive use. This is a research release with known shakiness, prompt sensitivity and imperfect word/timing/burst control, not a safety-certified production voice.