comfytts-models / reference-limits.json
Prompt-Pirate's picture
Upload folder using huggingface_hub
66b5948 verified
Raw
History Blame Contribute Delete
15.8 kB
{
"generated": "2026-08-29T13:14:06.253Z",
"engines": {
"chatterbox-23lang-v1": {
"min": null,
"ideal": null,
"max": 10,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 β€” 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three."
},
"chatterbox-clone-v1": {
"min": null,
"ideal": null,
"max": 10,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 β€” 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three."
},
"chatterbox-sts-v1": {
"min": null,
"ideal": null,
"max": 10,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 β€” 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three."
},
"cosyvoice3-clone-v1": {
"min": null,
"ideal": null,
"max": 30,
"max_is_hard": true,
"trailing_silence_ms": null,
"evidence": "engines/cosyvoice/impl/cosyvoice/cli/frontend.py:97 asserts speech.shape[1] / 16000 <= 30 ('do not support extract speech token for audio longer than 30s'). MEASURED: a 31.49s clip produced -120dB silence while ComfyUI reported the prompt executed successfully."
},
"dia2-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"dots-tts-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"dramabox-clone-v1": {
"min": null,
"ideal": null,
"max": 10,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "This graph pins ref_duration = 10 on DramaBoxEngineNode, and that value IS the reference window: vendor/src/inference_server.py:414 calls decode_audio_from_file(voice_ref, device, 0.0, ref_duration), i.e. it decodes only the first ref_duration seconds, and :424 sets target_samples = int(ref_duration * sampling_rate) so a shorter clip is padded up to it. Audio past 10s is never read. NOTE this ceiling is OUR PIN, not an engine hard limit - the node accepts ref_duration 3.0-30.0 (nodes/engines/dramabox_engine_node.py:92-96, tooltip 'Seconds used from the beginning of the voice reference'), so raising the pin would raise the window."
},
"echo-tts-clone-v1": {
"min": null,
"ideal": null,
"max": 300,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "engines/adapters/echo_tts_adapter.py:35 declares MAX_REF_SECONDS = 300 (# 5 minutes), and _prepare_reference_audio at :264-267 TRUNCATES silently: max_samples = int(self.MAX_REF_SECONDS * sample_rate); if audio.shape[-1] > max_samples: audio = audio[:, :max_samples]. It cuts, it never refuses, so the ceiling is soft."
},
"f5tts-clone-v1": {
"min": null,
"ideal": null,
"max": 12,
"max_is_hard": false,
"trailing_silence_ms": 1000,
"evidence": "engines/f5_tts/infer/README.md:11 (upstream F5-TTS guidance): 'Use reference audio <12s and leave proper silence space (e.g. 1s) at the end. Otherwise there is a risk of truncating in the middle of word, leading to suboptimal generation.' Soft: it degrades and may truncate, it does not refuse."
},
"firered2-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"fish2-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"higgs-v2-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"higgs-v3-clone-v1": {
"min": null,
"ideal": null,
"max": 100,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "engines/higgs_audio_v3/native.py:661 defaults max_reference_seconds=100.0 and :674-676 TRUNCATES the reference to the first 100s (wav[:max_samples]) rather than refusing - so the ceiling is soft. NOTE for any cutter: :672 also strips silence edges itself (trim_silence_edges at -42dB), so trailing silence is REMOVED by this engine and must not be relied on the way F5-TTS requires it."
},
"indextts2-clone-v1": {
"min": null,
"ideal": null,
"max": 15,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "THE ENGINE ITSELF TRUNCATES AT 15s: engines/index_tts/indextts/infer_v2.py:679 calls self._load_and_cut_audio(spk_audio_prompt, 15, verbose), and _load_and_cut_audio at :516-527 cuts to int(max_audio_length_seconds * sr) with the message 'Audio too long ..., truncating to N samples'. Audio past 15s is therefore DISCARDED, never heard by the model. Separately, engines/adapters/index_tts_adapter.py:503-510 warns about OOM risk above 30s and 60s - that is the adapter's memory advice on the file it loads, not the cloning window, and 15 is the smaller and more useful number."
},
"indextts2-emotion-v1": {
"min": null,
"ideal": null,
"max": 15,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "Same engine as indextts2-clone-v1, which truncates at 15s: engines/index_tts/indextts/infer_v2.py:679 calls _load_and_cut_audio(spk_audio_prompt, 15, verbose), cutting to int(15 * sr) at :516-527 ('Audio too long ..., truncating'). The adapter's separate >30s/>60s OOM warnings (index_tts_adapter.py:503-510) are about memory on the loaded file, not the cloning window."
},
"kitten-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"kokoro-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"longcat-clone-v1": {
"min": 3,
"ideal": [
3,
15
],
"max": 15,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "The node's own prompt_audio tooltip (nodes/voice_clone_node.py:91-96, confirmed live via /object_info): 'Reference audio to clone the voice from. 3-15 seconds gives the best results.' Soft because nothing truncates or refuses at 15s. SEPARATELY AND MORE IMPORTANTLY, THE REFERENCE EATS THE OUTPUT LENGTH: voice_clone_node.py:222 reads max_duration = model.config.max_wav_duration, :267 gives the text only `max_duration - prompt_time`, and :277 clamps the total. MEASURED from the downloaded 3.5B-bf16 config.json: max_wav_duration = 60. So reference and generated speech SHARE A 60-SECOND BUDGET - a 15s reference leaves 45s of output, a 30s one leaves 30s. Same shared-budget shape as ZONOS2's max_seqlen, and it makes staying near the 3-15s window worth more than the tooltip implies."
},
"moss-tts-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"mossnano-clone-v1": {
"min": null,
"ideal": null,
"max": null,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "No documented limits found in the runtime source (onnx_tts_runtime.py encodes the whole reference through the codec with no truncation or assert - checked for the standard shapes at build 2026-08-16). Absent numbers mean UNKNOWN, never unlimited; the 9.78s standard test reference cloned well at the audition and at build. Upstream demo prompts are ~5-15s."
},
"mossnano-preset-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"neutts-air-v1": {
"min": 3,
"ideal": [
3,
15
],
"max": 20,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "Upstream README recommends 3-15s references. The real bound is ARITHMETIC (the LongCat shape): max_context is 2048 tokens SHARED between the reference codes (~50/s - the 9.78s standard clip encodes to 488 codes, measured), the phonemized text and the generated speech (~50 tokens/s), so a long reference eats the output budget; 20s of reference (~1000 codes) halves what can be spoken. No assert exists - over-length degrades to a truncated render, never a refusal."
},
"omnivoice-clone-v1": {
"min": null,
"ideal": [
3,
10
],
"max": 20,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "site-packages/omnivoice/models/omnivoice.py:802 warns above 20s ('slower generation, higher memory usage, and degraded voice cloning quality. We recommend trimming it to 3-10s'). CAVEAT that makes app-side cutting necessary: its own auto-trim at :779 runs ONLY when ref_text is None, and Parrot ALWAYS supplies a transcript (sidecar or filename), so the engine's built-in safety trim never fires for us."
},
"omnivoice-design-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"orpheus-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"piper-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"pocket-clone-v1": {
"min": null,
"ideal": null,
"max": 30,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "pocket_tts SDK 2.1.0, TTSModel.get_state_for_audio_prompt: the truncate option exists to 'truncate long audio prompts to 30 seconds' and 'prevent memory issues with very long inputs' (docstring). The graph pins truncate_prompt true, so a clip over 30s CLONES FROM ITS FIRST 30 SECONDS (the node mirrors the SDK's own Path-branch truncate for tensor inputs and logs when it fires) - soft, never a failure. No minimum is documented; the standard 9.78s test reference cloned well at build."
},
"pocket-preset-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"qwen3-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"qwen3-design-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"qwen3-preset-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"sesame-csm-v1": {
"min": null,
"ideal": null,
"max": 120,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "The pack builds the CSM backbone with max_seq_len=2048 (csm_nodes.py llama3_2_1B kwargs) and Mimi frames at 12.5/s, so context + generated output SHARE a ~163s total budget - the reference EATS the output budget, the LongCat shape. 120s leaves ~40s of speech; behaviour when the combined budget is exceeded is UNVERIFIED (no assert found in the vendored generator), so max_is_hard stays false and the number is arithmetic, not an observed refusal. Short conversational refs (5-30s) are the upstream usage pattern, but no engine-side minimum exists - that range is convention, not evidence."
},
"silero-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"soprano-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"spark-clone-v1": {
"min": 3,
"ideal": [
3,
10
],
"max": 10,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "Training-data guidance, not an assert: the official Space discussion (huggingface.co/spaces/Mobvoi/Offical-Spark-TTS discussions/1) says reference clips should be under 10 seconds, ideally 1-3 sentences, because the training data was short clips. The pack itself imposes NO length limit anywhere (AILab_SparkTTS_Core.py and sparktts/models/audio_tokenizer.py checked for MAX constants, slicing and truncation messages - none exist; BiCodecTokenizer.get_ref_clip takes a fixed ref_segment_duration window for the speaker embedding and the semantic tokens use the whole clip; ref_segment_duration is 6 seconds, read from the downloaded model's own config.yaml 2026-08-16 - so the timbre embedding only ever sees 6s regardless of clip length). Over 10s merely degrades - a longer clip is not refused."
},
"spark-design-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"step-editx-clone-v1": {
"min": 3,
"ideal": [
3,
30
],
"max": 30,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "docs/Dev reports/Step_Audio_EditX_Implementation_Plan.md:36 - 'Zero-shot voice cloning (3-30s reference audio)'. HONEST CAVEAT: that is the pack's own DESIGN DOCUMENT, not an assert or a runtime check - no enforcement was found in the shipped engine code, which is why max_is_hard is false. The separate 0.5-30s figures in the same docs belong to the AUDIO EDIT path (input_audio), not to the clone reference, and are deliberately not recorded here."
},
"supertonic-tts-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"vibevoice-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"voxcpm2-clone-v1": {
"min": null,
"ideal": null,
"max": 50,
"max_is_hard": true,
"trailing_silence_ms": null,
"evidence": "voxcpm2_nodes.py:22 defines MAX_REFERENCE_AUDIO_SECONDS = 50.0, and :152-159 _validate_reference_audio_duration RAISES ValueError above it - it does NOT truncate. OWNER-HIT IN PRACTICE and captured in ComfyUI's own log (user/comfyui_8188.log): 'ValueError: Reference audio is 58.7s - max allowed is 50s. Trim the audio and try again.' The render is refused outright before any generation, so this is a genuine hard ceiling - the second in this registry after cosyvoice3, and unlike cosyvoice3 it fails LOUDLY rather than returning silence. NO min or ideal is recorded: the pack states no lower bound or sweet spot anywhere, and inventing one would feed the Auto-Bind cutter a guess."
},
"voxtream-clone-v1": {
"min": 1,
"ideal": [
5,
10
],
"max": 20,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "voxtream 0.2.4 generator.json config: max_prompt_sec 20 / min_prompt_sec 1; run.py -pa help: '5-10 sec of target voice. Max 20 sec'. Behaviour past 20s is UNMEASURED (soft declaration); the 9.78s standard test reference sits in the ideal band and cloned well at the audition and at build."
},
"zonos-clone-v1": {
"unknown": true,
"note": "no researched block β€” absent means UNKNOWN, never unlimited"
},
"zonos2-clone-v1": {
"min": 5,
"ideal": [
5,
30
],
"max": 60,
"max_is_hard": false,
"trailing_silence_ms": null,
"evidence": "runtime.py:31-33 define MAX_REFERENCE_SECONDS = 60.0, RECOMMENDED_REFERENCE_MIN_SECONDS = 5.0 and RECOMMENDED_REFERENCE_MAX_SECONDS = 30.0. runtime.py:176-185 clips anything longer than 60s to the first 60s and emits a logger.warning; runtime.py:186-192 warns (without failing) below 5s. Both limits are advisory in the engine itself - it degrades, it never refuses - so max_is_hard is false."
},
"gptsovits": {
"min": 3,
"ideal": [
3,
10
],
"max": 10,
"max_is_hard": true,
"evidence": "api_v2 enforces 3-10s with HTTP 400 (live probe, v25.54.0)"
},
"elevenlabs": {
"na": true,
"note": "direct voice_id route (ElevenLabs cloud voices) β€” not a reference-audio cloner, no reference-length limit applies"
}
}
}