{ "generated": "2026-08-29T13:14:06.253Z", "engines": { "chatterbox-23lang-v1": { "min": null, "ideal": null, "max": 10, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 — 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three." }, "chatterbox-clone-v1": { "min": null, "ideal": null, "max": 10, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 — 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three." }, "chatterbox-sts-v1": { "min": null, "ideal": null, "max": 10, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 — 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three." }, "cosyvoice3-clone-v1": { "min": null, "ideal": null, "max": 30, "max_is_hard": true, "trailing_silence_ms": null, "evidence": "engines/cosyvoice/impl/cosyvoice/cli/frontend.py:97 asserts speech.shape[1] / 16000 <= 30 ('do not support extract speech token for audio longer than 30s'). MEASURED: a 31.49s clip produced -120dB silence while ComfyUI reported the prompt executed successfully." }, "dia2-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "dots-tts-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "dramabox-clone-v1": { "min": null, "ideal": null, "max": 10, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "This graph pins ref_duration = 10 on DramaBoxEngineNode, and that value IS the reference window: vendor/src/inference_server.py:414 calls decode_audio_from_file(voice_ref, device, 0.0, ref_duration), i.e. it decodes only the first ref_duration seconds, and :424 sets target_samples = int(ref_duration * sampling_rate) so a shorter clip is padded up to it. Audio past 10s is never read. NOTE this ceiling is OUR PIN, not an engine hard limit - the node accepts ref_duration 3.0-30.0 (nodes/engines/dramabox_engine_node.py:92-96, tooltip 'Seconds used from the beginning of the voice reference'), so raising the pin would raise the window." }, "echo-tts-clone-v1": { "min": null, "ideal": null, "max": 300, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "engines/adapters/echo_tts_adapter.py:35 declares MAX_REF_SECONDS = 300 (# 5 minutes), and _prepare_reference_audio at :264-267 TRUNCATES silently: max_samples = int(self.MAX_REF_SECONDS * sample_rate); if audio.shape[-1] > max_samples: audio = audio[:, :max_samples]. It cuts, it never refuses, so the ceiling is soft." }, "f5tts-clone-v1": { "min": null, "ideal": null, "max": 12, "max_is_hard": false, "trailing_silence_ms": 1000, "evidence": "engines/f5_tts/infer/README.md:11 (upstream F5-TTS guidance): 'Use reference audio <12s and leave proper silence space (e.g. 1s) at the end. Otherwise there is a risk of truncating in the middle of word, leading to suboptimal generation.' Soft: it degrades and may truncate, it does not refuse." }, "firered2-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "fish2-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "higgs-v2-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "higgs-v3-clone-v1": { "min": null, "ideal": null, "max": 100, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "engines/higgs_audio_v3/native.py:661 defaults max_reference_seconds=100.0 and :674-676 TRUNCATES the reference to the first 100s (wav[:max_samples]) rather than refusing - so the ceiling is soft. NOTE for any cutter: :672 also strips silence edges itself (trim_silence_edges at -42dB), so trailing silence is REMOVED by this engine and must not be relied on the way F5-TTS requires it." }, "indextts2-clone-v1": { "min": null, "ideal": null, "max": 15, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "THE ENGINE ITSELF TRUNCATES AT 15s: engines/index_tts/indextts/infer_v2.py:679 calls self._load_and_cut_audio(spk_audio_prompt, 15, verbose), and _load_and_cut_audio at :516-527 cuts to int(max_audio_length_seconds * sr) with the message 'Audio too long ..., truncating to N samples'. Audio past 15s is therefore DISCARDED, never heard by the model. Separately, engines/adapters/index_tts_adapter.py:503-510 warns about OOM risk above 30s and 60s - that is the adapter's memory advice on the file it loads, not the cloning window, and 15 is the smaller and more useful number." }, "indextts2-emotion-v1": { "min": null, "ideal": null, "max": 15, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "Same engine as indextts2-clone-v1, which truncates at 15s: engines/index_tts/indextts/infer_v2.py:679 calls _load_and_cut_audio(spk_audio_prompt, 15, verbose), cutting to int(15 * sr) at :516-527 ('Audio too long ..., truncating'). The adapter's separate >30s/>60s OOM warnings (index_tts_adapter.py:503-510) are about memory on the loaded file, not the cloning window." }, "kitten-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "kokoro-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "longcat-clone-v1": { "min": 3, "ideal": [ 3, 15 ], "max": 15, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "The node's own prompt_audio tooltip (nodes/voice_clone_node.py:91-96, confirmed live via /object_info): 'Reference audio to clone the voice from. 3-15 seconds gives the best results.' Soft because nothing truncates or refuses at 15s. SEPARATELY AND MORE IMPORTANTLY, THE REFERENCE EATS THE OUTPUT LENGTH: voice_clone_node.py:222 reads max_duration = model.config.max_wav_duration, :267 gives the text only `max_duration - prompt_time`, and :277 clamps the total. MEASURED from the downloaded 3.5B-bf16 config.json: max_wav_duration = 60. So reference and generated speech SHARE A 60-SECOND BUDGET - a 15s reference leaves 45s of output, a 30s one leaves 30s. Same shared-budget shape as ZONOS2's max_seqlen, and it makes staying near the 3-15s window worth more than the tooltip implies." }, "moss-tts-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "mossnano-clone-v1": { "min": null, "ideal": null, "max": null, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "No documented limits found in the runtime source (onnx_tts_runtime.py encodes the whole reference through the codec with no truncation or assert - checked for the standard shapes at build 2026-08-16). Absent numbers mean UNKNOWN, never unlimited; the 9.78s standard test reference cloned well at the audition and at build. Upstream demo prompts are ~5-15s." }, "mossnano-preset-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "neutts-air-v1": { "min": 3, "ideal": [ 3, 15 ], "max": 20, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "Upstream README recommends 3-15s references. The real bound is ARITHMETIC (the LongCat shape): max_context is 2048 tokens SHARED between the reference codes (~50/s - the 9.78s standard clip encodes to 488 codes, measured), the phonemized text and the generated speech (~50 tokens/s), so a long reference eats the output budget; 20s of reference (~1000 codes) halves what can be spoken. No assert exists - over-length degrades to a truncated render, never a refusal." }, "omnivoice-clone-v1": { "min": null, "ideal": [ 3, 10 ], "max": 20, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "site-packages/omnivoice/models/omnivoice.py:802 warns above 20s ('slower generation, higher memory usage, and degraded voice cloning quality. We recommend trimming it to 3-10s'). CAVEAT that makes app-side cutting necessary: its own auto-trim at :779 runs ONLY when ref_text is None, and Parrot ALWAYS supplies a transcript (sidecar or filename), so the engine's built-in safety trim never fires for us." }, "omnivoice-design-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "orpheus-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "piper-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "pocket-clone-v1": { "min": null, "ideal": null, "max": 30, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "pocket_tts SDK 2.1.0, TTSModel.get_state_for_audio_prompt: the truncate option exists to 'truncate long audio prompts to 30 seconds' and 'prevent memory issues with very long inputs' (docstring). The graph pins truncate_prompt true, so a clip over 30s CLONES FROM ITS FIRST 30 SECONDS (the node mirrors the SDK's own Path-branch truncate for tensor inputs and logs when it fires) - soft, never a failure. No minimum is documented; the standard 9.78s test reference cloned well at build." }, "pocket-preset-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "qwen3-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "qwen3-design-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "qwen3-preset-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "sesame-csm-v1": { "min": null, "ideal": null, "max": 120, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "The pack builds the CSM backbone with max_seq_len=2048 (csm_nodes.py llama3_2_1B kwargs) and Mimi frames at 12.5/s, so context + generated output SHARE a ~163s total budget - the reference EATS the output budget, the LongCat shape. 120s leaves ~40s of speech; behaviour when the combined budget is exceeded is UNVERIFIED (no assert found in the vendored generator), so max_is_hard stays false and the number is arithmetic, not an observed refusal. Short conversational refs (5-30s) are the upstream usage pattern, but no engine-side minimum exists - that range is convention, not evidence." }, "silero-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "soprano-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "spark-clone-v1": { "min": 3, "ideal": [ 3, 10 ], "max": 10, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "Training-data guidance, not an assert: the official Space discussion (huggingface.co/spaces/Mobvoi/Offical-Spark-TTS discussions/1) says reference clips should be under 10 seconds, ideally 1-3 sentences, because the training data was short clips. The pack itself imposes NO length limit anywhere (AILab_SparkTTS_Core.py and sparktts/models/audio_tokenizer.py checked for MAX constants, slicing and truncation messages - none exist; BiCodecTokenizer.get_ref_clip takes a fixed ref_segment_duration window for the speaker embedding and the semantic tokens use the whole clip; ref_segment_duration is 6 seconds, read from the downloaded model's own config.yaml 2026-08-16 - so the timbre embedding only ever sees 6s regardless of clip length). Over 10s merely degrades - a longer clip is not refused." }, "spark-design-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "step-editx-clone-v1": { "min": 3, "ideal": [ 3, 30 ], "max": 30, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "docs/Dev reports/Step_Audio_EditX_Implementation_Plan.md:36 - 'Zero-shot voice cloning (3-30s reference audio)'. HONEST CAVEAT: that is the pack's own DESIGN DOCUMENT, not an assert or a runtime check - no enforcement was found in the shipped engine code, which is why max_is_hard is false. The separate 0.5-30s figures in the same docs belong to the AUDIO EDIT path (input_audio), not to the clone reference, and are deliberately not recorded here." }, "supertonic-tts-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "vibevoice-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "voxcpm2-clone-v1": { "min": null, "ideal": null, "max": 50, "max_is_hard": true, "trailing_silence_ms": null, "evidence": "voxcpm2_nodes.py:22 defines MAX_REFERENCE_AUDIO_SECONDS = 50.0, and :152-159 _validate_reference_audio_duration RAISES ValueError above it - it does NOT truncate. OWNER-HIT IN PRACTICE and captured in ComfyUI's own log (user/comfyui_8188.log): 'ValueError: Reference audio is 58.7s - max allowed is 50s. Trim the audio and try again.' The render is refused outright before any generation, so this is a genuine hard ceiling - the second in this registry after cosyvoice3, and unlike cosyvoice3 it fails LOUDLY rather than returning silence. NO min or ideal is recorded: the pack states no lower bound or sweet spot anywhere, and inventing one would feed the Auto-Bind cutter a guess." }, "voxtream-clone-v1": { "min": 1, "ideal": [ 5, 10 ], "max": 20, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "voxtream 0.2.4 generator.json config: max_prompt_sec 20 / min_prompt_sec 1; run.py -pa help: '5-10 sec of target voice. Max 20 sec'. Behaviour past 20s is UNMEASURED (soft declaration); the 9.78s standard test reference sits in the ideal band and cloned well at the audition and at build." }, "zonos-clone-v1": { "unknown": true, "note": "no researched block — absent means UNKNOWN, never unlimited" }, "zonos2-clone-v1": { "min": 5, "ideal": [ 5, 30 ], "max": 60, "max_is_hard": false, "trailing_silence_ms": null, "evidence": "runtime.py:31-33 define MAX_REFERENCE_SECONDS = 60.0, RECOMMENDED_REFERENCE_MIN_SECONDS = 5.0 and RECOMMENDED_REFERENCE_MAX_SECONDS = 30.0. runtime.py:176-185 clips anything longer than 60s to the first 60s and emits a logger.warning; runtime.py:186-192 warns (without failing) below 5s. Both limits are advisory in the engine itself - it degrades, it never refuses - so max_is_hard is false." }, "gptsovits": { "min": 3, "ideal": [ 3, 10 ], "max": 10, "max_is_hard": true, "evidence": "api_v2 enforces 3-10s with HTTP 400 (live probe, v25.54.0)" }, "elevenlabs": { "na": true, "note": "direct voice_id route (ElevenLabs cloud voices) — not a reference-audio cloner, no reference-length limit applies" } } }