File size: 15,783 Bytes
85ccdea
66b5948
85ccdea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
{
 "generated": "2026-08-29T13:14:06.253Z",
 "engines": {
  "chatterbox-23lang-v1": {
   "min": null,
   "ideal": null,
   "max": 10,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 β€” 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three."
  },
  "chatterbox-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 10,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 β€” 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three."
  },
  "chatterbox-sts-v1": {
   "min": null,
   "ideal": null,
   "max": 10,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "engines/chatterbox/models/s3gen/s3gen.py:131-132 β€” 'if ref_wav.size(1) > 10 * ref_sr: print(\"WARNING: cosydec received ref longer than 10s\")'. Soft: it warns and continues. s3gen is shared by the clone, 23-lang and voice-changer routes, so the limit applies to all three."
  },
  "cosyvoice3-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 30,
   "max_is_hard": true,
   "trailing_silence_ms": null,
   "evidence": "engines/cosyvoice/impl/cosyvoice/cli/frontend.py:97 asserts speech.shape[1] / 16000 <= 30 ('do not support extract speech token for audio longer than 30s'). MEASURED: a 31.49s clip produced -120dB silence while ComfyUI reported the prompt executed successfully."
  },
  "dia2-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "dots-tts-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "dramabox-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 10,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "This graph pins ref_duration = 10 on DramaBoxEngineNode, and that value IS the reference window: vendor/src/inference_server.py:414 calls decode_audio_from_file(voice_ref, device, 0.0, ref_duration), i.e. it decodes only the first ref_duration seconds, and :424 sets target_samples = int(ref_duration * sampling_rate) so a shorter clip is padded up to it. Audio past 10s is never read. NOTE this ceiling is OUR PIN, not an engine hard limit - the node accepts ref_duration 3.0-30.0 (nodes/engines/dramabox_engine_node.py:92-96, tooltip 'Seconds used from the beginning of the voice reference'), so raising the pin would raise the window."
  },
  "echo-tts-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 300,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "engines/adapters/echo_tts_adapter.py:35 declares MAX_REF_SECONDS = 300 (# 5 minutes), and _prepare_reference_audio at :264-267 TRUNCATES silently: max_samples = int(self.MAX_REF_SECONDS * sample_rate); if audio.shape[-1] > max_samples: audio = audio[:, :max_samples]. It cuts, it never refuses, so the ceiling is soft."
  },
  "f5tts-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 12,
   "max_is_hard": false,
   "trailing_silence_ms": 1000,
   "evidence": "engines/f5_tts/infer/README.md:11 (upstream F5-TTS guidance): 'Use reference audio <12s and leave proper silence space (e.g. 1s) at the end. Otherwise there is a risk of truncating in the middle of word, leading to suboptimal generation.' Soft: it degrades and may truncate, it does not refuse."
  },
  "firered2-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "fish2-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "higgs-v2-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "higgs-v3-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 100,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "engines/higgs_audio_v3/native.py:661 defaults max_reference_seconds=100.0 and :674-676 TRUNCATES the reference to the first 100s (wav[:max_samples]) rather than refusing - so the ceiling is soft. NOTE for any cutter: :672 also strips silence edges itself (trim_silence_edges at -42dB), so trailing silence is REMOVED by this engine and must not be relied on the way F5-TTS requires it."
  },
  "indextts2-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 15,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "THE ENGINE ITSELF TRUNCATES AT 15s: engines/index_tts/indextts/infer_v2.py:679 calls self._load_and_cut_audio(spk_audio_prompt, 15, verbose), and _load_and_cut_audio at :516-527 cuts to int(max_audio_length_seconds * sr) with the message 'Audio too long ..., truncating to N samples'. Audio past 15s is therefore DISCARDED, never heard by the model. Separately, engines/adapters/index_tts_adapter.py:503-510 warns about OOM risk above 30s and 60s - that is the adapter's memory advice on the file it loads, not the cloning window, and 15 is the smaller and more useful number."
  },
  "indextts2-emotion-v1": {
   "min": null,
   "ideal": null,
   "max": 15,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "Same engine as indextts2-clone-v1, which truncates at 15s: engines/index_tts/indextts/infer_v2.py:679 calls _load_and_cut_audio(spk_audio_prompt, 15, verbose), cutting to int(15 * sr) at :516-527 ('Audio too long ..., truncating'). The adapter's separate >30s/>60s OOM warnings (index_tts_adapter.py:503-510) are about memory on the loaded file, not the cloning window."
  },
  "kitten-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "kokoro-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "longcat-clone-v1": {
   "min": 3,
   "ideal": [
    3,
    15
   ],
   "max": 15,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "The node's own prompt_audio tooltip (nodes/voice_clone_node.py:91-96, confirmed live via /object_info): 'Reference audio to clone the voice from. 3-15 seconds gives the best results.' Soft because nothing truncates or refuses at 15s. SEPARATELY AND MORE IMPORTANTLY, THE REFERENCE EATS THE OUTPUT LENGTH: voice_clone_node.py:222 reads max_duration = model.config.max_wav_duration, :267 gives the text only `max_duration - prompt_time`, and :277 clamps the total. MEASURED from the downloaded 3.5B-bf16 config.json: max_wav_duration = 60. So reference and generated speech SHARE A 60-SECOND BUDGET - a 15s reference leaves 45s of output, a 30s one leaves 30s. Same shared-budget shape as ZONOS2's max_seqlen, and it makes staying near the 3-15s window worth more than the tooltip implies."
  },
  "moss-tts-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "mossnano-clone-v1": {
   "min": null,
   "ideal": null,
   "max": null,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "No documented limits found in the runtime source (onnx_tts_runtime.py encodes the whole reference through the codec with no truncation or assert - checked for the standard shapes at build 2026-08-16). Absent numbers mean UNKNOWN, never unlimited; the 9.78s standard test reference cloned well at the audition and at build. Upstream demo prompts are ~5-15s."
  },
  "mossnano-preset-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "neutts-air-v1": {
   "min": 3,
   "ideal": [
    3,
    15
   ],
   "max": 20,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "Upstream README recommends 3-15s references. The real bound is ARITHMETIC (the LongCat shape): max_context is 2048 tokens SHARED between the reference codes (~50/s - the 9.78s standard clip encodes to 488 codes, measured), the phonemized text and the generated speech (~50 tokens/s), so a long reference eats the output budget; 20s of reference (~1000 codes) halves what can be spoken. No assert exists - over-length degrades to a truncated render, never a refusal."
  },
  "omnivoice-clone-v1": {
   "min": null,
   "ideal": [
    3,
    10
   ],
   "max": 20,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "site-packages/omnivoice/models/omnivoice.py:802 warns above 20s ('slower generation, higher memory usage, and degraded voice cloning quality. We recommend trimming it to 3-10s'). CAVEAT that makes app-side cutting necessary: its own auto-trim at :779 runs ONLY when ref_text is None, and Parrot ALWAYS supplies a transcript (sidecar or filename), so the engine's built-in safety trim never fires for us."
  },
  "omnivoice-design-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "orpheus-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "piper-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "pocket-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 30,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "pocket_tts SDK 2.1.0, TTSModel.get_state_for_audio_prompt: the truncate option exists to 'truncate long audio prompts to 30 seconds' and 'prevent memory issues with very long inputs' (docstring). The graph pins truncate_prompt true, so a clip over 30s CLONES FROM ITS FIRST 30 SECONDS (the node mirrors the SDK's own Path-branch truncate for tensor inputs and logs when it fires) - soft, never a failure. No minimum is documented; the standard 9.78s test reference cloned well at build."
  },
  "pocket-preset-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "qwen3-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "qwen3-design-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "qwen3-preset-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "sesame-csm-v1": {
   "min": null,
   "ideal": null,
   "max": 120,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "The pack builds the CSM backbone with max_seq_len=2048 (csm_nodes.py llama3_2_1B kwargs) and Mimi frames at 12.5/s, so context + generated output SHARE a ~163s total budget - the reference EATS the output budget, the LongCat shape. 120s leaves ~40s of speech; behaviour when the combined budget is exceeded is UNVERIFIED (no assert found in the vendored generator), so max_is_hard stays false and the number is arithmetic, not an observed refusal. Short conversational refs (5-30s) are the upstream usage pattern, but no engine-side minimum exists - that range is convention, not evidence."
  },
  "silero-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "soprano-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "spark-clone-v1": {
   "min": 3,
   "ideal": [
    3,
    10
   ],
   "max": 10,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "Training-data guidance, not an assert: the official Space discussion (huggingface.co/spaces/Mobvoi/Offical-Spark-TTS discussions/1) says reference clips should be under 10 seconds, ideally 1-3 sentences, because the training data was short clips. The pack itself imposes NO length limit anywhere (AILab_SparkTTS_Core.py and sparktts/models/audio_tokenizer.py checked for MAX constants, slicing and truncation messages - none exist; BiCodecTokenizer.get_ref_clip takes a fixed ref_segment_duration window for the speaker embedding and the semantic tokens use the whole clip; ref_segment_duration is 6 seconds, read from the downloaded model's own config.yaml 2026-08-16 - so the timbre embedding only ever sees 6s regardless of clip length). Over 10s merely degrades - a longer clip is not refused."
  },
  "spark-design-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "step-editx-clone-v1": {
   "min": 3,
   "ideal": [
    3,
    30
   ],
   "max": 30,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "docs/Dev reports/Step_Audio_EditX_Implementation_Plan.md:36 - 'Zero-shot voice cloning (3-30s reference audio)'. HONEST CAVEAT: that is the pack's own DESIGN DOCUMENT, not an assert or a runtime check - no enforcement was found in the shipped engine code, which is why max_is_hard is false. The separate 0.5-30s figures in the same docs belong to the AUDIO EDIT path (input_audio), not to the clone reference, and are deliberately not recorded here."
  },
  "supertonic-tts-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "vibevoice-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "voxcpm2-clone-v1": {
   "min": null,
   "ideal": null,
   "max": 50,
   "max_is_hard": true,
   "trailing_silence_ms": null,
   "evidence": "voxcpm2_nodes.py:22 defines MAX_REFERENCE_AUDIO_SECONDS = 50.0, and :152-159 _validate_reference_audio_duration RAISES ValueError above it - it does NOT truncate. OWNER-HIT IN PRACTICE and captured in ComfyUI's own log (user/comfyui_8188.log): 'ValueError: Reference audio is 58.7s - max allowed is 50s. Trim the audio and try again.' The render is refused outright before any generation, so this is a genuine hard ceiling - the second in this registry after cosyvoice3, and unlike cosyvoice3 it fails LOUDLY rather than returning silence. NO min or ideal is recorded: the pack states no lower bound or sweet spot anywhere, and inventing one would feed the Auto-Bind cutter a guess."
  },
  "voxtream-clone-v1": {
   "min": 1,
   "ideal": [
    5,
    10
   ],
   "max": 20,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "voxtream 0.2.4 generator.json config: max_prompt_sec 20 / min_prompt_sec 1; run.py -pa help: '5-10 sec of target voice. Max 20 sec'. Behaviour past 20s is UNMEASURED (soft declaration); the 9.78s standard test reference sits in the ideal band and cloned well at the audition and at build."
  },
  "zonos-clone-v1": {
   "unknown": true,
   "note": "no researched block β€” absent means UNKNOWN, never unlimited"
  },
  "zonos2-clone-v1": {
   "min": 5,
   "ideal": [
    5,
    30
   ],
   "max": 60,
   "max_is_hard": false,
   "trailing_silence_ms": null,
   "evidence": "runtime.py:31-33 define MAX_REFERENCE_SECONDS = 60.0, RECOMMENDED_REFERENCE_MIN_SECONDS = 5.0 and RECOMMENDED_REFERENCE_MAX_SECONDS = 30.0. runtime.py:176-185 clips anything longer than 60s to the first 60s and emits a logger.warning; runtime.py:186-192 warns (without failing) below 5s. Both limits are advisory in the engine itself - it degrades, it never refuses - so max_is_hard is false."
  },
  "gptsovits": {
   "min": 3,
   "ideal": [
    3,
    10
   ],
   "max": 10,
   "max_is_hard": true,
   "evidence": "api_v2 enforces 3-10s with HTTP 400 (live probe, v25.54.0)"
  },
  "elevenlabs": {
   "na": true,
   "note": "direct voice_id route (ElevenLabs cloud voices) β€” not a reference-audio cloner, no reference-length limit applies"
  }
 }
}