logo
multimodal inference · text / image / audio → text + audio
Clone Voice
请自然朗读下面这句话。建议 3 到 6 秒,环境尽量安静。录完后会自动做 1.5 倍速、提取 `CAM++` speaker embedding 与 `Mimi` ref codes,然后立即加入音色列表。
今天阳光很好,欢迎来到我的语音克隆小实验。
点击 start 开始录音
live voice · bidirectional streaming
connecting
init
00:00:00
latest · awaiting first turnspeak naturally — model responds with voice and text
session transcript 0 turns