C1Tech/OmniVoice
OmniVoice is an advanced Text-To-Speech (TTS) and voice generation model developed by C1Tech. Designed specifically to bring expressive, highly natural, and context-aware speech synthesis, this model leverages a high-quality custom dataset to achieve rich prosody and fluid conversational delivery across multiple languages.
Key Features
- Natural Prosody: Fine-tuned to capture natural accentuation, cadence, and intonation patterns for highly realistic speech.
- Zero-Shot Voice Cloning: Prompt the model with short audio references (
.wav) to accurately replicate tone and speaker characteristics. - Bilingual Support: Robust performance on Persian and English text, with smooth handling of embedded words and mixed-language phrases.
Samples
Usage
pip install omnivoice
import soundfile as sf
import torch
from omnivoice import OmniVoice
model = OmniVoice.from_pretrained(
"C1Tech/OmniVoice_Persian_TTS",
device_map="cuda",
dtype=torch.float16
)
ref_audio = "REFERENCE.wav"
ref_text = "REFERENCE TEXT"
target_text = "سلام آرش. حالت چطوره؟"
# Timed inference run
with torch.inference_mode():
audio = model.generate(
text=target_text,
ref_audio=ref_audio,
ref_text=ref_text,
speed=1.1,
num_step=50,
guidance_scale=3,
)
sf.write(
"output.wav",
audio[0],
24000,
)
Responsible Usage
سلب مسئولیت (Disclaimer)
این مدل صرفاً برای اهداف پژوهشی و آموزشی در حوزهی تولید گفتار منتشر شده است. شرکت C1Tech هیچگونه مسئولیتی را در قبال نحوهی استفاده، سوءاستفاده، بهرهبرداری غیراخلاقی یا غیرمجاز از این مدل توسط کاربران یا اشخاص ثالث نمیپذیرد. مسئولیت کامل استفاده از خروجیهای تولیدشده توسط این مدل، از جمله رعایت قوانین و مقررات جاری، حقوق مالکیت معنوی، حریم خصوصی افراد و اصول اخلاقی، صرفاً بر عهدهی کاربر نهایی است. هرگونه استفاده از این مدل برای جعل هویت صوتی افراد بدون رضایت صریح و مستند آنها، تولید اخبار یا محتوای گمراهکننده، فریب یا کلاهبرداری، و یا هر فعالیت غیرقانونی دیگر، بهشدت ممنوع بوده و خارج از حدود مجوز استفاده از این مدل است. شرکت C1Tech در برابر هرگونه خسارت مستقیم یا غیرمستقیم ناشی از استفاده از این مدل، تحت هیچ شرایطی مسئولیتی نخواهد داشت.
Direct intended uses
The OmniVoice model is intended for research purposes exploring highly realistic audio dialogue generation and advanced speech synthesis.
Out-of-scope uses
Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in any other way that is prohibited by the Apache 2.0 License. Use to generate any text transcript. Furthermore, this release is not intended or licensed for any of the following scenarios:
- Voice impersonation without explicit, recorded consent – cloning a real individual’s voice for satire, advertising, ransom, social‑engineering, or authentication bypass.
- Disinformation or impersonation – creating audio presented as genuine recordings of real people or events.
- Real‑time or low‑latency voice conversion – telephone or video‑conference “live deep‑fake” applications.
- Unsupported language – the model is primarily trained on Persian and English data; outputs in other languages are unsupported and may be unintelligible or inaccurate.
- Generation of background ambience, Foley, or music – OmniVoice is speech‑only and will not produce coherent non‑speech audio.
Risks and limitations
While efforts have been made to optimize it through various techniques, it may still produce outputs that are unexpected, biased, or inaccurate. OmniVoice inherits any biases, errors, or omissions produced by its base components.
- Potential for Deepfakes and Disinformation: High-quality synthetic speech can be misused to create convincing fake audio content for impersonation, fraud, or spreading disinformation. Users must ensure transcripts are reliable, check content accuracy, and avoid using generated content in misleading ways. Users are expected to deploy the models in a lawful manner, in full compliance with all applicable laws and regulations in the relevant jurisdictions. It is best practice to disclose the use of AI when sharing AI-generated content.
- Language Limitations: Transcripts in languages other than Persian or English may result in unexpected audio outputs.
- Non-Speech Audio: The model focuses solely on speech synthesis and does not handle background noise, music, or other sound effects.
- Overlapping Speech: The current model does not explicitly model or generate overlapping speech segments in conversations.
We specialize in cutting-edge Voice & Audio Intelligence—from state-of-the-art Speech-to-Text (STT) and natural Text-to-Speech (TTS) to Voice Verification, Audio Intelligence, and domain-adapted LLMs.
Need higher accuracy, lower latency, or custom-trained voice models for your enterprise?
📬 Contact Sales: info@c1tech.group
🔗 Explore Dashboard: https://ai.c1tech.group/dashboard ```
- Downloads last month
- 5