You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

C1Tech OmniVoice Banner

C1Tech/OmniVoice

OmniVoice is an advanced Text-To-Speech (TTS) and voice generation model developed by C1Tech. Designed specifically to bring expressive, highly natural, and context-aware speech synthesis, this model leverages a high-quality custom dataset to achieve rich prosody and fluid conversational delivery across multiple languages.

Key Features

  • Natural Prosody: Fine-tuned to capture natural accentuation, cadence, and intonation patterns for highly realistic speech.
  • Zero-Shot Voice Cloning: Prompt the model with short audio references (.wav) to accurately replicate tone and speaker characteristics.
  • Bilingual Support: Robust performance on Persian and English text, with smooth handling of embedded words and mixed-language phrases.

Samples

Usage

pip install omnivoice
import soundfile as sf
import torch
from omnivoice import OmniVoice

model = OmniVoice.from_pretrained(
    "C1Tech/OmniVoice_Persian_TTS",
    device_map="cuda",
    dtype=torch.float16
)


ref_audio = "REFERENCE.wav"
ref_text = "REFERENCE TEXT"

target_text = "سلام آرش. حالت چطوره؟"

# Timed inference run
with torch.inference_mode():
    audio = model.generate(
        text=target_text,
        ref_audio=ref_audio,
        ref_text=ref_text,
        speed=1.1,
        num_step=50,
        guidance_scale=3,
    )

sf.write(
    "output.wav",
    audio[0],
    24000,
)

Responsible Usage

سلب مسئولیت (Disclaimer)

این مدل صرفاً برای اهداف پژوهشی و آموزشی در حوزه‌ی تولید گفتار منتشر شده است. شرکت C1Tech هیچ‌گونه مسئولیتی را در قبال نحوه‌ی استفاده، سوءاستفاده، بهره‌برداری غیراخلاقی یا غیرمجاز از این مدل توسط کاربران یا اشخاص ثالث نمی‌پذیرد. مسئولیت کامل استفاده از خروجی‌های تولیدشده توسط این مدل، از جمله رعایت قوانین و مقررات جاری، حقوق مالکیت معنوی، حریم خصوصی افراد و اصول اخلاقی، صرفاً بر عهده‌ی کاربر نهایی است. هرگونه استفاده از این مدل برای جعل هویت صوتی افراد بدون رضایت صریح و مستند آن‌ها، تولید اخبار یا محتوای گمراه‌کننده، فریب یا کلاه‌برداری، و یا هر فعالیت غیرقانونی دیگر، به‌شدت ممنوع بوده و خارج از حدود مجوز استفاده از این مدل است. شرکت C1Tech در برابر هرگونه خسارت مستقیم یا غیرمستقیم ناشی از استفاده از این مدل، تحت هیچ شرایطی مسئولیتی نخواهد داشت.

Direct intended uses

The OmniVoice model is intended for research purposes exploring highly realistic audio dialogue generation and advanced speech synthesis.

Out-of-scope uses

Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in any other way that is prohibited by the Apache 2.0 License. Use to generate any text transcript. Furthermore, this release is not intended or licensed for any of the following scenarios:

  • Voice impersonation without explicit, recorded consent – cloning a real individual’s voice for satire, advertising, ransom, social‑engineering, or authentication bypass.
  • Disinformation or impersonation – creating audio presented as genuine recordings of real people or events.
  • Real‑time or low‑latency voice conversion – telephone or video‑conference “live deep‑fake” applications.
  • Unsupported language – the model is primarily trained on Persian and English data; outputs in other languages are unsupported and may be unintelligible or inaccurate.
  • Generation of background ambience, Foley, or music – OmniVoice is speech‑only and will not produce coherent non‑speech audio.

Risks and limitations

While efforts have been made to optimize it through various techniques, it may still produce outputs that are unexpected, biased, or inaccurate. OmniVoice inherits any biases, errors, or omissions produced by its base components.

  • Potential for Deepfakes and Disinformation: High-quality synthetic speech can be misused to create convincing fake audio content for impersonation, fraud, or spreading disinformation. Users must ensure transcripts are reliable, check content accuracy, and avoid using generated content in misleading ways. Users are expected to deploy the models in a lawful manner, in full compliance with all applicable laws and regulations in the relevant jurisdictions. It is best practice to disclose the use of AI when sharing AI-generated content.
  • Language Limitations: Transcripts in languages other than Persian or English may result in unexpected audio outputs.
  • Non-Speech Audio: The model focuses solely on speech synthesis and does not handle background noise, music, or other sound effects.
  • Overlapping Speech: The current model does not explicitly model or generate overlapping speech segments in conversations.

We specialize in cutting-edge Voice & Audio Intelligence—from state-of-the-art Speech-to-Text (STT) and natural Text-to-Speech (TTS) to Voice Verification, Audio Intelligence, and domain-adapted LLMs.

Need higher accuracy, lower latency, or custom-trained voice models for your enterprise?

📬 Contact Sales: info@c1tech.group

🔗 Explore Dashboard: https://ai.c1tech.group/dashboard ```

Downloads last month
5
Safetensors
Model size
0.6B params
Tensor type
I64
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for C1Tech/OmniVoice_Persian_TTS

Finetuned
Qwen/Qwen3-0.6B
Finetuned
k2-fsa/OmniVoice
Finetuned
(61)
this model