elphi / README.md
berkergultepe's picture
Elphi Türkçe v1: EMA Lightning (weights 7a6ba1a, code v1.0.4) as Core ML
ad47c7f verified
|
Raw History Blame Contribute Delete
5.9 kB
---
license: apache-2.0
language:
- tr
pipeline_tag: text-to-speech
library_name: coreml
base_model: canberkkkkkk/ema-lightning
tags:
- text-to-speech
- tts
- turkish
- coreml
- on-device
- ema-lightning
---
# Elphi Türkçe — EMA Lightning for Core ML
**Elphi Türkçe is the Core ML distribution of EMA Lightning for the Elphi app.** It lets Elphi read Turkish
answers aloud on iPhone, iPad and Mac entirely on the device: once downloaded it works offline, and no text or
audio is sent anywhere.
**We did not train the underlying text-to-speech model.** EMA Lightning was created by **Canberk Aslan**. This
repository holds his published weights converted to Core ML for Apple devices — nothing more. The Core ML
conversion, packaging and validation were done by **Bosphorus Intelligence LLC**.
This distribution is not affiliated with, sponsored by or endorsed by the author of EMA Lightning.
## The original model
| | |
|---|---|
| Model | EMA Lightning by Canberk Aslan |
| Hugging Face | [canberkkkkkk/ema-lightning](https://huggingface.co/canberkkkkkk/ema-lightning) |
| GitHub | [canberk7/ema-lightning](https://github.com/canberk7/ema-lightning) |
| Weights converted | `canberkkkkkk/ema-lightning` at revision `7a6ba1ad216bb2f1da9863f80ac8770a6a807632` (`ema.pt` SHA-256 `95aec03dafbe0e1d69bca774ab597c779464729a14bc99bfcb52480090c7dfe6`, `decoder.pt` SHA-256 `9595819b173f411340f63d11332695121a97f8bf1f6d8b6fef0b21cf99c7ad67`) |
| Code followed | `canberk7/ema-lightning` at commit `12797c5f4dbcced8e4b3392b12bba2fe8b9f5ec6` (v1.0.4) |
| License | Apache License 2.0, for the weights and the code (see `LICENSE`) |
How the model was designed and trained is described in the original model card; this repository makes no
claims beyond it.
## Files
| File | Bytes | SHA-256 |
|---|---:|---|
| `EMADecoder.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 171,303 | `a952147fbfe0bd43e878a61eae3d8718e2307dcf8475664e8d75c06ca2883ce5` |
| `EMADecoder.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 11,984,304 | `ed8e4a1701ae8b7b2c1298db38ce34d83fbf107f13f82b77281c1a87f6fcbab9` |
| `EMADecoder.mlpackage/Manifest.json` | 617 | `bbf2a573a9fc80dddc5d80f03e9a35d33b6d2a291dfa741aa1fd6bd8290b1acd` |
| `EMASound.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 275,635 | `b38839114c7b430312881e27d007f955273697ede50739d2e8fcfed90e965b65` |
| `EMASound.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 15,319,296 | `3081db93ad88e7ed0fbe3864e66c3e0cbd8723adac88e0eada64a41ebe92b827` |
| `EMASound.mlpackage/Manifest.json` | 617 | `e558217deb6ec55e3aaf91f587fc0743e1596c257db76de43797b4f7e951d872` |
| `EMAText.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 29,210 | `24ec2f1acc76eb34fa8979e34e2b29d6fb668533794050957e251ac8e2b3b1ee` |
| `EMAText.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 4,792,000 | `f2ec3166e629929b159fbe12b06cb445a202a5a16e35e78b3b5c30dbefa60db3` |
| `EMAText.mlpackage/Manifest.json` | 617 | `94ca43b45ecdbefc8d250ded9b325132ecb241ce4df42c5d17ee4c67003b755b` |
| `voice.json` | 1,443 | `fba9baf368a1176f9b43d9c37c56a8497684c0d0c14c4ef074732523f14300d8` |
32,575,042 bytes in all.
- `EMAText.mlpackage` — the text stage: letter ids → letter features and per-letter durations.
- `EMASound.mlpackage` — the sound stage: the aligner and the four flow-matching steps → 64-dimensional latents at 25 Hz.
- `EMADecoder.mlpackage` — the original decoder: latents → 48 kHz mono audio.
- `voice.json` — the alphabet, the constants and the pinned sources.
Core ML ML Programs with float32 weights, for iOS, iPadOS and macOS 26 and later; Elphi runs them on the CPU.
## Changes from the original
As the Apache License 2.0 asks, the changes made to the original:
- The acoustic model was split into a text stage and a sound stage; the decoder is unchanged. The index math the
original computes inside its forward pass (word timeline, letter and frame positions, the aligner's window) is
computed by the app, exactly as the original computes it.
- The stages were traced with PyTorch 2.7.0 and converted with coremltools 9.0 to ML Programs, then packaged
deterministically (the same bytes on every build).
- The weights' values are unchanged: no retraining, no fine-tuning, no quantization.
## Validation
Against the original PyTorch model on the CPU, with identical noise, on 12 reference pieces: identical frame
timelines, and audio within 48.0–67.4 dB SNR of the original (median 61.3 dB) — rounding-level differences.
On an iPhone 11 (A13): first audio in 60–80 ms, speech made about 20× faster than it plays.
## How Elphi uses it
Elphi downloads these files only when the person asks for Elphi Türkçe, checks every file against its pinned size
and SHA-256, compiles the models on the device and keeps them there. The download sends no text. Text
normalization (numbers, dates, amounts) uses [normalizer-tr](https://github.com/erdemtuna/normalizer-tr) 0.4.0
by Erdem Tuna (Apache License 2.0), which ships inside the app and is not part of this repository.
## Responsible use
EMA Lightning makes synthetic speech, and — as the original model card asks — people who hear it should know
that. Elphi presents Elphi Türkçe as an AI-generated voice. Do not present it as a real person's voice or use it
for impersonation or scams. Numbers, dates and abbreviations go through automatic normalization, which can
misread unusual input.
## License and attribution
Apache License 2.0 — `LICENSE` is the original's license text; `NOTICE` carries the attribution. Provided "as
is", without warranty of any kind.
## Citation of the original model
```bibtex
@misc{aslan2026emalightning,
title = {EMA Lightning: Tiny, Fast and Accurate Turkish Text to Speech},
author = {Aslan, Canberk},
year = {2026},
howpublished = {\url{https://huggingface.co/canberkkkkkk/ema-lightning}}
}
```