--- license: apache-2.0 language: - tr pipeline_tag: text-to-speech library_name: coreml base_model: canberkkkkkk/ema-lightning tags: - text-to-speech - tts - turkish - coreml - on-device - ema-lightning --- # Elphi Türkçe — EMA Lightning for Core ML **Elphi Türkçe is the Core ML distribution of EMA Lightning for the Elphi app.** It lets Elphi read Turkish answers aloud on iPhone, iPad and Mac entirely on the device: once downloaded it works offline, and no text or audio is sent anywhere. **We did not train the underlying text-to-speech model.** EMA Lightning was created by **Canberk Aslan**. This repository holds his published weights converted to Core ML for Apple devices — nothing more. The Core ML conversion, packaging and validation were done by **Bosphorus Intelligence LLC**. This distribution is not affiliated with, sponsored by or endorsed by the author of EMA Lightning. ## The original model | | | |---|---| | Model | EMA Lightning by Canberk Aslan | | Hugging Face | [canberkkkkkk/ema-lightning](https://huggingface.co/canberkkkkkk/ema-lightning) | | GitHub | [canberk7/ema-lightning](https://github.com/canberk7/ema-lightning) | | Weights converted | `canberkkkkkk/ema-lightning` at revision `7a6ba1ad216bb2f1da9863f80ac8770a6a807632` (`ema.pt` SHA-256 `95aec03dafbe0e1d69bca774ab597c779464729a14bc99bfcb52480090c7dfe6`, `decoder.pt` SHA-256 `9595819b173f411340f63d11332695121a97f8bf1f6d8b6fef0b21cf99c7ad67`) | | Code followed | `canberk7/ema-lightning` at commit `12797c5f4dbcced8e4b3392b12bba2fe8b9f5ec6` (v1.0.4) | | License | Apache License 2.0, for the weights and the code (see `LICENSE`) | How the model was designed and trained is described in the original model card; this repository makes no claims beyond it. ## Files | File | Bytes | SHA-256 | |---|---:|---| | `EMADecoder.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 171,303 | `a952147fbfe0bd43e878a61eae3d8718e2307dcf8475664e8d75c06ca2883ce5` | | `EMADecoder.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 11,984,304 | `ed8e4a1701ae8b7b2c1298db38ce34d83fbf107f13f82b77281c1a87f6fcbab9` | | `EMADecoder.mlpackage/Manifest.json` | 617 | `bbf2a573a9fc80dddc5d80f03e9a35d33b6d2a291dfa741aa1fd6bd8290b1acd` | | `EMASound.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 275,635 | `b38839114c7b430312881e27d007f955273697ede50739d2e8fcfed90e965b65` | | `EMASound.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 15,319,296 | `3081db93ad88e7ed0fbe3864e66c3e0cbd8723adac88e0eada64a41ebe92b827` | | `EMASound.mlpackage/Manifest.json` | 617 | `e558217deb6ec55e3aaf91f587fc0743e1596c257db76de43797b4f7e951d872` | | `EMAText.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 29,210 | `24ec2f1acc76eb34fa8979e34e2b29d6fb668533794050957e251ac8e2b3b1ee` | | `EMAText.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 4,792,000 | `f2ec3166e629929b159fbe12b06cb445a202a5a16e35e78b3b5c30dbefa60db3` | | `EMAText.mlpackage/Manifest.json` | 617 | `94ca43b45ecdbefc8d250ded9b325132ecb241ce4df42c5d17ee4c67003b755b` | | `voice.json` | 1,443 | `fba9baf368a1176f9b43d9c37c56a8497684c0d0c14c4ef074732523f14300d8` | 32,575,042 bytes in all. - `EMAText.mlpackage` — the text stage: letter ids → letter features and per-letter durations. - `EMASound.mlpackage` — the sound stage: the aligner and the four flow-matching steps → 64-dimensional latents at 25 Hz. - `EMADecoder.mlpackage` — the original decoder: latents → 48 kHz mono audio. - `voice.json` — the alphabet, the constants and the pinned sources. Core ML ML Programs with float32 weights, for iOS, iPadOS and macOS 26 and later; Elphi runs them on the CPU. ## Changes from the original As the Apache License 2.0 asks, the changes made to the original: - The acoustic model was split into a text stage and a sound stage; the decoder is unchanged. The index math the original computes inside its forward pass (word timeline, letter and frame positions, the aligner's window) is computed by the app, exactly as the original computes it. - The stages were traced with PyTorch 2.7.0 and converted with coremltools 9.0 to ML Programs, then packaged deterministically (the same bytes on every build). - The weights' values are unchanged: no retraining, no fine-tuning, no quantization. ## Validation Against the original PyTorch model on the CPU, with identical noise, on 12 reference pieces: identical frame timelines, and audio within 48.0–67.4 dB SNR of the original (median 61.3 dB) — rounding-level differences. On an iPhone 11 (A13): first audio in 60–80 ms, speech made about 20× faster than it plays. ## How Elphi uses it Elphi downloads these files only when the person asks for Elphi Türkçe, checks every file against its pinned size and SHA-256, compiles the models on the device and keeps them there. The download sends no text. Text normalization (numbers, dates, amounts) uses [normalizer-tr](https://github.com/erdemtuna/normalizer-tr) 0.4.0 by Erdem Tuna (Apache License 2.0), which ships inside the app and is not part of this repository. ## Responsible use EMA Lightning makes synthetic speech, and — as the original model card asks — people who hear it should know that. Elphi presents Elphi Türkçe as an AI-generated voice. Do not present it as a real person's voice or use it for impersonation or scams. Numbers, dates and abbreviations go through automatic normalization, which can misread unusual input. ## License and attribution Apache License 2.0 — `LICENSE` is the original's license text; `NOTICE` carries the attribution. Provided "as is", without warranty of any kind. ## Citation of the original model ```bibtex @misc{aslan2026emalightning, title = {EMA Lightning: Tiny, Fast and Accurate Turkish Text to Speech}, author = {Aslan, Canberk}, year = {2026}, howpublished = {\url{https://huggingface.co/canberkkkkkk/ema-lightning}} } ```