File size: 12,637 Bytes
4a5f792 9ef05f3 4a5f792 9ef05f3 4a5f792 9ef05f3 4a5f792 9ef05f3 4a5f792 b5f9085 4a5f792 98682c3 4a5f792 98682c3 4a5f792 b5f9085 4a5f792 b5f9085 4a5f792 b5f9085 4a5f792 b5f9085 4a5f792 b5f9085 4a5f792 b5f9085 4a5f792 9ef05f3 4a5f792 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 | ---
language:
- ak
- ar
- az
- ca
- cs
- cy
- ee
- es
- ff
- fr
- ga
- gn
- ha
- he
- hr
- ht
- hu
- ig
- ku
- ln
- lt
- lv
- mi
- pl
- pt
- qu
- ro
- sk
- sl
- sm
- sr
- tk
- tr
- uz
- vi
- wo
- yo
license: apache-2.0
tags:
- token-classification
- diacritics
- tone-restoration
- vocalization
- on-device
- edge-ai
- core-ml
- onnx
- webassembly
- multilingual
pipeline_tag: token-classification
---
# Mark-38M
A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (`ak`), Arabic (`ar`), Azerbaijani (`az`), Catalan (`ca`), Czech (`cs`), Welsh (`cy`), Ewe (`ee`), Spanish (`es`), Pulaar (`ff`), French (`fr`), Irish (`ga`), Guaraní (`gn`), Hausa (`ha`), Hebrew (`he`), Croatian (`hr`), Haitian Creole (`ht`), Hungarian (`hu`), Igbo (`ig`), Kurdish (`ku`), Lingala (`ln`), Lithuanian (`lt`), Latvian (`lv`), Māori (`mi`), Polish (`pl`), Portuguese (`pt`), Quechua (`qu`), Romanian (`ro`), Slovak (`sk`), Slovenian (`sl`), Samoan (`sm`), Serbian (`sr`), Turkmen (`tk`), Turkish (`tr`), Uzbek (`uz`), Vietnamese (`vi`), Wolof (`wo`), and Yorùbá (`yo`).
On the official academic held-out **Yorùbá YAD test set** (3,330 sentences, 142k characters), it achieves a **15.88% Diacritic Error Rate (DER)**, **19.38% Word Error Rate (WER)**, and **5.58% Character Error Rate (CER)** with **0.0139% text corruption** (zero invented or dropped words), against 69.10% DER for the unmarked input.
Across the 37-language joint evaluation suite, it achieves **93.69% macro marked-position accuracy** with a composite score of **0.8419**. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes.
The model is exported into native on-device formats: a **41.77 MB INT8 ONNX graph** for CPU and WebAssembly, and a compiled **Core ML package** for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with **99.64% character parity** (822 of 825 characters over the 74 evaluation probes).
---
## The Task
Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints:
1. **Tone Languages** (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like `ba` can represent `bá` (to accompany/meet), `bà` (to perch/alight), or `ba` (to hide). Missing tones invert negation, tense, and aspect.
2. **Abjads** (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (*ḥarakāt*: *fatḥah*, *ḍammah*, *kasrah*, *sukūn*, *shaddah*) or case nunation (*tanwīn*). Syntax (*i'rab*), voice (active vs passive), and word semantics require contextual disambiguation across the sentence.
3. **Latin Orthographies** (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (`comio` vs `comió`), nominal cases, and distinct phonemes (`c` vs `ç`, `s` vs `ş`, `a` vs `ă`).
### Why Bytes, Not Subwords or Characters
Subword tokenizers (BPE, WordPiece) break down on diacritized text:
* Unmarked input and marked output share base characters but differ in byte sequences.
* Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably.
* Subword vocabularies cannot generalize to rare combining tone stacks (e.g. `M:0300+0301`).
Mark operates directly on the **256 UTF-8 byte values**. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary:
* `KEEP`: Maintain the underlying byte.
* `P:src->dst`: Substitute single character (e.g. `P:a->á`).
* `M:hex`: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (`NFC`).
**Invariant Guarantee**: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words.
---
## Results
### Official Yorùbá YAD Test Benchmark (3,330 Sentences)
Evaluated on the academic gold-standard [Yorùbá YAD test corpus](https://arxiv.org/abs/2004.14811) (Asahiah et al., Orife):
| System | DER (Diacritic Error) | WER (Word Error) | CER (Char Error) | Hallucination / Invention | Inference Latency |
| :--- | :---: | :---: | :---: | :---: | :---: |
| **Mark-38M** | **15.88%** | **19.38%** | **5.58%** | **0.0139%** | **50.07 ms / sent** |
| Unmarked Input (Identity Baseline) | 69.10% | 78.40% | 21.20% | 0.0000% | — |
| Absolute Improvement | *-53.22%* | *-59.02%* | *-15.62%* | *Zero text drift* | *20 sent/sec (Apple Silicon)* |
### 37-Language Joint Accuracy Report
Held-out evaluation report across all 37 languages:
| Code | Language | Marked-Position Accuracy | Code | Language | Marked-Position Accuracy |
| :--- | :--- | :---: | :--- | :--- | :---: |
| `az` | Azerbaijani | **99.65%** | `ee` | Ewe | **97.39%** |
| `fr` | French | **99.47%** | `sl` | Slovenian | **97.39%** |
| `tr` | Turkish | **99.02%** | `sk` | Slovak | **97.21%** |
| `gn` | Guaraní | **88.61%** | `lv` | Latvian | **96.86%** |
| `es` | Spanish | **98.22%** | `ku` | Kurdish (Kurmanji) | **96.51%** |
| `pt` | Portuguese | **98.09%** | `ig` | Igbo | **96.30%** |
| `ga` | Irish | **98.07%** | `hu` | Hungarian | **95.60%** |
| `sr` | Serbian | **98.07%** | `cs` | Czech | **95.47%** |
| `pl` | Polish | **97.89%** | `ht` | Haitian Creole | **95.47%** |
| `ak` | Akan (Twi) | **97.88%** | `ar` | Arabic | **94.86%** |
| `hr` | Croatian | **97.84%** | `ca` | Catalan | **94.85%** |
| `tk` | Turkmen | **97.80%** | `wo` | Wolof | **93.07%** |
| `lt` | Lithuanian | **97.79%** | `ha` | Hausa | **92.18%** |
| `ro` | Romanian | **97.41%** | `sm` | Samoan | **90.98%** |
| `vi` | Vietnamese | **90.98%** | `qu` | Quechua | **89.78%** |
| `he` | Hebrew | **89.40%** | `uz` | Uzbek | **89.15%** |
| `yo` | Yorùbá | **87.39%** | `ff` | Pulaar (Fula) | **85.56%** |
| `mi` | Māori | **81.15%** | `ln` | Lingala | **78.56%** |
| `cy` | Welsh | **74.68%** | | | |
* **Macro marked-position accuracy**: **93.69%**
* **Worst-case language**: Welsh (`cy`) at **74.68%**
* **Composite score**: **0.8419**
---
## Quantization & Edge Footprint
| Format | Precision | File Size | Recommended Target | Latency (CPU / Apple NE) |
| :--- | :--- | ---: | :--- | :---: |
| `mark_int8.onnx` | Dynamic INT8 | **41.77 MB** | Edge CPU, Mobile, Browser (WASM) | 71.49 ms (4-thread CPU) |
| `mark_fp16.onnx` | Float16 | **80.06 MB** | Mobile GPUs, WebGPU | 25.10 ms (GPU) |
| `mark.mlpackage` | 8-bit Core ML | **42.10 MB** | Apple Neural Engine (iOS, macOS) | 12.40 ms (ANE) |
| `mark_fp32.onnx` | Float32 | **158.84 MB** | Reference server baseline | 135.15 ms (1-thread CPU) |
---
## Install & Integration
### Python (ONNX Runtime)
```python
from mark import restore
# Standard inference
restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo")
# -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."
# Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4)
stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4)
# -> "El niño comió jamón en la mañana."
```
### iOS / macOS (Swift Package Manager)
Add to your `Package.swift`:
```swift
dependencies: [
.package(url: "https://huggingface.co/Mythologic/mark", branch: "main")
]
```
Declare dependency in your target:
```swift
.product(name: "Mark", package: "mark")
```
Direct inference on Apple Neural Engine (Swift 6.3):
```swift
import Mark
let mark = try Mark()
// Yorùbá
let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo")
// -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."
// Arabic
let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar")
// -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ"
// Spanish (with calibrated margin: 0.4 for zero typing jitter)
let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4)
// -> "El niño comió jamón en la mañana."
```
### Browser & Node.js (npm)
Install via npm:
```bash
npm i @mythologic/mark
```
Inference via ONNX Runtime Web (WASM / WebGPU):
```typescript
import { Mark, restore } from "@mythologic/mark";
// Direct helper
const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo");
// -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé.
const arabic = await restore("ذهب الولد الى المدرسة", "ar");
// -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ
// Reusable instance with calibrated anti-flicker margin
const mark = new Mark();
await mark.init();
const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 });
// -> El niño comió jamón en la mañana.
```
---
## Architecture
| Attribute | Specification |
| :--- | :--- |
| **Parameters, total** | **38,735,217** (38.7M) |
| **Encoder layers** | 16 Transformer encoder layers |
| **Hidden dimension** | 384 |
| **Intermediate dimension** | 1,536 (SwiGLU feed-forward) |
| **Attention heads** | 8 heads (head dimension 48) |
| **Positions** | Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form) |
| **Convolution stem** | 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192) |
| **Vocabulary** | 257 (256 raw UTF-8 bytes + language conditioning) |
| **Classifier head** | Linear projection to 1,073 tag classes |
| **Context window** | 512 bytes with sentence/whitespace sliding window |
---
## Training Setup
* **Dataset volume**: ~786M tokens across 37 languages.
* **Loss formulation**: Class-Balanced focal loss (`cb`, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0).
* **Sampling**: Temperature-scaled power law ($\tau = 0.5$) over language shards.
* **Schedule**: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence).
* **Optimizer**: Muon (matrix trunk) + AdamW (embeddings & classifier).
* **Hardware**: Dedicated high-throughput accelerator cluster.
---
## Limitations & Failure Modes
1. **Zero-Context Homographs**: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish `si` [if] vs `sí` [yes], Arabic `qtr` -> `qaṭara` vs `qaṭr`) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency.
2. **Tail Language Shard Volume**: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed **95–99% accuracy**. Low-resource tail languages in the joint mixture (Welsh `cy` at 74.68%, Lingala `ln` at 78.56%) exhibit lower precision due to smaller corpus volume.
3. **Dialectal Writing**: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy.
4. **Window Chunk Boundaries**: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss.
---
## Files
| File | Format | Size | Description |
| :--- | :--- | ---: | :--- |
| `mark_int8.onnx` | ONNX (INT8) | 41.77 MB | Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly |
| `mark_fp16.onnx` | ONNX (FP16) | 80.06 MB | Half-precision graph for GPUs and Neural Engines |
| `mark_fp32.onnx` | ONNX (FP32) | 158.84 MB | Full-precision reference model |
| `mark.mlpackage.zip` | Core ML | 36.16 MB | Compiled Core ML package for Apple Neural Engine |
| `config.json` | JSON | 1 KB | Model architectural hyperparameters |
| `tags.json` | JSON | 40 KB | 1,073 tag operation mappings |
| `languages.json` | JSON | 1 KB | 37 ISO 639-1 language codes |
---
## Author & Citation
Developed by **Ainouche Abderahmane** and **Mythologic**.
Released under the **Apache 2.0** License.
```bibtex
@software{ainouche_mythologic_mark_2026,
title = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages},
author = {Ainouche, Abderahmane and Mythologic},
year = {2026},
url = {https://huggingface.co/mythologic/mark},
note = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages}
}
```
Part of [Mythologic](https://huggingface.co/mythologic).
|