File size: 12,637 Bytes
4a5f792
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ef05f3
4a5f792
 
 
9ef05f3
4a5f792
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ef05f3
 
4a5f792
9ef05f3
4a5f792
 
 
 
 
b5f9085
 
 
 
 
 
 
 
 
 
 
 
 
 
4a5f792
 
 
 
 
 
98682c3
4a5f792
 
 
 
 
 
98682c3
4a5f792
 
 
 
 
 
 
 
 
 
b5f9085
4a5f792
 
 
b5f9085
 
4a5f792
b5f9085
 
4a5f792
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5f9085
4a5f792
b5f9085
4a5f792
 
b5f9085
4a5f792
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ef05f3
 
 
4a5f792
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
---
language:
  - ak
  - ar
  - az
  - ca
  - cs
  - cy
  - ee
  - es
  - ff
  - fr
  - ga
  - gn
  - ha
  - he
  - hr
  - ht
  - hu
  - ig
  - ku
  - ln
  - lt
  - lv
  - mi
  - pl
  - pt
  - qu
  - ro
  - sk
  - sl
  - sm
  - sr
  - tk
  - tr
  - uz
  - vi
  - wo
  - yo
license: apache-2.0
tags:
  - token-classification
  - diacritics
  - tone-restoration
  - vocalization
  - on-device
  - edge-ai
  - core-ml
  - onnx
  - webassembly
  - multilingual
pipeline_tag: token-classification
---

# Mark-38M

A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (`ak`), Arabic (`ar`), Azerbaijani (`az`), Catalan (`ca`), Czech (`cs`), Welsh (`cy`), Ewe (`ee`), Spanish (`es`), Pulaar (`ff`), French (`fr`), Irish (`ga`), Guaraní (`gn`), Hausa (`ha`), Hebrew (`he`), Croatian (`hr`), Haitian Creole (`ht`), Hungarian (`hu`), Igbo (`ig`), Kurdish (`ku`), Lingala (`ln`), Lithuanian (`lt`), Latvian (`lv`), Māori (`mi`), Polish (`pl`), Portuguese (`pt`), Quechua (`qu`), Romanian (`ro`), Slovak (`sk`), Slovenian (`sl`), Samoan (`sm`), Serbian (`sr`), Turkmen (`tk`), Turkish (`tr`), Uzbek (`uz`), Vietnamese (`vi`), Wolof (`wo`), and Yorùbá (`yo`).

On the official academic held-out **Yorùbá YAD test set** (3,330 sentences, 142k characters), it achieves a **15.88% Diacritic Error Rate (DER)**, **19.38% Word Error Rate (WER)**, and **5.58% Character Error Rate (CER)** with **0.0139% text corruption** (zero invented or dropped words), against 69.10% DER for the unmarked input.

Across the 37-language joint evaluation suite, it achieves **93.69% macro marked-position accuracy** with a composite score of **0.8419**. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes.

The model is exported into native on-device formats: a **41.77 MB INT8 ONNX graph** for CPU and WebAssembly, and a compiled **Core ML package** for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with **99.64% character parity** (822 of 825 characters over the 74 evaluation probes).

---

## The Task

Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints:

1. **Tone Languages** (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like `ba` can represent `bá` (to accompany/meet), `bà` (to perch/alight), or `ba` (to hide). Missing tones invert negation, tense, and aspect.
2. **Abjads** (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (*ḥarakāt*: *fatḥah*, *ḍammah*, *kasrah*, *sukūn*, *shaddah*) or case nunation (*tanwīn*). Syntax (*i'rab*), voice (active vs passive), and word semantics require contextual disambiguation across the sentence.
3. **Latin Orthographies** (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (`comio` vs `comió`), nominal cases, and distinct phonemes (`c` vs `ç`, `s` vs `ş`, `a` vs `ă`).

### Why Bytes, Not Subwords or Characters

Subword tokenizers (BPE, WordPiece) break down on diacritized text:
* Unmarked input and marked output share base characters but differ in byte sequences.
* Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably.
* Subword vocabularies cannot generalize to rare combining tone stacks (e.g. `M:0300+0301`).

Mark operates directly on the **256 UTF-8 byte values**. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary:
* `KEEP`: Maintain the underlying byte.
* `P:src->dst`: Substitute single character (e.g. `P:a->á`).
* `M:hex`: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (`NFC`).

**Invariant Guarantee**: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words.

---

## Results

### Official Yorùbá YAD Test Benchmark (3,330 Sentences)

Evaluated on the academic gold-standard [Yorùbá YAD test corpus](https://arxiv.org/abs/2004.14811) (Asahiah et al., Orife):

| System | DER (Diacritic Error) | WER (Word Error) | CER (Char Error) | Hallucination / Invention | Inference Latency |
| :--- | :---: | :---: | :---: | :---: | :---: |
| **Mark-38M** | **15.88%** | **19.38%** | **5.58%** | **0.0139%** | **50.07 ms / sent** |
| Unmarked Input (Identity Baseline) | 69.10% | 78.40% | 21.20% | 0.0000% | — |
| Absolute Improvement | *-53.22%* | *-59.02%* | *-15.62%* | *Zero text drift* | *20 sent/sec (Apple Silicon)* |

### 37-Language Joint Accuracy Report

Held-out evaluation report across all 37 languages:

| Code | Language | Marked-Position Accuracy | Code | Language | Marked-Position Accuracy |
| :--- | :--- | :---: | :--- | :--- | :---: |
| `az` | Azerbaijani | **99.65%** | `ee` | Ewe | **97.39%** |
| `fr` | French | **99.47%** | `sl` | Slovenian | **97.39%** |
| `tr` | Turkish | **99.02%** | `sk` | Slovak | **97.21%** |
| `gn` | Guaraní | **88.61%** | `lv` | Latvian | **96.86%** |
| `es` | Spanish | **98.22%** | `ku` | Kurdish (Kurmanji) | **96.51%** |
| `pt` | Portuguese | **98.09%** | `ig` | Igbo | **96.30%** |
| `ga` | Irish | **98.07%** | `hu` | Hungarian | **95.60%** |
| `sr` | Serbian | **98.07%** | `cs` | Czech | **95.47%** |
| `pl` | Polish | **97.89%** | `ht` | Haitian Creole | **95.47%** |
| `ak` | Akan (Twi) | **97.88%** | `ar` | Arabic | **94.86%** |
| `hr` | Croatian | **97.84%** | `ca` | Catalan | **94.85%** |
| `tk` | Turkmen | **97.80%** | `wo` | Wolof | **93.07%** |
| `lt` | Lithuanian | **97.79%** | `ha` | Hausa | **92.18%** |
| `ro` | Romanian | **97.41%** | `sm` | Samoan | **90.98%** |
| `vi` | Vietnamese | **90.98%** | `qu` | Quechua | **89.78%** |
| `he` | Hebrew | **89.40%** | `uz` | Uzbek | **89.15%** |
| `yo` | Yorùbá | **87.39%** | `ff` | Pulaar (Fula) | **85.56%** |
| `mi` | Māori | **81.15%** | `ln` | Lingala | **78.56%** |
| `cy` | Welsh | **74.68%** | | | |

* **Macro marked-position accuracy**: **93.69%**
* **Worst-case language**: Welsh (`cy`) at **74.68%**
* **Composite score**: **0.8419**

---

## Quantization & Edge Footprint

| Format | Precision | File Size | Recommended Target | Latency (CPU / Apple NE) |
| :--- | :--- | ---: | :--- | :---: |
| `mark_int8.onnx` | Dynamic INT8 | **41.77 MB** | Edge CPU, Mobile, Browser (WASM) | 71.49 ms (4-thread CPU) |
| `mark_fp16.onnx` | Float16 | **80.06 MB** | Mobile GPUs, WebGPU | 25.10 ms (GPU) |
| `mark.mlpackage` | 8-bit Core ML | **42.10 MB** | Apple Neural Engine (iOS, macOS) | 12.40 ms (ANE) |
| `mark_fp32.onnx` | Float32 | **158.84 MB** | Reference server baseline | 135.15 ms (1-thread CPU) |

---

## Install & Integration

### Python (ONNX Runtime)

```python
from mark import restore

# Standard inference
restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo")
# -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."

# Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4)
stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4)
# -> "El niño comió jamón en la mañana."
```

### iOS / macOS (Swift Package Manager)

Add to your `Package.swift`:

```swift
dependencies: [
    .package(url: "https://huggingface.co/Mythologic/mark", branch: "main")
]
```

Declare dependency in your target:

```swift
.product(name: "Mark", package: "mark")
```

Direct inference on Apple Neural Engine (Swift 6.3):

```swift
import Mark

let mark = try Mark()

// Yorùbá
let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo")
// -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."

// Arabic
let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar")
// -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ"

// Spanish (with calibrated margin: 0.4 for zero typing jitter)
let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4)
// -> "El niño comió jamón en la mañana."
```

### Browser & Node.js (npm)

Install via npm:

```bash
npm i @mythologic/mark
```

Inference via ONNX Runtime Web (WASM / WebGPU):

```typescript
import { Mark, restore } from "@mythologic/mark";

// Direct helper
const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo");
// -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé.

const arabic = await restore("ذهب الولد الى المدرسة", "ar");
// -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ

// Reusable instance with calibrated anti-flicker margin
const mark = new Mark();
await mark.init();
const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 });
// -> El niño comió jamón en la mañana.
```

---

## Architecture

| Attribute | Specification |
| :--- | :--- |
| **Parameters, total** | **38,735,217** (38.7M) |
| **Encoder layers** | 16 Transformer encoder layers |
| **Hidden dimension** | 384 |
| **Intermediate dimension** | 1,536 (SwiGLU feed-forward) |
| **Attention heads** | 8 heads (head dimension 48) |
| **Positions** | Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form) |
| **Convolution stem** | 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192) |
| **Vocabulary** | 257 (256 raw UTF-8 bytes + language conditioning) |
| **Classifier head** | Linear projection to 1,073 tag classes |
| **Context window** | 512 bytes with sentence/whitespace sliding window |

---

## Training Setup

* **Dataset volume**: ~786M tokens across 37 languages.
* **Loss formulation**: Class-Balanced focal loss (`cb`, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0).
* **Sampling**: Temperature-scaled power law ($\tau = 0.5$) over language shards.
* **Schedule**: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence).
* **Optimizer**: Muon (matrix trunk) + AdamW (embeddings & classifier).
* **Hardware**: Dedicated high-throughput accelerator cluster.

---

## Limitations & Failure Modes

1. **Zero-Context Homographs**: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish `si` [if] vs `sí` [yes], Arabic `qtr` -> `qaṭara` vs `qaṭr`) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency.
2. **Tail Language Shard Volume**: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed **95–99% accuracy**. Low-resource tail languages in the joint mixture (Welsh `cy` at 74.68%, Lingala `ln` at 78.56%) exhibit lower precision due to smaller corpus volume.
3. **Dialectal Writing**: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy.
4. **Window Chunk Boundaries**: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss.

---

## Files

| File | Format | Size | Description |
| :--- | :--- | ---: | :--- |
| `mark_int8.onnx` | ONNX (INT8) | 41.77 MB | Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly |
| `mark_fp16.onnx` | ONNX (FP16) | 80.06 MB | Half-precision graph for GPUs and Neural Engines |
| `mark_fp32.onnx` | ONNX (FP32) | 158.84 MB | Full-precision reference model |
| `mark.mlpackage.zip` | Core ML | 36.16 MB | Compiled Core ML package for Apple Neural Engine |
| `config.json` | JSON | 1 KB | Model architectural hyperparameters |
| `tags.json` | JSON | 40 KB | 1,073 tag operation mappings |
| `languages.json` | JSON | 1 KB | 37 ISO 639-1 language codes |

---

## Author & Citation

Developed by **Ainouche Abderahmane** and **Mythologic**.

Released under the **Apache 2.0** License.

```bibtex
@software{ainouche_mythologic_mark_2026,
  title  = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages},
  author = {Ainouche, Abderahmane and Mythologic},
  year   = {2026},
  url    = {https://huggingface.co/mythologic/mark},
  note   = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages}
}
```

Part of [Mythologic](https://huggingface.co/mythologic).