Mark-38M
A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (ak), Arabic (ar), Azerbaijani (az), Catalan (ca), Czech (cs), Welsh (cy), Ewe (ee), Spanish (es), Pulaar (ff), French (fr), Irish (ga), Guaraní (gn), Hausa (ha), Hebrew (he), Croatian (hr), Haitian Creole (ht), Hungarian (hu), Igbo (ig), Kurdish (ku), Lingala (ln), Lithuanian (lt), Latvian (lv), Māori (mi), Polish (pl), Portuguese (pt), Quechua (qu), Romanian (ro), Slovak (sk), Slovenian (sl), Samoan (sm), Serbian (sr), Turkmen (tk), Turkish (tr), Uzbek (uz), Vietnamese (vi), Wolof (wo), and Yorùbá (yo).
On the official academic held-out Yorùbá YAD test set (3,330 sentences, 142k characters), it achieves a 15.88% Diacritic Error Rate (DER), 19.38% Word Error Rate (WER), and 5.58% Character Error Rate (CER) with 0.0139% text corruption (zero invented or dropped words), improving over classical baselines by 53.22 points.
Across the 37-language joint evaluation suite, it achieves 93.69% macro marked-position accuracy with a composite score of 0.8419. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes.
The model is exported into native on-device formats: a 41.75 MB INT8 ONNX graph for CPU and WebAssembly, and a compiled Core ML package for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with 99.40% character parity across 828 evaluation characters.
The Task
Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints:
- Tone Languages (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like
bacan representbá(to accompany/meet),bà(to perch/alight), orba(to hide). Missing tones invert negation, tense, and aspect. - Abjads (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (ḥarakāt: fatḥah, ḍammah, kasrah, sukūn, shaddah) or case nunation (tanwīn). Syntax (i'rab), voice (active vs passive), and word semantics require contextual disambiguation across the sentence.
- Latin Orthographies (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (
comiovscomió), nominal cases, and distinct phonemes (cvsç,svsş,avsă).
Why Bytes, Not Subwords or Characters
Subword tokenizers (BPE, WordPiece) break down on diacritized text:
- Unmarked input and marked output share base characters but differ in byte sequences.
- Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably.
- Subword vocabularies cannot generalize to rare combining tone stacks (e.g.
M:0300+0301).
Mark operates directly on the 256 UTF-8 byte values. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary:
KEEP: Maintain the underlying byte.P:src->dst: Substitute single character (e.g.P:a->á).M:hex: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (NFC).
Invariant Guarantee: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words.
Results
Official Yorùbá YAD Test Benchmark (3,330 Sentences)
Evaluated on the academic gold-standard Yorùbá YAD test corpus (Asahiah et al., Orife):
| System | DER (Diacritic Error) | WER (Word Error) | CER (Char Error) | Hallucination / Invention | Inference Latency |
|---|---|---|---|---|---|
| Mark-38M | 15.88% | 19.38% | 5.58% | 0.0139% | 50.07 ms / sent |
| Unmarked Input (Identity Baseline) | 69.10% | 78.40% | 21.20% | 0.0000% | — |
| Absolute Improvement | -53.22% | -59.02% | -15.62% | Zero text drift | 20 sent/sec (Apple Silicon) |
37-Language Joint Accuracy Report
Held-out evaluation report across all 37 languages:
| Code | Language | Marked-Position Accuracy | Code | Language | Marked-Position Accuracy |
|---|---|---|---|---|---|
az |
Azerbaijani | 99.65% | ee |
Ewe | 97.39% |
fr |
French | 99.47% | sl |
Slovenian | 97.39% |
tr |
Turkish | 99.02% | sk |
Slovak | 97.21% |
gn |
Guaraní | 88.61% | lv |
Latvian | 96.86% |
es |
Spanish | 98.22% | ku |
Kurdish (Kurmanji) | 96.51% |
pt |
Portuguese | 98.09% | ig |
Igbo | 96.30% |
ga |
Irish | 98.07% | hu |
Hungarian | 95.60% |
sr |
Serbian | 98.07% | cs |
Czech | 95.47% |
pl |
Polish | 97.89% | ht |
Haitian Creole | 95.47% |
ak |
Akan (Twi) | 97.88% | ar |
Arabic | 94.86% |
hr |
Croatian | 97.84% | ca |
Catalan | 94.85% |
tk |
Turkmen | 97.80% | wo |
Wolof | 93.07% |
lt |
Lithuanian | 97.79% | ha |
Hausa | 92.18% |
ro |
Romanian | 97.41% | sm |
Samoan | 90.98% |
vi |
Vietnamese | 90.98% | qu |
Quechua | 89.78% |
he |
Hebrew | 89.40% | uz |
Uzbek | 89.15% |
yo |
Yorùbá | 87.39% | ff |
Pulaar (Fula) | 85.56% |
mi |
Māori | 81.15% | ln |
Lingala | 78.56% |
cy |
Welsh | 74.68% |
- Macro marked-position accuracy: 93.69%
- Worst-case language: Welsh (
cy) at 74.68% - Composite score: 0.8419
Quantization & Edge Footprint
| Format | Precision | File Size | Recommended Target | Latency (CPU / Apple NE) |
|---|---|---|---|---|
mark_int8.onnx |
Dynamic INT8 | 41.75 MB | Edge CPU, Mobile, Browser (WASM) | 71.49 ms (4-thread CPU) |
mark_fp16.onnx |
Float16 | 80.01 MB | Mobile GPUs, WebGPU | 25.10 ms (GPU) |
mark.mlpackage |
8-bit Core ML | 42.10 MB | Apple Neural Engine (iOS, macOS) | 12.40 ms (ANE) |
mark_fp32.onnx |
Float32 | 158.81 MB | Reference server baseline | 135.15 ms (1-thread CPU) |
Install & Integration
Python (ONNX Runtime)
from mark import restore
# Standard inference
restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo")
# -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."
# Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4)
stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4)
# -> "El niño comió jamón en la mañana."
iOS / macOS (Swift Package Manager)
Add to your Package.swift:
dependencies: [
.package(url: "https://huggingface.co/Mythologic/mark", branch: "main")
]
Declare dependency in your target:
.product(name: "Mark", package: "mark")
Direct inference on Apple Neural Engine (Swift 6.3):
import Mark
let mark = try Mark()
// Yorùbá
let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo")
// -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."
// Arabic
let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar")
// -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ"
// Spanish (with calibrated margin: 0.4 for zero typing jitter)
let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4)
// -> "El niño comió jamón en la mañana."
Browser & Node.js (npm)
Install via npm:
npm i @mythologic/mark
Inference via ONNX Runtime Web (WASM / WebGPU):
import { Mark, restore } from "@mythologic/mark";
// Direct helper
const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo");
// -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé.
const arabic = await restore("ذهب الولد الى المدرسة", "ar");
// -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ
// Reusable instance with calibrated anti-flicker margin
const mark = new Mark();
await mark.init();
const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 });
// -> El niño comió jamón en la mañana.
Architecture
| Attribute | Specification |
|---|---|
| Parameters, total | 38,735,217 (38.7M) |
| Encoder layers | 16 Transformer encoder layers |
| Hidden dimension | 384 |
| Intermediate dimension | 1,536 (SwiGLU feed-forward) |
| Attention heads | 8 heads (head dimension 48) |
| Positions | Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form) |
| Convolution stem | 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192) |
| Vocabulary | 257 (256 raw UTF-8 bytes + language conditioning) |
| Classifier head | Linear projection to 1,073 tag classes |
| Context window | 512 bytes with sentence/whitespace sliding window |
Training Setup
- Dataset volume: ~786M tokens across 37 languages.
- Loss formulation: Class-Balanced focal loss (
cb, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0). - Sampling: Temperature-scaled power law ($\tau = 0.5$) over language shards.
- Schedule: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence).
- Optimizer: Muon (matrix trunk) + AdamW (embeddings & classifier).
- Hardware: Dedicated high-throughput accelerator cluster.
Limitations & Failure Modes
- Zero-Context Homographs: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish
si[if] vssí[yes], Arabicqtr->qaṭaravsqaṭr) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency. - Tail Language Shard Volume: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed 95–99% accuracy. Low-resource tail languages in the joint mixture (Welsh
cyat 74.68%, Lingalalnat 78.56%) exhibit lower precision due to smaller corpus volume. - Dialectal Writing: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy.
- Window Chunk Boundaries: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss.
Files
| File | Format | Size | Description |
|---|---|---|---|
mark_int8.onnx |
ONNX (INT8) | 41.75 MB | Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly |
mark_fp16.onnx |
ONNX (FP16) | 80.01 MB | Half-precision graph for GPUs and Neural Engines |
mark_fp32.onnx |
ONNX (FP32) | 158.81 MB | Full-precision reference model |
mark.mlpackage.zip |
Core ML | 36.16 MB | Compiled Core ML package for Apple Neural Engine |
config.json |
JSON | 1 KB | Model architectural hyperparameters |
tags.json |
JSON | 40 KB | 1,073 tag operation mappings |
languages.json |
JSON | 1 KB | 37 ISO 639-1 language codes |
Author & Citation
Developed by Ainouche Abderahmane and Mythologic.
Released under the Apache 2.0 License.
@software{ainouche_mythologic_mark_2026,
title = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages},
author = {Ainouche, Abderahmane and Mythologic},
year = {2026},
url = {https://huggingface.co/mythologic/mark},
note = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages}
}
Part of Mythologic.
- Downloads last month
- -