diacnet-2.0

diacnet-2.0 Highlights

diacnet-2.0 restores diacritics, tone marks and Arabic vowel marks in 11 languages from a single 582M byte-level model, with the following key features:

  • The most accurate Olaverse diacritizer on 8 of 10 languages. With output alignment (below), its error rate is lower than diactag-2.0 on Yorùbá, Hausa, Polish, Turkish, Portuguese, Spanish, French and Italian.
  • Yorùbá, rebuilt: error rate 0.174 → 0.070 against diacnet-1.1 (2.5× lower), and 12% of sentences exactly right (1.1: 0.8%).
  • Now with Arabic: full diacritization (<ara>, DER 0.095 on Classical Arabic) and a no-case-endings mode (<ara-nocase>, DER 0.072) from the same model.
  • Never changes your text: the alignment helper keeps every input letter and takes only the marks, so output compliance is 100% by construction. Raw output additionally fixes typos, if you want that.
  • Language optional: with <auto>, the error rate stays within 0.0005 of the explicit tag.
  • Meaning hints: add [g: word=meaning] for words whose marks depend on meaning (Vietnamese ambiguous-word accuracy 0.63 → 0.90).
  • Partial input: marks already in the text are kept (99.9–100%) and the rest are filled in.

Model Overview

diacnet-2.0 has the following features:

  • Type: byte-level encoder-decoder (text to text), fine-tuned from google/byt5-base
  • Architecture: ByT5-base: 18 encoder + 6 decoder layers, d_model 1536, 12 heads
  • Parameters: 582M; weights 2.3 GB (fp32 safetensors)
  • Languages and tags: <yor> Yorùbá, <ibo> Igbo, <hau> Hausa, <vie> Vietnamese, <pol> Polish, <tur> Turkish, <por> Portuguese, <spa> Spanish, <fra> French, <ita> Italian, <ara> / <ara-nocase> Arabic, <auto>
  • Input: <tag> [g: word=meaning] text, up to about 300 characters per call (the helper below splits longer text)
  • Output: the text with its diacritics restored

Quickstart

import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

repo = "olaverse/diacnet-2.0"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, dtype=torch.bfloat16).eval()   # float32 on CPU

import difflib
import unicodedata as ud

_FOLD = str.maketrans("ɓɗƙƴđıłƁƊƘƳĐŁ", "bdkydilBDKYDL")   # letters with no combining form
_LETTER = {"\u0653", "\u0654", "\u0655"}                  # Arabic madda / hamza are spelling, not marks

def _units(text):
    units = []                                  # [letter, letter + marks]
    for c in ud.normalize("NFD", text):
        if units and ud.combining(c):
            if c in _LETTER:
                units[-1][0] += c
            units[-1][1] += c
        else:
            units.append([c, c])
    return [(ud.normalize("NFC", b).translate(_FOLD), ud.normalize("NFC", f)) for b, f in units]

def align(source, output):
    """`source` with the marks diacnet put on every letter it kept; letters it changed,
    dropped or added fall back to `source`, so your text itself never changes."""
    src, out = _units(source), _units(output)
    res = [f for _, f in src]
    sm = difflib.SequenceMatcher(None, [b for b, _ in src], [b for b, _ in out], autojunk=False)
    for a, b, n in sm.get_matching_blocks():
        res[a:a + n] = [f for _, f in out[b:b + n]]
    return "".join(res)

def chunks(text, n=300):
    out, cur = [], ""
    for w in text.split(" "):
        if cur and len(cur) + 1 + len(w) > n:
            out.append(cur); cur = w
        else:
            cur = f"{cur} {w}" if cur else w
    return out + [cur]

@torch.no_grad()
def diacritize(text, tag="auto", hint="", aligned=True):
    parts = []
    for c in chunks(text):
        x = f"<{tag}> {hint + ' ' if hint else ''}{c}"
        enc = tok(x, return_tensors="pt").to(model.device)
        out = model.generate(**enc, max_new_tokens=2 * enc["input_ids"].shape[1] + 16, num_beams=1)
        y = tok.decode(out[0], skip_special_tokens=True).strip()
        parts.append(align(c, y) if aligned else y)
    return " ".join(parts)

diacritize("O so fun ara re pe oun ko ni isoro kankan.", "yor")
# 'Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.'
diacritize("Dimkpa abuo a ga-ezute onwe ha ozo na mba Saudi Arabia.", "ibo")
# 'Dimkpa abụọ a ga-ezute onwe ha ọzọ na mba Saudi Arabia.'
diacritize("Cutar kan dauki tsawon kwana 14 zuwa 21 kafin ya warke.", "hau")
# 'Cutar kan ɗauki tsawon kwana 14 zuwa 21 kafin ya warke.'
diacritize("Karl bat dau lam bai tap ve nha cua minh.", "vie")
# 'Karl bắt đầu làm bài tập về nhà của mình.'
diacritize("Facebook zawiesil jedno z moich szesciu kont.", "pol")
# 'Facebook zawiesił jedno z moich sześciu kont.'
diacritize("No pude sujetarme a la cuerda mas tiempo.")          # <auto>
# 'No pude sujetarme a la cuerda más tiempo.'
diacritize("وهذا قول مرغوب عنه .", "ara")
# 'وَهَذَا قَوْلٌ مَرْغُوبٌ عَنْهُ .'

For many sentences, batch inputs of similar length (padding="longest") and run on a GPU in bf16.

With the olaverse library

The olaverse library (v0.4.0+) wraps all of the above (chunking, batching, alignment, hints) in one call:

pip install "olaverse[deeplearning]"
from olaverse.nlp import Diacritizer

d = Diacritizer(model="diacnet-2.0", lang="yor", device="cuda")   # "cpu", "mps" or "auto"; bf16 on CUDA
d.restore("O so fun ara re pe oun ko ni isoro kankan.")
# 'Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.'

d.restore("وهذا قول مرغوب عنه .", lang="ara", case_endings=False)
# 'وَهَذَا قَوْل مَرْغُوب عَنْه .'

d.restore("Chi ay chi that su ranh vao nhung buoi toi sau khi da cho con ngu say.",
          lang="vie", hints={"ranh": "free (time)"})
# 'Chị ấy chỉ thật sự rảnh vào những buổi tối sau khi đã cho con ngủ say.'

Diacritizer(model="diacnet-2.0").restore("No pude sujetarme a la cuerda mas tiempo.")   # no lang -> <auto>
# 'No pude sujetarme a la cuerda más tiempo.'

texts = ["O so fun ara re pe oun ko ni isoro kankan.", "وهذا قول مرغوب عنه .",
         "No pude sujetarme a la cuerda mas tiempo."]
d.restore_batch(texts, lang=["yor", "ara", "auto"])   # one language per text, or one string for all
# ['Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.', 'وَهَذَا قَوْلٌ مَرْغُوبٌ عَنْهُ .',
#  'No pude sujetarme a la cuerda más tiempo.']

Output is aligned by default (aligned=False returns the raw output). hints also accepts a list of "word=meaning" strings or a ready "[g: ...]" block.

Tags and Modes

Language tags and <auto>

Give the language tag when you know it. <auto> lets the model infer the language from the text; it costs at most 0.0005 DER on diacbench and nothing measurable on Yorùbá, Igbo, Hausa, Vietnamese or Arabic.

Arabic: <ara> and <ara-nocase>

<ara> restores everything, including the grammatical case ending on each word's last letter (for teaching, classical and religious text, and text-to-speech). <ara-nocase> leaves word-final vowels off (shadda is kept), which is how most modern Arabic is vowelled when it is vowelled at all.

diacritize("وهذا قول مرغوب عنه .", "ara-nocase")   # 'وَهَذَا قَوْل مَرْغُوب عَنْه .'

Hamza letters (أ إ آ ؤ ئ) are spelling, not diacritics, and are kept as typed.

Meaning hints: [g: word=meaning]

Some words take different marks for different meanings (Vietnamese ranh: rảnh "free", rành "skilled", ranh "border"; Yorùbá ogun: ogun "war", ogún "twenty"). A hint in English steers the choice; separate several hints with |.

s = "Chi ay chi that su ranh vao nhung buoi toi sau khi da cho con ngu say."
diacritize(s, "vie", hint="[g: ranh=free (time)]")
# 'Chị ấy chỉ thật sự rảnh vào những buổi tối sau khi đã cho con ngủ say.'
diacritize(s, "vie")
# 'Chị ấy chỉ thật sự rành vào những buổi tối ...'   (without the hint: "skilled")

Hints help most where the sentence alone is not enough (Vietnamese, Turkish, Polish). When a hint contradicts a clear sentence context, the model usually follows the context.

Partial input

Text that already carries some marks is completed, not undone: marks you typed are kept and the rest are filled in.

diacritize("Ki lo ṣe ti inu rẹ ko ni dun?", "yor")   # underdots typed, tones missing
# 'Kí ló ṣe tí inú rẹ̀ kò ní dùn?'

Raw vs aligned output

aligned=True (recommended) guarantees that only marks change. aligned=False returns the model's own text, which also corrects typos it was trained to fix (swapped, missing or doubled letters), but may change a letter in a few percent of sentences.

Evaluation

Metrics. DER = share of mark-bearing letters (letters with more than one legal outcome in the language) whose marks differ from the reference. Strict scoring: an output that changes, adds or drops a letter counts every mark-bearing letter of that sentence as wrong. Exact = whole sentence correct. Compliance = output strips back to the input. Greedy decoding. Scored with the same code and letter definitions as the diactag-2.0 card, so the numbers compare directly.

diacbench (1,000 sentences per language)

diacbench: diacnet and diactag

Show table

DER on diacbench v1 references; diacnet outputs aligned.

Language diacnet-1.1 diacnet-mini-2.0 diacnet-2.0 diactag-2.0 diacnet-2.0 on diacbench v2 Exact (diacnet-2.0) Raw compliance (diacnet-2.0)
Yorùbá yor 0.1743 0.0815 0.0698 0.0787 0.0645 0.123 0.993
Igbo ibo 0.0219 0.0202 0.0185 0.0162 0.0160 0.461 0.934
Hausa hau 0.0063 0.0044 0.0041 0.0044 0.0023 0.745 0.958
Vietnamese vie 0.0316 0.0207 0.0143 0.0141 0.0140 0.635 0.999
Polish pol 0.0055 0.0024 0.0014 0.0035 0.0013 0.969 1.000
Turkish tur 0.0076 0.0042 0.0030 0.0035 0.0030 0.966 0.999
Portuguese por 0.0042 0.0027 0.0021 0.0032 0.0020 0.948 0.999
Spanish spa 0.0064 0.0031 0.0027 0.0038 0.0025 0.942 0.999
French fra 0.0028 0.0013 0.0012 0.0017 0.0012 0.965 1.000
Italian ita 0.0006 0.0003 0.0002 0.0005 0.0002 0.994 1.000
Mean 0.0261 0.0141 0.0117 0.0130 0.0107 0.775

Output alignment

raw vs aligned

Most of the raw error on Igbo, Hausa and Arabic comes from sentences where the model changed a letter, not from wrong marks. Alignment removes it.

Show table
Test Raw DER Aligned DER Raw compliance
Yorùbá 0.0780 0.0698 0.993
Igbo 0.0926 0.0185 0.934
Hausa 0.0485 0.0041 0.958
Vietnamese 0.0158 0.0143 0.999
Polish 0.0014 0.0014 1.000
Turkish 0.0037 0.0030 0.999
Portuguese 0.0029 0.0021 0.999
Spanish 0.0040 0.0027 0.999
French 0.0012 0.0012 1.000
Italian 0.0002 0.0002 1.000
Arabic, Classical (full) 0.0961 0.0955 0.998
Arabic, WikiNews (MSA) 0.3442 0.1542 0.733
Arabic, SadeedDiac-25 benchmark 0.2702 0.1313 0.778

Arabic

Arabic

Show table

DER over Arabic letters only, strict, on the same sentences for every system; ranked by WikiNews.

System WikiNews (MSA) SadeedDiac-25 benchmark Classical (diacbench ar) Classical (SadeedDiac-25 Fadel)
CATT-EO (Abjad AI) 0.0396 0.0473 0.0140 0.0132
diactag-2.0 0.0858 0.0994 0.0587 0.0625
Shakkala v3 0.1056 0.1068 0.0365 0.0456
Fine-Tashkeel 0.1492 0.2165 0.0136 0.0207
diacnet-2.0 (aligned) 0.1542 0.1313 0.0955 0.0855
diacnet-mini-2.0 (aligned) 0.2189 0.1685 0.1104 0.1082
Tashkeel-700M 0.2961 0.3087 0.0392 0.0445

No-case-endings mode (<ara-nocase>) on Classical Arabic: DER 0.0715. diacnet was not trained on Tashkeela, the corpus the Classical test sets come from (most other systems were), so the two Modern Standard Arabic columns are the fairer comparison. Test sets: SadeedDiac-25 (WikiNewsTruth, 146 sentences; benchmark, 454; Fadel test, 600) and diacbench ar (1,000).

Meaning hints

gloss hints

Show table

Accuracy on the ambiguous word in 580 held-out sentences whose ambiguous spellings never appeared in training. The last column gives the hint of the word's other meaning instead and counts how often the output switches.

Language No hint With hint Switches with the other meaning's hint (n)
Vietnamese 0.63 0.90 0.23 (92)
Turkish 0.75 0.84 0.50 (32)
Polish 0.93 0.98 0.00 (18)
Spanish 1.00 1.00 0.08 (12)
French 1.00 1.00 –
Yorùbá 0.44 0.44 0.37 (169)
Arabic 0.70 0.72 0.36 (28)

<auto> and partial input

Show tables

DER on the first 300 diacbench sentences per language (aligned):

Language Explicit tag <auto>
Yorùbá 0.0702 0.0685
Igbo 0.0160 0.0158
Hausa 0.0023 0.0022
Vietnamese 0.0140 0.0136
Polish 0.0020 0.0018
Turkish 0.0040 0.0035
Portuguese 0.0024 0.0025
Spanish 0.0027 0.0033
French 0.0014 0.0015
Italian 0.0002 0.0004
Arabic 0.0921 0.0927

Partial input, first 300 sentences (Yorùbá "underdots only": tones removed, underdots kept; "half the marks": a random half of the marked letters kept):

Language Marks given DER (aligned) DER from fully stripped input Given marks kept
Hausa half the marks 0.0021 0.0023 1.000
Igbo half the marks 0.0162 0.0160 1.000
Vietnamese half the marks 0.0145 0.0140 1.000
Yorùbá half the marks 0.0832 0.0702 0.999
Yorùbá underdots only 0.0771 0.0702 1.000

Best Practices

  1. Use align() unless you want typo correction: it costs nothing and removes letter changes.
  2. Give the language tag when you know it; <auto> is a safe default for mixed input.
  3. Split long text into sentences or chunks of about 300 characters (the model was trained on units of that size).
  4. Greedy decoding (num_beams=1), max_new_tokens about twice the input length.
  5. Choosing between Olaverse diacritizers:
diacnet-2.0 diacnet-mini-2.0 diactag-2.0
Parameters 582M 300M 37.9M
Approach generative (text to text) generative (text to text) character tagger
Strongest on Yorùbá and the European languages close to diacnet-2.0 at half the size Igbo, Vietnamese, Modern Standard Arabic
Extras meaning hints, typo correction meaning hints, typo correction int8 ONNX for CPU, never alters letters

Training

Trained from google/byt5-base for one epoch on about 2.66M (input, target) pairs: a language-balanced 2.5M-pair sample of diacnet-1.1-train (Yorùbá pairs kept only if well tone-marked, the cause of 1.1's Yorùbá regression) plus the teacher-marked set below, repeated 4 times. AdamW, learning rate 2e-4, 3% warmup then linear decay, weight decay 0.01, effective batch 128 (8 × 16 gradient accumulation), bf16, inputs and targets up to 768 bytes, length-grouped batches. Inputs follow 1.1's format and augmentations: <auto> 12%, partial marking 30%, meaning hints, character noise 7%.

Training Data and Licence

  • diacnet-1.1-train: FineWeb-2 and Wikipedia sentences in the ten Latin-script languages (ODC-By 1.0 / CC BY-SA 4.0).
  • About 37k teacher-marked pairs: real Yorùbá, Igbo, Hausa and Arabic news and web text re-marked in full by a large language model and kept only when its letters were unchanged, plus sentences built around words whose marks depend on meaning (Arabic in full and no-case-endings form). Classical Arabic from Fadel et al. (2019, MIT) was used only to confirm that rare spellings exist, not as training text.
  • Benchmark sentences were removed from the training data.

Released under Apache-2.0.

Citation

@misc{diacnet-2.0,
  title  = {diacnet-2.0},
  author = {Olaverse},
  year   = {2026},
  url    = {https://huggingface.co/olaverse/diacnet-2.0}
}
Downloads last month
2
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for olaverse/diacnet-2.0

Base model

google/byt5-base
Finetuned
(87)
this model

Datasets used to train olaverse/diacnet-2.0

Collection including olaverse/diacnet-2.0