- is-mongolian — Mongolian Cyrillic (Khalkha) text classifier
- 1. What "Mongolian" means here
- 2. Files in this repository
- 3. Quickstart
- 4. Training data
- 5. Training procedure
- 6. Evaluation — human gold set
- 7. Domain shift and the auxiliary-data ablation
- 8. Quantization and CPU performance
- 9. Limitations
- 10. Ethical considerations and privacy
- 11. Reproducibility
- 12. Citation and licence
- 1. What "Mongolian" means here
is-mongolian — Mongolian Cyrillic (Khalkha) text classifier
A CPU-first ONNX classifier that answers one question about a piece of text:
Is this written in Mongolian Cyrillic, as the language is used in Mongolia (Khalkha, ISO 639-3
khk)?
It is a three-class model — MN / NON-MN / ABSTAIN — fine-tuned from
FacebookAI/xlm-roberta-base (278 M parameters) and shipped as an int8 ONNX graph
(278.7 MB) that runs on a plain CPU with onnxruntime, no GPU and no network access.
| Task | Text classification (MN vs NON-MN), + ABSTAIN |
| Base model | FacebookAI/xlm-roberta-base (MIT) |
| Parameters | 278,048,262 |
| Input | raw UTF-8 text, up to 192 sub-word tokens |
| Output | 3 sentence logits + per-token O/B-MN/I-MN logits |
| Calibration | temperature T = 0.597175 (fitted on the validation split) |
| Macro F1 (human gold set) | 0.9544 (95 % CI 0.9484 – 0.9602) |
| Size / runtime | 278.7 MB int8 · 1110.1 MB fp32 · CPUExecutionProvider |
| License | Apache-2.0 (model card text, weights, inference code) |
1. What "Mongolian" means here
This is the single most important thing to read before using the model. The scope is deliberate and narrow:
- In scope: Cyrillic Mongolian as written in Mongolia — Khalkha Mongolian (
khk). Texts that mix Mongolian with Latin (@mentions, URLs, emoji, English words) are in scope as long as the Mongolian content is Cyrillic Mongolian. - Out of scope (classified
NON-MNby design):- Buryat and Yakut/Sakha — closely related Mongolic and Siberian-Turkic
languages written in Cyrillic. They share most of the Mongolian Cyrillic letter
inventory, so a naive "has
ө/ү" rule cannot separate them. They are treated as not Mongolian for the intended use case. - Traditional (vertical) Mongolian script — entirely outside the scope of this model (see §9).
- Kazakh, Kyrgyz, Uzbek, Tatar, Bashkir, Russian and Ukrainian.
- Buryat and Yakut/Sakha — closely related Mongolic and Siberian-Turkic
languages written in Cyrillic. They share most of the Mongolian Cyrillic letter
inventory, so a naive "has
- The model classifies text, not accounts, authors, or users. No user-level information is used at inference time.
The negative classes are not arbitrary: the model was explicitly trained to separate Khalkha Mongolian from the Cyrillic languages it is most often confused with, including the neighbouring Mongolic/Turkic ones listed above.
2. Files in this repository
| Path | Contents |
|---|---|
model.onnx |
Recommended. int8 dynamically-quantized graph, 278.7 MB (Gather/embedding included in quantization). |
model_config.json |
Runtime configuration: label order, mn_index, abstain_index, temperature, max length, rule thresholds. |
tokenizer.json, tokenizer_config.json, special_tokens_map.json |
SentencePiece tokenizer (identical to xlm-roberta-base). |
fp32/model_fp32.onnx |
Full-precision graph, 1110.1 MB. Highest fidelity; 4× larger. |
fp32/model_config.json |
Config for the fp32 graph. |
audit/model_config.json |
Audit / filter mode: same weights, but rule-level abstention enabled (min_cyrillic = 5). |
v1.0/ |
The previous model version (trained without auxiliary web data), for comparison and reproduction of the ablation in §7. |
README.mn.md |
The same card in Mongolian. |
Sub-folders reference the tokenizer and (for audit/) the weights at the repository
root with relative paths, so nothing is duplicated.
3. Quickstart
Python (ONNX Runtime, CPU, offline)
import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
snapshot = "aimongolia/is-mongolian" # or a local directory
cfg = json.load(open(f"{snapshot}/model_config.json"))
tok = Tokenizer.from_file(f"{snapshot}/{cfg['tokenizer']}")
tok.enable_truncation(max_length=cfg["max_length"])
tok.enable_padding() # dynamic: pad to the longest sequence in the batch
sess = ort.InferenceSession(
f"{snapshot}/{cfg['onnx_file']}",
providers=["CPUExecutionProvider"],
)
def classify(text: str) -> tuple[str, float]:
ids = np.asarray([tok.encode(text).ids], dtype=np.int64)
mask = np.ones_like(ids)
logits = sess.run(None, {"input_ids": ids, "attention_mask": mask})[0][0]
p = np.exp(logits / cfg["temperature"])
p /= p.sum()
i = int(p.argmax())
return cfg["labels"][i], float(p[i])
classify("Энэ өдөр цаг агаар сайхан байна.")
# ('MN', 0.99...)
classify("Сегодня хорошая погода.")
# ('NON-MN', 0.99...)
ℹ️ Padding. The reference implementation pads each batch to its longest sequence (dynamic padding), so short posts do not pay for a full 192-token forward pass — see §8. Int8 graphs retain a small (~1–2 %) batch-size sensitivity in either mode; if you need strictly batch-invariant output, pad to the full
max_length(enable_padding(length=cfg["max_length"])) instead and accept the higher latency.
The reference implementation used for all numbers in this card additionally exposes the
rule thresholds (--min-cyrillic, --abstain-tau) and a CLI:
mn-lang --model aimongolia/is-mongolian detect "Сайн байна уу?" --verbose
cat tweets.txt | mn-lang --model aimongolia/is-mongolian detect --stdin --json
The standalone recipe above is self-sufficient: onnxruntime + tokenizers + the three
files at the repository root are all that is required.
Downloading
Repository IDs are resolved through huggingface_hub.snapshot_download, so the model
is cached locally and inference afterwards needs no network access.
4. Training data
| Source | Use | Licence / terms |
|---|---|---|
| Social-media text (Mongolian Cyrillic tweets and replies) | ~113 k training examples | Text is not redistributed; only derived model weights are published |
| Wikipedia text in Buryat, Sakha, Kazakh, Kyrgyz, Uzbek | Negative examples | CC BY-SA (attribution: Wikimedia contributors) |
| Human-labelled gold set (5,046 items, two independent annotation passes) | Evaluation only, never training | Curated in-house, not redistributed |
cis-lmu/GlotCC-V1 (CommonCrawl, glotlid-labelled) |
~113 k auxiliary training examples: 56,632 Mongolian and 56,633 non-Mongolian (9 languages) | CC0-1.0 |
Notes:
- No source text is redistributed in this repository. The published artefacts are model weights, tokenizer files and metadata only. Tweet IDs, usernames and texts are not included anywhere in this repository or its documentation.
- The gold set is a human-labelled benchmark and is kept completely disjoint from training data (exact-hash and near-duplicate screening).
- The GlotCC auxiliary data is used only in the training split; it never enters validation or evaluation (machine labels are not ground truth).
- The evaluation split below is the human gold set (
test_b): 5,046 items, of which 399 were markedABSTAINby the annotators (both passes). Metrics are reported on the 4,647 items with a definite human label.
5. Training procedure
| Hyper-parameter | Value |
|---|---|
| Base checkpoint | FacebookAI/xlm-roberta-base (a frozen local copy was used, so training is offline-reproducible) |
| Max length | 192 |
| Batch size × grad. accum. | 32 × 2 |
| Epochs | 3 |
| Learning rate | 2e-5 (linear warmup, 6 % warmup ratio) |
| Weight decay | 0.01 |
| Label smoothing | 0.05 |
| Token-level loss weight | 0.3 (auxiliary O/B-MN/I-MN head) |
| Precision | AMP fp16 (single T4) |
| Seeds | 42, 1337, 2024 — seed 42 is the published checkpoint |
| Model selection | best validation macro F1 (never the test set) |
The published checkpoint is seed 42 (best step 7500, validation macro F1 0.9981 on the in-distribution split). Across the three seeds the gold macro F1 is 0.9544 / 0.9312 / 0.9539 (mean 0.9465, spread 0.0232).
6. Evaluation — human gold set
Only test_b (human-labelled) is used for the headline numbers. Metrics are computed on
the 4,647 items with a definite human label, after temperature calibration
(T = 0.597175).
| Metric | Value |
|---|---|
| Macro F1 | 0.9544 (95 % CI 0.9484 – 0.9602) |
| Accuracy | 0.9559 |
| Balanced accuracy | 0.9531 |
| Cohen's κ (vs. majority human label) | 0.9088 |
| Expected calibration error | 0.0367 (Brier 0.0810) |
| Class | Precision | Recall | F1 | Support |
|---|---|---|---|---|
MN |
0.9186 | 0.9688 | 0.9431 | 2,727 |
NON-MN |
0.8295 | 0.9375 | 0.8802 | 1,920 |
ABSTAIN |
— | — | — | 399 (never predicted) |
Context on the ceiling. Two independent human annotation passes over the same items agree at κ = 0.9289 (99.31 % agreement on items that both passes labelled definitely). The model reaches κ = 0.9088 against the majority human label — i.e. it is close to, but below, the level at which the annotators agree with each other. Do not expect higher agreement without a task redefinition.
Baselines on the same gold set (rule-based, no training)
A well-known shortcut for this language pair is "does the text contain ө or ү?".
It is a poor detector, and it is worth comparing against before trusting any accuracy
figure on Mongolian text:
| Baseline | Macro F1 |
|---|---|
Contains ө/ү |
0.4849 |
ө/ү + function words |
0.6156 |
ө/ү minus Turkic-family letters (ә ң җ ғ қ һ ҕ) |
0.5733 |
| Cyrillic script only | 0.3833 |
| Long-vowel pattern rule | 0.8711 |
| This model | 0.9544 |
The ө/ү shortcut has an AUC of 0.5488 on the gold set — statistically
indistinguishable from chance, despite being ~0.95 accurate on the weak labels that are
common in this domain. If you evaluate a Mongolian detector on auto-labelled data, this
is the trap you will fall into.
Rule-level abstention (audit/)
The ABSTAIN class is essentially unused by the model itself (it never wins the
argmax). Abstention is instead a two-stage decision:
| Stage | Setting | Abstain rate | Human-abstention recall | Human-abstention precision |
|---|---|---|---|---|
| Rule: fewer than 5 Cyrillic letters | min_cyrillic = 5 |
5.63 % | 0.584 | 0.820 |
| Probability threshold | abstain_tau = 0.90 |
~3.9 % | 0.100 | 0.201 |
| Both (union) | — | 9.00 % | 0.622 | 0.546 |
| Default (shipping) | abstain_tau = 0.0, min_cyrillic = 0 |
0 % | — | — |
So the rule is ~4× more precise than the model's own confidence at predicting where a
human annotator hesitated, because human ABSTAIN is dominated by
"too little Cyrillic evidence / truncated text", not by model uncertainty.
The default configuration ships with all rules off; use audit/ (or the
--min-cyrillic 5 flag) if you are triaging rather than classifying.
7. Domain shift and the auxiliary-data ablation
The evaluation set above is Twitter-like text. On formal web text (CommonCrawl
sentences from GlotCC-V1, 9 negative languages × 263–281 items and 516 Khalkha
positives, after removing every item that overlaps the training data), the model
versions behave very differently:
| Model version | Macro F1 (web) | Errors on Khalkha | Errors on the 9 negative languages |
|---|---|---|---|
v1.0 (social media only) |
0.5262 | 0 / 516 | 1,356 / 2,463 (55.0 %) |
| intermediate (Wikipedia negatives) | 0.9670 | 54 / 516 (10.5 %) | 0 / 2,463 |
this model (v5.0, + GlotCC aux) |
1.0000 | 0 / 516 | 0 / 2,463 |
The intermediate version had learned the style of social-media Mongolian well enough to reject legitimate Khalkha web sentences, while the first version accepted more than half of Buryat, Bashkir, Kazakh, Kyrgyz, Tatar and Yakut sentences as Mongolian. The auxiliary web data fixed both failure modes.
⚠️ This web set is in-distribution for the auxiliary training data, so 1.0000 here must not be read as a general capability claim. The human gold set in §6 is the measurement that matters.
v1.0/ in this repository is the first version, published so the ablation above can be
reproduced independently.
8. Quantization and CPU performance
Quantization loss (int8 vs the fp32 reference, on the human gold set):
| Precision | Size | Macro F1 | Δ vs fp32 | Row agreement |
|---|---|---|---|---|
| fp32 | 1110.1 MB | 0.9544 | — | — |
| int8 | 278.7 MB | 0.9563 | +0.0019 | 0.9883 |
| fp16 | 555.3 MB | 0.9544 | 0.0000 | 1.0000 |
Numbers in this table are from the frozen gold revision
n = 5,046used at release time. The gold set was later extended ton = 5,535; on that revision the same int8 bundle scores 0.9476 vs 0.9487 for fp32 (Δ = −0.0011, agreement 0.987). The Δ stays far inside the ≤ 0.005 target either way.
fp16 is not shipped: it is twice the size of int8 and was measured 1.51× slower than fp32 on CPU (fp16 operators are emulated), while providing no accuracy benefit.
Latency (ONNX Runtime CPUExecutionProvider, 16 threads, dynamic padding —
each batch is padded only to its longest sequence, Intel Xeon E5-2690 v3 @ 2.60 GHz,
unloaded host):
| Workload | p50 | p95 | p99 | Throughput |
|---|---|---|---|---|
| synthetic, batch 1 | 12.8 ms | 28.6 ms | 34.2 ms | ~78 texts/s |
| real tweets (n = 5,081), batch 1 | 15.4 ms | 26.3 ms | 33.3 ms | ~65 texts/s |
| synthetic, batch 32 | 6.94 ms / text | 8.95 ms / text | 9.04 ms / text | 144 texts/s |
Latency therefore scales with input length: short social-media posts run in ~13–15 ms,
while a full 192-token input costs ~85 ms. Set pad_to_max_length=True if you need
fixed-length, batch-invariant padding (at the cost of ~200 ms per request). int8 output
itself carries a ~1–2 % batch-size sensitivity — an inherent quantization noise, also
present with fixed padding. Measure on your own hardware; a re-calibrated temperature
is not required, but re-validate if you truncate more text than the model was trained on.
9. Limitations
- Cyrillic only. Traditional (vertical) Mongolian script, and transliterated
Mongolian in Latin script, are out of scope and will be classified
NON-MN(or fall into the rule-level abstention if enabled). - Buryat and Yakut are intentionally
NON-MN. They are mutually intelligible with Mongolian in places and share theө/үletters. If your application needs to distinguish "Mongolic family" from "Mongolian of Mongolia", this model will not do it out of the box. - Trained mostly on social-media text. Formal/technical prose is handled well
(§7), but very short fragments (
"за за тэгье"), song lyrics, and transliterated-in-formal-register texts are the observed error modes. - Label noise in the original data. The underlying training labels were derived from a per-user flag in the source corpus rather than a per-text judgement, and auditing showed that the positive class contained a substantial share of closely related Turkic-language text. That is precisely why the auxiliary negative data and the human gold benchmark were introduced; the residual risk is that some of this bias remains in the weights.
MN/NON-MNis not a quality judgement. The model does not detect spam, abuse, or language quality, and should not be used as a moderation signal on its own.ABSTAINis a confidence mechanism, not a model of hesitation. The model's probability ofABSTAINnever exceeds 3.3e-5; abstention must be produced by the rule/threshold stages described in §6.- Non-determinism: int8 dynamic quantization can shift a small share of
predictions (1.2 % of rows differ from fp32). Use
fp32/if you need byte-level reproducibility against the PyTorch reference.
10. Ethical considerations and privacy
- Training text came from publicly visible social-media posts and public dumps. No post text, post ID, username or user identifier is redistributed in this repository or in the model card.
- Publishing derived weights fine-tuned on user-generated content carries a memorisation risk. The model is a text classifier, not a generative model, and the training texts were hash- and near-duplicate-screened against all evaluation sets, but no formal membership-inference audit was performed.
- The model must not be used to infer the nationality, ethnicity or identity of a person or account. Language identification is not identity.
- No personally identifying information was used to train or evaluate the model.
11. Reproducibility
- Evaluation split: human gold set
test_b(5,046 items) with a publishedMANIFEST.json(SHA-256). - Every number in this card is recorded in
benchmarks/results.jsonlwith the code fingerprint, seed, dataset version and git revision that produced it. - Model selection used the validation split only; the test split was evaluated once per configuration.
- The auxiliary web corpus is
cis-lmu/GlotCC-V1(CC0-1.0) — public and re-downloadable. - The
v1.0/sub-folder allows the ablation table in §7 to be reproduced.
12. Citation and licence
@misc{aimongolia_is_mongolian,
title = {is-mongolian: a CPU-first Mongolian Cyrillic (Khalkha) text classifier},
author = {{aimongolia}},
year = {2026},
howpublished = {\url{https://huggingface.co/aimongolia/is-mongolian}}
}
Released under the Apache License 2.0 (LICENSE). The base model
FacebookAI/xlm-roberta-base is MIT-licensed; the GlotCC-V1 auxiliary corpus is
CC0-1.0; the Wikipedia negative examples are CC BY-SA (Wikimedia contributors).
Model tree for aimongolia/is-mongolian
Base model
FacebookAI/xlm-roberta-baseDataset used to train aimongolia/is-mongolian
Evaluation results
- Macro F1 on Human-labelled gold set (test_b)self-reported0.954
- Accuracy on Human-labelled gold set (test_b)self-reported0.956
- Cohen's kappa on Human-labelled gold set (test_b)self-reported0.909