is-mongolian — Mongolian Cyrillic (Khalkha) text classifier

A CPU-first ONNX classifier that answers one question about a piece of text:

Is this written in Mongolian Cyrillic, as the language is used in Mongolia (Khalkha, ISO 639-3 khk)?

It is a three-class model — MN / NON-MN / ABSTAIN — fine-tuned from FacebookAI/xlm-roberta-base (278 M parameters) and shipped as an int8 ONNX graph (278.7 MB) that runs on a plain CPU with onnxruntime, no GPU and no network access.

Task Text classification (MN vs NON-MN), + ABSTAIN
Base model FacebookAI/xlm-roberta-base (MIT)
Parameters 278,048,262
Input raw UTF-8 text, up to 192 sub-word tokens
Output 3 sentence logits + per-token O/B-MN/I-MN logits
Calibration temperature T = 0.597175 (fitted on the validation split)
Macro F1 (human gold set) 0.9544 (95 % CI 0.9484 – 0.9602)
Size / runtime 278.7 MB int8 · 1110.1 MB fp32 · CPUExecutionProvider
License Apache-2.0 (model card text, weights, inference code)

1. What "Mongolian" means here

This is the single most important thing to read before using the model. The scope is deliberate and narrow:

  • In scope: Cyrillic Mongolian as written in Mongolia — Khalkha Mongolian (khk). Texts that mix Mongolian with Latin (@mentions, URLs, emoji, English words) are in scope as long as the Mongolian content is Cyrillic Mongolian.
  • Out of scope (classified NON-MN by design):
    • Buryat and Yakut/Sakha — closely related Mongolic and Siberian-Turkic languages written in Cyrillic. They share most of the Mongolian Cyrillic letter inventory, so a naive "has ө/ү" rule cannot separate them. They are treated as not Mongolian for the intended use case.
    • Traditional (vertical) Mongolian script — entirely outside the scope of this model (see §9).
    • Kazakh, Kyrgyz, Uzbek, Tatar, Bashkir, Russian and Ukrainian.
  • The model classifies text, not accounts, authors, or users. No user-level information is used at inference time.

The negative classes are not arbitrary: the model was explicitly trained to separate Khalkha Mongolian from the Cyrillic languages it is most often confused with, including the neighbouring Mongolic/Turkic ones listed above.


2. Files in this repository

Path Contents
model.onnx Recommended. int8 dynamically-quantized graph, 278.7 MB (Gather/embedding included in quantization).
model_config.json Runtime configuration: label order, mn_index, abstain_index, temperature, max length, rule thresholds.
tokenizer.json, tokenizer_config.json, special_tokens_map.json SentencePiece tokenizer (identical to xlm-roberta-base).
fp32/model_fp32.onnx Full-precision graph, 1110.1 MB. Highest fidelity; 4× larger.
fp32/model_config.json Config for the fp32 graph.
audit/model_config.json Audit / filter mode: same weights, but rule-level abstention enabled (min_cyrillic = 5).
v1.0/ The previous model version (trained without auxiliary web data), for comparison and reproduction of the ablation in §7.
README.mn.md The same card in Mongolian.

Sub-folders reference the tokenizer and (for audit/) the weights at the repository root with relative paths, so nothing is duplicated.


3. Quickstart

Python (ONNX Runtime, CPU, offline)

import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer

snapshot = "aimongolia/is-mongolian"          # or a local directory
cfg = json.load(open(f"{snapshot}/model_config.json"))
tok = Tokenizer.from_file(f"{snapshot}/{cfg['tokenizer']}")
tok.enable_truncation(max_length=cfg["max_length"])
tok.enable_padding()  # dynamic: pad to the longest sequence in the batch

sess = ort.InferenceSession(
    f"{snapshot}/{cfg['onnx_file']}",
    providers=["CPUExecutionProvider"],
)

def classify(text: str) -> tuple[str, float]:
    ids = np.asarray([tok.encode(text).ids], dtype=np.int64)
    mask = np.ones_like(ids)
    logits = sess.run(None, {"input_ids": ids, "attention_mask": mask})[0][0]
    p = np.exp(logits / cfg["temperature"])
    p /= p.sum()
    i = int(p.argmax())
    return cfg["labels"][i], float(p[i])

classify("Энэ өдөр цаг агаар сайхан байна.")
# ('MN', 0.99...)
classify("Сегодня хорошая погода.")
# ('NON-MN', 0.99...)

ℹ️ Padding. The reference implementation pads each batch to its longest sequence (dynamic padding), so short posts do not pay for a full 192-token forward pass — see §8. Int8 graphs retain a small (~1–2 %) batch-size sensitivity in either mode; if you need strictly batch-invariant output, pad to the full max_length (enable_padding(length=cfg["max_length"])) instead and accept the higher latency.

The reference implementation used for all numbers in this card additionally exposes the rule thresholds (--min-cyrillic, --abstain-tau) and a CLI:

mn-lang --model aimongolia/is-mongolian detect "Сайн байна уу?" --verbose
cat tweets.txt | mn-lang --model aimongolia/is-mongolian detect --stdin --json

The standalone recipe above is self-sufficient: onnxruntime + tokenizers + the three files at the repository root are all that is required.

Downloading

Repository IDs are resolved through huggingface_hub.snapshot_download, so the model is cached locally and inference afterwards needs no network access.


4. Training data

Source Use Licence / terms
Social-media text (Mongolian Cyrillic tweets and replies) ~113 k training examples Text is not redistributed; only derived model weights are published
Wikipedia text in Buryat, Sakha, Kazakh, Kyrgyz, Uzbek Negative examples CC BY-SA (attribution: Wikimedia contributors)
Human-labelled gold set (5,046 items, two independent annotation passes) Evaluation only, never training Curated in-house, not redistributed
cis-lmu/GlotCC-V1 (CommonCrawl, glotlid-labelled) ~113 k auxiliary training examples: 56,632 Mongolian and 56,633 non-Mongolian (9 languages) CC0-1.0

Notes:

  • No source text is redistributed in this repository. The published artefacts are model weights, tokenizer files and metadata only. Tweet IDs, usernames and texts are not included anywhere in this repository or its documentation.
  • The gold set is a human-labelled benchmark and is kept completely disjoint from training data (exact-hash and near-duplicate screening).
  • The GlotCC auxiliary data is used only in the training split; it never enters validation or evaluation (machine labels are not ground truth).
  • The evaluation split below is the human gold set (test_b): 5,046 items, of which 399 were marked ABSTAIN by the annotators (both passes). Metrics are reported on the 4,647 items with a definite human label.

5. Training procedure

Hyper-parameter Value
Base checkpoint FacebookAI/xlm-roberta-base (a frozen local copy was used, so training is offline-reproducible)
Max length 192
Batch size × grad. accum. 32 × 2
Epochs 3
Learning rate 2e-5 (linear warmup, 6 % warmup ratio)
Weight decay 0.01
Label smoothing 0.05
Token-level loss weight 0.3 (auxiliary O/B-MN/I-MN head)
Precision AMP fp16 (single T4)
Seeds 42, 1337, 2024 — seed 42 is the published checkpoint
Model selection best validation macro F1 (never the test set)

The published checkpoint is seed 42 (best step 7500, validation macro F1 0.9981 on the in-distribution split). Across the three seeds the gold macro F1 is 0.9544 / 0.9312 / 0.9539 (mean 0.9465, spread 0.0232).


6. Evaluation — human gold set

Only test_b (human-labelled) is used for the headline numbers. Metrics are computed on the 4,647 items with a definite human label, after temperature calibration (T = 0.597175).

Metric Value
Macro F1 0.9544 (95 % CI 0.9484 – 0.9602)
Accuracy 0.9559
Balanced accuracy 0.9531
Cohen's κ (vs. majority human label) 0.9088
Expected calibration error 0.0367 (Brier 0.0810)
Class Precision Recall F1 Support
MN 0.9186 0.9688 0.9431 2,727
NON-MN 0.8295 0.9375 0.8802 1,920
ABSTAIN — — — 399 (never predicted)

Context on the ceiling. Two independent human annotation passes over the same items agree at κ = 0.9289 (99.31 % agreement on items that both passes labelled definitely). The model reaches κ = 0.9088 against the majority human label — i.e. it is close to, but below, the level at which the annotators agree with each other. Do not expect higher agreement without a task redefinition.

Baselines on the same gold set (rule-based, no training)

A well-known shortcut for this language pair is "does the text contain ө or ү?". It is a poor detector, and it is worth comparing against before trusting any accuracy figure on Mongolian text:

Baseline Macro F1
Contains ө/ү 0.4849
ө/ү + function words 0.6156
ө/ү minus Turkic-family letters (ә ң җ ғ қ һ ҕ) 0.5733
Cyrillic script only 0.3833
Long-vowel pattern rule 0.8711
This model 0.9544

The ө/ү shortcut has an AUC of 0.5488 on the gold set — statistically indistinguishable from chance, despite being ~0.95 accurate on the weak labels that are common in this domain. If you evaluate a Mongolian detector on auto-labelled data, this is the trap you will fall into.

Rule-level abstention (audit/)

The ABSTAIN class is essentially unused by the model itself (it never wins the argmax). Abstention is instead a two-stage decision:

Stage Setting Abstain rate Human-abstention recall Human-abstention precision
Rule: fewer than 5 Cyrillic letters min_cyrillic = 5 5.63 % 0.584 0.820
Probability threshold abstain_tau = 0.90 ~3.9 % 0.100 0.201
Both (union) — 9.00 % 0.622 0.546
Default (shipping) abstain_tau = 0.0, min_cyrillic = 0 0 % — —

So the rule is ~4× more precise than the model's own confidence at predicting where a human annotator hesitated, because human ABSTAIN is dominated by "too little Cyrillic evidence / truncated text", not by model uncertainty. The default configuration ships with all rules off; use audit/ (or the --min-cyrillic 5 flag) if you are triaging rather than classifying.


7. Domain shift and the auxiliary-data ablation

The evaluation set above is Twitter-like text. On formal web text (CommonCrawl sentences from GlotCC-V1, 9 negative languages × 263–281 items and 516 Khalkha positives, after removing every item that overlaps the training data), the model versions behave very differently:

Model version Macro F1 (web) Errors on Khalkha Errors on the 9 negative languages
v1.0 (social media only) 0.5262 0 / 516 1,356 / 2,463 (55.0 %)
intermediate (Wikipedia negatives) 0.9670 54 / 516 (10.5 %) 0 / 2,463
this model (v5.0, + GlotCC aux) 1.0000 0 / 516 0 / 2,463

The intermediate version had learned the style of social-media Mongolian well enough to reject legitimate Khalkha web sentences, while the first version accepted more than half of Buryat, Bashkir, Kazakh, Kyrgyz, Tatar and Yakut sentences as Mongolian. The auxiliary web data fixed both failure modes.

⚠️ This web set is in-distribution for the auxiliary training data, so 1.0000 here must not be read as a general capability claim. The human gold set in §6 is the measurement that matters.

v1.0/ in this repository is the first version, published so the ablation above can be reproduced independently.


8. Quantization and CPU performance

Quantization loss (int8 vs the fp32 reference, on the human gold set):

Precision Size Macro F1 Δ vs fp32 Row agreement
fp32 1110.1 MB 0.9544 — —
int8 278.7 MB 0.9563 +0.0019 0.9883
fp16 555.3 MB 0.9544 0.0000 1.0000

Numbers in this table are from the frozen gold revision n = 5,046 used at release time. The gold set was later extended to n = 5,535; on that revision the same int8 bundle scores 0.9476 vs 0.9487 for fp32 (Δ = −0.0011, agreement 0.987). The Δ stays far inside the ≤ 0.005 target either way.

fp16 is not shipped: it is twice the size of int8 and was measured 1.51× slower than fp32 on CPU (fp16 operators are emulated), while providing no accuracy benefit.

Latency (ONNX Runtime CPUExecutionProvider, 16 threads, dynamic padding — each batch is padded only to its longest sequence, Intel Xeon E5-2690 v3 @ 2.60 GHz, unloaded host):

Workload p50 p95 p99 Throughput
synthetic, batch 1 12.8 ms 28.6 ms 34.2 ms ~78 texts/s
real tweets (n = 5,081), batch 1 15.4 ms 26.3 ms 33.3 ms ~65 texts/s
synthetic, batch 32 6.94 ms / text 8.95 ms / text 9.04 ms / text 144 texts/s

Latency therefore scales with input length: short social-media posts run in ~13–15 ms, while a full 192-token input costs ~85 ms. Set pad_to_max_length=True if you need fixed-length, batch-invariant padding (at the cost of ~200 ms per request). int8 output itself carries a ~1–2 % batch-size sensitivity — an inherent quantization noise, also present with fixed padding. Measure on your own hardware; a re-calibrated temperature is not required, but re-validate if you truncate more text than the model was trained on.


9. Limitations

  1. Cyrillic only. Traditional (vertical) Mongolian script, and transliterated Mongolian in Latin script, are out of scope and will be classified NON-MN (or fall into the rule-level abstention if enabled).
  2. Buryat and Yakut are intentionally NON-MN. They are mutually intelligible with Mongolian in places and share the ө/ү letters. If your application needs to distinguish "Mongolic family" from "Mongolian of Mongolia", this model will not do it out of the box.
  3. Trained mostly on social-media text. Formal/technical prose is handled well (§7), but very short fragments ("за за тэгье"), song lyrics, and transliterated-in-formal-register texts are the observed error modes.
  4. Label noise in the original data. The underlying training labels were derived from a per-user flag in the source corpus rather than a per-text judgement, and auditing showed that the positive class contained a substantial share of closely related Turkic-language text. That is precisely why the auxiliary negative data and the human gold benchmark were introduced; the residual risk is that some of this bias remains in the weights.
  5. MN/NON-MN is not a quality judgement. The model does not detect spam, abuse, or language quality, and should not be used as a moderation signal on its own.
  6. ABSTAIN is a confidence mechanism, not a model of hesitation. The model's probability of ABSTAIN never exceeds 3.3e-5; abstention must be produced by the rule/threshold stages described in §6.
  7. Non-determinism: int8 dynamic quantization can shift a small share of predictions (1.2 % of rows differ from fp32). Use fp32/ if you need byte-level reproducibility against the PyTorch reference.

10. Ethical considerations and privacy

  • Training text came from publicly visible social-media posts and public dumps. No post text, post ID, username or user identifier is redistributed in this repository or in the model card.
  • Publishing derived weights fine-tuned on user-generated content carries a memorisation risk. The model is a text classifier, not a generative model, and the training texts were hash- and near-duplicate-screened against all evaluation sets, but no formal membership-inference audit was performed.
  • The model must not be used to infer the nationality, ethnicity or identity of a person or account. Language identification is not identity.
  • No personally identifying information was used to train or evaluate the model.

11. Reproducibility

  • Evaluation split: human gold set test_b (5,046 items) with a published MANIFEST.json (SHA-256).
  • Every number in this card is recorded in benchmarks/results.jsonl with the code fingerprint, seed, dataset version and git revision that produced it.
  • Model selection used the validation split only; the test split was evaluated once per configuration.
  • The auxiliary web corpus is cis-lmu/GlotCC-V1 (CC0-1.0) — public and re-downloadable.
  • The v1.0/ sub-folder allows the ablation table in §7 to be reproduced.

12. Citation and licence

@misc{aimongolia_is_mongolian,
  title        = {is-mongolian: a CPU-first Mongolian Cyrillic (Khalkha) text classifier},
  author       = {{aimongolia}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/aimongolia/is-mongolian}}
}

Released under the Apache License 2.0 (LICENSE). The base model FacebookAI/xlm-roberta-base is MIT-licensed; the GlotCC-V1 auxiliary corpus is CC0-1.0; the Wikipedia negative examples are CC BY-SA (Wikimedia contributors).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aimongolia/is-mongolian

Quantized
(36)
this model

Dataset used to train aimongolia/is-mongolian

Evaluation results

  • Macro F1 on Human-labelled gold set (test_b)
    self-reported
    0.954
  • Accuracy on Human-labelled gold set (test_b)
    self-reported
    0.956
  • Cohen's kappa on Human-labelled gold set (test_b)
    self-reported
    0.909