SynRank quality classifiers

fastText classifiers that score a document by how much it looks like clean, knowledge-dense text in its language. From the ACL 2026 paper How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective.

Point one at a raw web corpus, keep the top-scoring documents, and you get a pretraining corpus that, in the paper's controlled comparison, matches what human experts produce with hand-written rules on knowledge benchmarks (XARC-E/XARC-C). You do not need an expert who speaks the language.

File Language
synrank_id.ftz Indonesian
synrank_sw.ftz Swahili
synrank_tr.ftz Turkish
synrank_vi.ftz Vietnamese

Each is the round-1 model — the second iteration, which has seen real high-scoring web text as well as synthetic positives. Round 0 has only seen machine-translated positives and over-selects for translationese. On 2026-09-28, each packaged file was checked against the original training outputs: its complete serialized arguments, dictionary and unquantized output matrix match iterative_model_1_epoch.bin, and differ from rounds 0 and 2. See the provenance record and machine-readable evidence for the checks and limits. Use SHA256SUMS to verify downloaded weight files.

Quantized with fastText product quantization for distribution: 7.71 GB → 969 MB each. Most of the unquantized size is fastText's 2,000,000 hashed n-gram buckets × 896 dims (≈7.17 GB). The 151,647 × 896 vocabulary embeddings add only ≈0.54 GB. Quantization can change both absolute scores and document rankings.

Recompute your threshold. Use a percentile of the scores produced by the model you actually run, and evaluate selection overlap on your corpus before substituting a quantized classifier for the original. A cutoff from the unquantized model need not transfer.

Usage

Install the scoring dependencies:

pip install fasttext transformers huggingface_hub

The example downloads the Indonesian classifier from tyzhu/mulsyn_classifiers. For another language, choose its filename from the table above.

Documents must be encoded the same way they were during training: tokenized with the Qwen2 tokenizer, with each token id written as the pseudo-word w<id>. Feeding raw text will produce meaningless scores.

import fasttext
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer

model_path = hf_hub_download(
    repo_id="tyzhu/mulsyn_classifiers",
    filename="synrank_id.ftz",
)
model = fasttext.load_model(model_path)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2-0.5B")

def quality_score(text: str) -> float:
    text = " ".join(text.replace("\\n", "\n").split())  # same flattening as mulsyn.io.extract_text
    ids = tokenizer(text, truncation=True, max_length=512)["input_ids"]
    words = " ".join(f"w{i}" for i in ids)
    labels, scores = model.predict([words], k=2)  # list input also works with NumPy 2
    return float(dict(zip(labels[0], scores[0]))["__label__positive"])

print(quality_score("Ilmu pengetahuan membantu kita memahami dunia."))

The returned score is the probability of __label__positive, even when the negative label is the most likely prediction. Rank documents within the classifier's language and choose a percentile cutoff for your corpus.

The paper measures ~980 documents/second on a standard CPU, so filtering a web-scale corpus is a matter of CPU parallelism rather than GPUs.

How they were trained

Positives: 10k documents from Cosmopedia machine-translated into the target language. The paper adds the highest-scoring real documents to the positives and filters with the round-1 model. With the MULSYN training code's defaults, training runs three rounds, promoting the top and bottom 10k of a 200k-document held-out pool into the positive and negative sets each round.

fastText: dim 896 (the Qwen 0.5B hidden size), lr 0.1, 5 epochs, word n-grams up to 3, min count 3, initialized from Qwen 0.5B input embeddings. The code defaults to Qwen/Qwen2-0.5B. A 2026-09-28 artifact audit found that sampled rows of the original exported embeddings and trained input matrices match Qwen2-0.5B; they do not match Qwen2.5-0.5B. Use Qwen2-0.5B for these files, even though the paper text names Qwen2.5-0.5B. Both have hidden size 896.

Held-out P/R/F1 on 10k synthetic positives + 10k noisy negatives, measured on the original unquantized models. The packaged files derive from round 1; these measurements were not rerun on the quantized files.

Language Round 0 P / R / F1 Round 1 (source) P / R / F1
Indonesian 99.50 / 99.80 / 99.65 91.52 / 99.95 / 95.55
Swahili 98.75 / 99.82 / 99.28 92.46 / 99.97 / 96.07
Turkish 98.98 / 99.75 / 99.36 91.60 / 99.96 / 95.60
Vietnamese 99.52 / 99.74 / 99.63 92.09 / 99.91 / 95.84

The round-0 column matches paper Table 4. The round-1 values match the saved val_stats_epoch_1.json files for the original round-1 models; they are not in the paper. Precision here measures agreement with the synthetic-versus-noisy labels. Lower precision means more noisy-set documents receive a positive prediction; it does not establish that those documents are fluent or useful for pretraining.

Neither column measures filter quality. The evidence for that is downstream, from 1.1B models pretrained in six languages. SynRank-filtered MADLAD beats the unfiltered corpus on XARC-C in all six languages and on XARC-E in five. It matches human-expert-cleaned corpora on those knowledge benchmarks (XARC-E 35.31 vs 35.32, XARC-C 24.56 vs 24.54), but not on XCOPA or XHellaswag.

Used in Sailor2

Several of the paper's authors are also authors of Sailor2. Sailor2 used a variant of this recipe to select the high-quality SEA web data in its Stage-2 (annealing) mixture (report §3.2). NLLB-3.3B translated high-quality English into each Southeast Asian language. One fastText classifier per language was trained on 10,000 translated positives (40% Cosmopedia, 40% MADLAD, 20% UltraChat) and 10,000 random CommonCrawl negatives. The top 20% of each language's CommonCrawl was kept for annealing (the blog says 10–20%). The Sailor2 organization card identifies sailor2/sea-commoncrawl-high-quality as that selection; it has 13 language folders and no Malay, so it does not match the Stage-2 table exactly. The models trained on the resulting Stage-2 data are in the Sailor2 collection (sail/Sailor2-1B, -8B, -20B and their chat versions).

These are not Sailor2's classifiers. Sailor2's classifiers are not public. The four models here are the paper's, trained with MADLAD-400 negatives. Indonesian and Vietnamese are Sailor2 languages; Swahili and Turkish are not.

Limitations

  • Language-specific. A classifier trained for Indonesian will not meaningfully score Turkish.
  • Selects for a style, not for truth. High-scoring documents are fluent and knowledge-dense in form. Fluent falsehoods score well.
  • Trained on MADLAD-400 / CommonCrawl web text. Scores on very different domains (code, chat logs, OCR output) are not calibrated.
  • Quantization can change scores and rankings. Check its effect on your corpus.

Citation

@inproceedings{zhu-etal-2026-synthetic,
    title = "How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective",
    author = "Zhu, Tongyao and Liu, Qian and Ma, Chang and Zhang, Jinghan and Dou, Longxu and He, Junxian and Chen, Shiqi",
    booktitle = "Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    year = "2026",
    url = "https://aclanthology.org/2026.acl-long.1002/",
}

If you also use the Sailor2 corpora or models, please cite Sailor2:

@misc{dou2025sailor2sailingsoutheastasia,
      title={Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs},
      author={Longxu Dou and Qian Liu and Fan Zhou and Changyu Chen and Zili Wang and Ziqi Jin and Zichen Liu and Tongyao Zhu and Cunxiao Du and Penghui Yang and Haonan Wang and Jiaheng Liu and Yongchi Zhao and Xiachong Feng and Xin Mao and Man Tsung Yeung and Kunat Pipatanakul and Fajri Koto and Min Si Thu and Hynek Kydl{\'\i}{\v{c}}ek and Zeyi Liu and Qunshu Lin and Sittipong Sripaisarnmongkol and Kridtaphad Sae-Khow and Nirattisai Thongchim and Taechawat Konkaew and Narong Borijindargoon and Anh Dao and Matichon Maneegard and Phakphum Artkaew and Zheng-Xin Yong and Quan Nguyen and Wannaphong Phatthiyaphaibun and Hoang H. Tran and Mike Zhang and Shiqi Chen and Tianyu Pang and Chao Du and Xinyi Wan and Wei Lu and Min Lin},
      year={2025},
      eprint={2502.12982},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2502.12982},
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for tyzhu/mulsyn_classifiers