Instructions to use tyzhu/mulsyn_classifiers with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use tyzhu/mulsyn_classifiers with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("tyzhu/mulsyn_classifiers", "model.bin")) - Notebooks
- Google Colab
- Kaggle
SynRank quality classifiers
fastText classifiers that score a document by how much it looks like clean, knowledge-dense text in its language. From the ACL 2026 paper How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective.
Point one at a raw web corpus, keep the top-scoring documents, and you get a pretraining corpus that, in the paper's controlled comparison, matches what human experts produce with hand-written rules on knowledge benchmarks (XARC-E/XARC-C). You do not need an expert who speaks the language.
| File | Language |
|---|---|
synrank_id.ftz |
Indonesian |
synrank_sw.ftz |
Swahili |
synrank_tr.ftz |
Turkish |
synrank_vi.ftz |
Vietnamese |
Each is the round-1 model — the second iteration, which has seen real
high-scoring web text as well as synthetic positives. Round 0 has only seen
machine-translated positives and over-selects for translationese.
On 2026-09-28, each packaged file was checked against the original training
outputs: its complete serialized arguments, dictionary and unquantized output
matrix match iterative_model_1_epoch.bin, and differ from rounds 0 and 2.
See the provenance record and
machine-readable evidence for the checks and limits.
Use SHA256SUMS to verify downloaded weight files.
Quantized with fastText product quantization for distribution: 7.71 GB → 969 MB each. Most of the unquantized size is fastText's 2,000,000 hashed n-gram buckets × 896 dims (≈7.17 GB). The 151,647 × 896 vocabulary embeddings add only ≈0.54 GB. Quantization can change both absolute scores and document rankings.
Recompute your threshold. Use a percentile of the scores produced by the model you actually run, and evaluate selection overlap on your corpus before substituting a quantized classifier for the original. A cutoff from the unquantized model need not transfer.
Usage
Install the scoring dependencies:
pip install fasttext transformers huggingface_hub
The example downloads the Indonesian classifier from
tyzhu/mulsyn_classifiers.
For another language, choose its filename from the table above.
Documents must be encoded the same way they were during training: tokenized with
the Qwen2 tokenizer, with each token id written as the pseudo-word w<id>.
Feeding raw text will produce meaningless scores.
import fasttext
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
model_path = hf_hub_download(
repo_id="tyzhu/mulsyn_classifiers",
filename="synrank_id.ftz",
)
model = fasttext.load_model(model_path)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2-0.5B")
def quality_score(text: str) -> float:
text = " ".join(text.replace("\\n", "\n").split()) # same flattening as mulsyn.io.extract_text
ids = tokenizer(text, truncation=True, max_length=512)["input_ids"]
words = " ".join(f"w{i}" for i in ids)
labels, scores = model.predict([words], k=2) # list input also works with NumPy 2
return float(dict(zip(labels[0], scores[0]))["__label__positive"])
print(quality_score("Ilmu pengetahuan membantu kita memahami dunia."))
The returned score is the probability of __label__positive, even when the
negative label is the most likely prediction. Rank documents within the
classifier's language and choose a percentile cutoff for your corpus.
The paper measures ~980 documents/second on a standard CPU, so filtering a web-scale corpus is a matter of CPU parallelism rather than GPUs.
How they were trained
Positives: 10k documents from Cosmopedia machine-translated into the target language. The paper adds the highest-scoring real documents to the positives and filters with the round-1 model. With the MULSYN training code's defaults, training runs three rounds, promoting the top and bottom 10k of a 200k-document held-out pool into the positive and negative sets each round.
fastText: dim 896 (the Qwen 0.5B hidden size), lr 0.1, 5 epochs, word n-grams
up to 3, min count 3, initialized from Qwen 0.5B input embeddings. The code
defaults to Qwen/Qwen2-0.5B. A 2026-09-28 artifact audit found that sampled
rows of the original exported embeddings and trained input matrices match
Qwen2-0.5B; they do not match Qwen2.5-0.5B. Use Qwen2-0.5B for these files,
even though the paper text names Qwen2.5-0.5B. Both have hidden size 896.
Held-out P/R/F1 on 10k synthetic positives + 10k noisy negatives, measured on the original unquantized models. The packaged files derive from round 1; these measurements were not rerun on the quantized files.
| Language | Round 0 P / R / F1 | Round 1 (source) P / R / F1 |
|---|---|---|
| Indonesian | 99.50 / 99.80 / 99.65 | 91.52 / 99.95 / 95.55 |
| Swahili | 98.75 / 99.82 / 99.28 | 92.46 / 99.97 / 96.07 |
| Turkish | 98.98 / 99.75 / 99.36 | 91.60 / 99.96 / 95.60 |
| Vietnamese | 99.52 / 99.74 / 99.63 | 92.09 / 99.91 / 95.84 |
The round-0 column matches paper Table 4. The round-1 values match the saved
val_stats_epoch_1.json files for the original round-1 models; they are not
in the paper.
Precision here measures agreement with the synthetic-versus-noisy labels.
Lower precision means more noisy-set documents receive a positive prediction;
it does not establish that those documents are fluent or useful for pretraining.
Neither column measures filter quality. The evidence for that is downstream, from 1.1B models pretrained in six languages. SynRank-filtered MADLAD beats the unfiltered corpus on XARC-C in all six languages and on XARC-E in five. It matches human-expert-cleaned corpora on those knowledge benchmarks (XARC-E 35.31 vs 35.32, XARC-C 24.56 vs 24.54), but not on XCOPA or XHellaswag.
Used in Sailor2
Several of the paper's authors are also authors of
Sailor2. Sailor2 used a variant of this
recipe to select the high-quality SEA web data in its Stage-2 (annealing)
mixture (report §3.2). NLLB-3.3B
translated high-quality English into each Southeast Asian language. One fastText
classifier per language was trained on 10,000 translated positives (40%
Cosmopedia, 40% MADLAD, 20% UltraChat) and 10,000 random CommonCrawl negatives.
The top 20% of each language's CommonCrawl was kept for annealing (the
blog
says 10–20%). The Sailor2 organization card identifies
sailor2/sea-commoncrawl-high-quality
as that selection; it has 13 language folders and no Malay, so it does not
match the Stage-2 table exactly.
The models trained on the resulting Stage-2 data are in the
Sailor2 collection
(sail/Sailor2-1B, -8B, -20B and their chat versions).
These are not Sailor2's classifiers. Sailor2's classifiers are not public. The four models here are the paper's, trained with MADLAD-400 negatives. Indonesian and Vietnamese are Sailor2 languages; Swahili and Turkish are not.
Limitations
- Language-specific. A classifier trained for Indonesian will not meaningfully score Turkish.
- Selects for a style, not for truth. High-scoring documents are fluent and knowledge-dense in form. Fluent falsehoods score well.
- Trained on MADLAD-400 / CommonCrawl web text. Scores on very different domains (code, chat logs, OCR output) are not calibrated.
- Quantization can change scores and rankings. Check its effect on your corpus.
Citation
@inproceedings{zhu-etal-2026-synthetic,
title = "How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective",
author = "Zhu, Tongyao and Liu, Qian and Ma, Chang and Zhang, Jinghan and Dou, Longxu and He, Junxian and Chen, Shiqi",
booktitle = "Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
year = "2026",
url = "https://aclanthology.org/2026.acl-long.1002/",
}
If you also use the Sailor2 corpora or models, please cite Sailor2:
@misc{dou2025sailor2sailingsoutheastasia,
title={Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs},
author={Longxu Dou and Qian Liu and Fan Zhou and Changyu Chen and Zili Wang and Ziqi Jin and Zichen Liu and Tongyao Zhu and Cunxiao Du and Penghui Yang and Haonan Wang and Jiaheng Liu and Yongchi Zhao and Xiachong Feng and Xin Mao and Man Tsung Yeung and Kunat Pipatanakul and Fajri Koto and Min Si Thu and Hynek Kydl{\'\i}{\v{c}}ek and Zeyi Liu and Qunshu Lin and Sittipong Sripaisarnmongkol and Kridtaphad Sae-Khow and Nirattisai Thongchim and Taechawat Konkaew and Narong Borijindargoon and Anh Dao and Matichon Maneegard and Phakphum Artkaew and Zheng-Xin Yong and Quan Nguyen and Wannaphong Phatthiyaphaibun and Hoang H. Tran and Mike Zhang and Shiqi Chen and Tianyu Pang and Chao Du and Xinyi Wan and Wei Lu and Min Lin},
year={2025},
eprint={2502.12982},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.12982},
}
- Downloads last month
- -