Language Detection (small, 141M): 183 languages

A fast language identifier for 183 languages (the FLORES-200 set), built on the multilingual mmBERT-small encoder. It works on sentences and on short strings (titles, chat messages, search queries), runs with the standard transformers text-classification pipeline, and ships ONNX files for CPU and the browser (transformers.js). Apache-2.0; trained only on openly licensed web text.

  • Labels are FLORES-200 codes: ISO 639-3 language + ISO 15924 script, e.g. eng_Latn, hin_Deva, srp_Cyrl; the full list is in labels.json. Chinese is one label, zho_Hani (Simplified and Traditional together), and Arabic is one label, arb_Arab: the dialect labels (acm_Arab, aeb_Arab, apc_Arab, ars_Arab, ary_Arab, arz_Arab) were too unreliable (they pulled Modern Standard Arabic away from the right answer) and are merged; Dyula (dyu_Latn) is merged into Bambara (bam_Latn).
  • You can restrict the prediction to the languages you expect (see Usage), which makes it more accurate on short text.
  • onnx/model_quantized.onnx (int8 embeddings, 269 MB) picks the same language as fp32 on 100.0% of 400 FLORES sentences and prefixes.

Usage

from transformers import pipeline

lid = pipeline("text-classification", model="Horizon-Labs/language-detection-small", top_k=3)
lid("Je voudrais réserver une table pour deux personnes.")
# [[{'label': 'fra_Latn', 'score': 0.99...}, ...]]

Restrict to candidate languages (recommended when you know the possible set):

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tok = AutoTokenizer.from_pretrained("Horizon-Labs/language-detection-small")
model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/language-detection-small").eval()
allowed = ["eng_Latn", "spa_Latn", "por_Latn", "fra_Latn"]
mask = torch.full((model.config.num_labels,), float("-inf"))
mask[[model.config.label2id[l] for l in allowed]] = 0
with torch.no_grad():
    logits = model(**tok(["obrigado pela ajuda"], return_tensors="pt")).logits + mask
print(model.config.id2label[int(logits.argmax(-1))])   # por_Latn

transformers.js:

import { pipeline } from "@huggingface/transformers";
const lid = await pipeline("text-classification", "Horizon-Labs/language-detection-small", { dtype: "q8" });
console.log(await lid("Dziękuję bardzo!", { top_k: 3 }));

Evaluation

FLORES-200 devtest (1,012 professionally translated sentences per language; evaluation only, never trained on - any training text containing a FLORES sentence was removed), on the 183 languages this model supports. "First 20-40 characters" cuts every sentence to a short prefix, to measure short-text behaviour. Every model is scored on the same sentences with the same script; GlotLID and fastText can predict more languages than these, which can only cost them. The same label merges (Arabic dialects, Dyula, Akan/Twi, Chinese scripts) are applied to every model's predictions.

this model (141M) GlotLID (fastText, 1.7 GB) fastText LID-218 (NLLB, CC-BY-NC)
FLORES-200 devtest, full sentences: accuracy 0.984 0.958 0.970
FLORES-200 devtest, full sentences: macro-F1 0.983 0.958 0.970
First 20-40 characters: accuracy 0.874 0.845 0.837
First 20-40 characters: macro-F1 0.866 0.852 0.831

Against the most-downloaded transformers language detector, on its 20 languages (it can only predict those 20, so the fair comparison is with our model restricted to the same 20):

papluca's 20 languages this model, all 183 labels this model, restricted to the 20 papluca/xlm-roberta-base-language-detection (278M)
full sentences (accuracy) 0.995 1.000 0.994
first 20-40 characters (accuracy) 0.953 0.991 0.973

Where it is weak

Languages below 0.90 accuracy on full sentences (mostly close varieties such as Bosnian/Croatian and Hindi-belt languages, or languages with little training text):

language code this model GlotLID
bos_Latn 0.46 0.46
taq_Latn 0.51 0.97
hrv_Latn 0.76 0.95
kin_Latn 0.81 0.95
mag_Deva 0.85 0.98
mos_Latn 0.90 0.98
zsm_Latn 0.90 0.88

Not supported: ace_Arab, acq_Arab, ajp_Arab, bug_Latn, cjk_Latn, kas_Deva, knc_Arab, min_Arab, nus_Latn, prs_Arab, taq_Tfng, yue_Hant (no or too little FineWeb-2 text; Cantonese yue_Hant because it could not be told apart from Mandarin and took Mandarin predictions).

Training

  • Data: FineWeb-2 and, for English, FineWeb (both ODC-BY): up to 6,000 samples per language from the first files of each language subset, as 1-3 sentence snippets and 10-60 character spans (about 1.6M samples).
  • Cleaning: FineWeb-2's language labels come from an automatic identifier (GlotLID), and its subsets contain off-language pages. We dropped samples whose script does not match the label's script, and non-English Latin-script samples that are mostly English function words (about 30k samples, 2%); Akan and Twi share one subset and are one label (aka_Latn). The model still inherits some label noise from FineWeb-2.
  • Model: mmBERT-small with a 183-way classification head, 2 epochs, max length 128 tokens.
  • Code: code/ in this repository.

Limitations

  • Close varieties are often confused (see the table above), e.g. Bosnian vs Croatian; Arabic dialects and Cantonese are not distinguished (Arabic dialects are reported as arb_Arab; Cantonese is not supported).
  • Accuracy drops on very short inputs (a few words); restrict the candidates when you can. On 64 everyday phrases we wrote ourselves (4 per language in 16 major languages, e.g. "Thank you very much for your help", "Bom dia"), it gets 56 right; the misses are two-word greetings that go to a close relative ("Goedemorgen" -> Limburgish, "Bom dia" -> Kabuverdianu), and confidence on short English phrases is often low even when the label is right.
  • Mixed-language text gets one label; romanized text (Hindi, Arabic, Russian written in Latin letters) is not a trained case.
  • Chinese Simplified vs Traditional is not distinguished.
Downloads last month
13
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Horizon-Labs/language-detection-small

Quantized
(279)
this model

Datasets used to train Horizon-Labs/language-detection-small

Space using Horizon-Labs/language-detection-small 1

Collection including Horizon-Labs/language-detection-small