Model Card for Model ID

This model is a reproduction of GlotLID on 125 languages using the Latn script, trained on the original GlotLID-C dataset for these languages, enriched by 1 million word-level examples per language. The word-level examples were obtained from splitting sentences from the dataset. It has also been trained with a bigger hashmap than GlotLID (2e6 instead of 1e6).

How to Get Started with the Model

import fasttext
from huggingface_hub import hf_hub_download

model_path = hf_hub_download(repo_id="paruwka/LiteLID", filename="model.ftz", cache_dir=None)
model = fasttext.load_model(model_path)
model.predict(['predicting', 'language'], k=3) # this will return a tuple:  (list of lists of top-k language labels, list of lists of their respective probabilities)

Model Details

model.ftz refers to v2, used throughout the paper. To try out other versions, simply append the version number to the filename.

versions:

v2: model.ftz or model_v2.ftz, trained on both words and sentences, with n-gram size 2-5, 2M hashmap buckets; the best-performing version to use with ILP4LID.

v0: model_v0.ftz, reproduces the training settings of GlotLID, using the Glot-C corpus of (Kargaran et al., 2024a) on a labelset reduced from 2102 to 125; performance near-identical to GlotLID on these labels.

v1: model_v1.ftz, the same settings, but n-gram size changed from 2-5 to 3-6, trained only on (up to) 1M words per language; slightly better performance on word-level scores, sentence-level prediction worsens.

Model Description

  • Model type: fasttext architecture
  • Language(s): ace, afr, als, ast, ayr, azj, bam, ban, bem, bjn, bug, cat, ceb, ces, cjk, crh, cym, dan, deu, dik, dyu, ekk, eng, epo, eus, ewe, fao, fij, fil, fin, fon, fra, fur, fuv, gaz, gla, gle, glg, gug, hat, hau, hin, hun, ibo, ilo, ind, isl, ita, jav, kab, kac, kam, kbp, kea, kik, kin, kmb, kmr, knc, kng, lij, lim, lin, lit, lmo, ltg, ltz, lua, lug, luo, lus, lvs, min, mlt, mos, mri, nld, nno, nob, npi, nso, nus, nya, oci, pag, pap, plt, pol, por, quy, ron, run, sag, scn, slk, slv, smo, sna, som, sot, spa, srd, ssw, sun, swe, swh, szl, taq, tpi, tsn, tso, tuk, tum, tur, twi, umb, uzn, vec, vie, war, wol, xho, yor, zsm, zul
  • Developed by: Joanna Radoła

Training Hyperparameters

v2: lr=0.8, epochs=1, dim=256, minn=2, maxn=5, bucket=2000000, loss='softmax'

v0: lr=0.8, epochs=1, dim=256, minn=2, maxn=5, bucket=1000000, loss='softmax'

v1: lr=0.8, epochs=1, dim=256, minn=3, maxn=6, bucket=1000000, loss='softmax'

Usage with ILP4LID

...

Evaluation

Refer to (Radoła et al., 2026)

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for paruwka/LiteLID

Base model

cis-lmu/glotlid
Finetuned
(2)
this model

Dataset used to train paruwka/LiteLID