Instructions to use paruwka/LiteLID with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use paruwka/LiteLID with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("paruwka/LiteLID", "model.bin")) - Notebooks
- Google Colab
- Kaggle
Model Card for Model ID
This model is a reproduction of GlotLID on 125 languages using the Latn script, trained on the original GlotLID-C dataset for these languages, enriched by 1 million word-level examples per language. The word-level examples were obtained from splitting sentences from the dataset. It has also been trained with a bigger hashmap than GlotLID (2e6 instead of 1e6).
How to Get Started with the Model
import fasttext
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(repo_id="paruwka/LiteLID", filename="model.ftz", cache_dir=None)
model = fasttext.load_model(model_path)
model.predict(['predicting', 'language'], k=3) # this will return a tuple: (list of lists of top-k language labels, list of lists of their respective probabilities)
Model Details
model.ftz refers to v2, used throughout the paper. To try out other versions, simply append the version number to the filename.
versions:
v2: model.ftz or model_v2.ftz, trained on both words and sentences, with n-gram size 2-5, 2M hashmap buckets; the best-performing version to use with ILP4LID.
v0: model_v0.ftz, reproduces the training settings of GlotLID, using the Glot-C corpus of (Kargaran et al., 2024a) on a labelset reduced from 2102 to 125; performance near-identical to GlotLID on these labels.
v1: model_v1.ftz, the same settings, but n-gram size changed from 2-5 to 3-6, trained only on (up to) 1M words per language; slightly better performance on word-level scores, sentence-level prediction worsens.
Model Description
- Model type: fasttext architecture
- Language(s): ace, afr, als, ast, ayr, azj, bam, ban, bem, bjn, bug, cat, ceb, ces, cjk, crh, cym, dan, deu, dik, dyu, ekk, eng, epo, eus, ewe, fao, fij, fil, fin, fon, fra, fur, fuv, gaz, gla, gle, glg, gug, hat, hau, hin, hun, ibo, ilo, ind, isl, ita, jav, kab, kac, kam, kbp, kea, kik, kin, kmb, kmr, knc, kng, lij, lim, lin, lit, lmo, ltg, ltz, lua, lug, luo, lus, lvs, min, mlt, mos, mri, nld, nno, nob, npi, nso, nus, nya, oci, pag, pap, plt, pol, por, quy, ron, run, sag, scn, slk, slv, smo, sna, som, sot, spa, srd, ssw, sun, swe, swh, szl, taq, tpi, tsn, tso, tuk, tum, tur, twi, umb, uzn, vec, vie, war, wol, xho, yor, zsm, zul
- Developed by: Joanna Radoła
Training Hyperparameters
v2: lr=0.8, epochs=1, dim=256, minn=2, maxn=5, bucket=2000000, loss='softmax'
v0: lr=0.8, epochs=1, dim=256, minn=2, maxn=5, bucket=1000000, loss='softmax'
v1: lr=0.8, epochs=1, dim=256, minn=3, maxn=6, bucket=1000000, loss='softmax'
Usage with ILP4LID
...
Evaluation
Refer to (Radoła et al., 2026)
- Downloads last month
- -
Model tree for paruwka/LiteLID
Base model
cis-lmu/glotlid