|
Download README.md from failed09/bashkir-lid: direct link, hf CLI and curl.
- Browser
- Download file 7.71 kB
-
https://huggingface.co/failed09/bashkir-lid/resolve/main/README.md
- Command line
-
hf download hf://failed09/bashkir-lid/README.md
-
curl -L -o README.md https://huggingface.co/failed09/bashkir-lid/resolve/main/README.md
7.71 kB
| language: | |
| - ba | |
| license: apache-2.0 | |
| pretty_name: Bashkir LID (Binary BA vs NON_BA) | |
| library_name: onnxruntime | |
| pipeline_tag: text-classification | |
| tags: | |
| - bashkir | |
| - language-identification | |
| - text-classification | |
| - onnx | |
| - onnxruntime | |
| - low-resource | |
| # Bashkir LID (Binary BA vs NON_BA) | |
| > Compact ONNX language gate that decides whether text is Bashkir, for corpus | |
| > filtering and local CPU inference. | |
| ## Overview | |
| A compact binary language gate for Bashkir (`ba` vs `non_ba`), designed for | |
| corpus filtering, OCR triage and local applications that need lightweight CPU | |
| inference. It was trained with hard negatives from sister Turkic languages, | |
| Russian and technical/noisy text, with special attention to Bashkir/Tatar | |
| separation. | |
| | At a glance | | | |
| | --- | --- | | |
| | Task | Binary language identification (`ba` vs `non_ba`) | | |
| | Default artifact | `model.onnx` + `lid.py` | | |
| | Source | Bashkir, Turkic, Russian and technical text (not redistributed) | | |
| | Version / license | v3 / Apache-2.0 | | |
| ## Contents | |
| ### Files and Configurations | |
| | File | Purpose | | |
| | --- | --- | | |
| | `model.onnx` | Exported classifier graph and weights (opset 17) | | |
| | `vectorizer.json` | Dual-channel character and word feature vocabularies | | |
| | `lid.py` | Portable ONNX Runtime adapter | | |
| | `config.json` | Runtime contract | | |
| | `META.json` | Model passport, validation results and artifact hashes | | |
| | `__init__.py` | Package marker | | |
| | `LICENSE` | Full license text | | |
| | `SHA256SUMS` | Release checksums | | |
| Training texts and the original `joblib` pipeline are not distributed. | |
| ### Model Architecture | |
| | Property | Description | | |
| | --- | --- | | |
| | Task | Binary language identification: `ba` or `non_ba` | | |
| | Features | Dual-channel TF-IDF: `char_wb` (2–5) + `word` (1–2), 240,000 features | | |
| | Classifier | Linear SGD classifier exported to ONNX with Sigmoid normalization | | |
| | Runtime | ONNX Runtime on CPU | | |
| | Preprocessing | Unicode lowercase and dual-channel sparse features, provided by `lid.py` | | |
| | Outputs | Class labels (`ba`, `non_ba`) and probabilities | | |
| The ONNX graph takes sparse feature tensors (`char_ids`, `char_counts`, | |
| `word_ids`, `word_counts`); the adapter handles preprocessing, batching and | |
| checksum validation. No custom ONNX operators or Transformers installation are | |
| required. | |
| ### Examples | |
| | Text | Prediction | `ba` probability | | |
| | --- | --- | ---: | | |
| | `Мин башҡорт телен яратам.` | `ba` | 0.998 | | |
| | `Бары бала бара.` | `ba` | 0.927 | | |
| | `Сегодня хорошая погода.` | `non_ba` | 0.044 | | |
| | `SELECT id FROM users WHERE active = TRUE;` | `non_ba` | 0.011 | | |
| Note the second row: valid Bashkir text does not always contain | |
| Bashkir-specific letters. | |
| ## Method | |
| The model is a dual-channel linear classifier: TF-IDF `char_wb` (2–5) and word | |
| (1–2) features are fused and passed to an SGD classifier exported to ONNX. It was | |
| trained with hard negatives from sister Turkic languages, Russian and | |
| technical/noisy text. Training texts are not distributed in this repository. | |
| ```text | |
| training data (Bashkir + hard negative Turkic/Russian/noise) | |
| ├── 1. extract character n-grams: char_wb (2–5) | |
| ├── 2. extract word n-grams: words (1–2) | |
| ├── 3. feature union & TF-IDF vectorization | |
| ├── 4. train SGDClassifier with log-loss | |
| └── 5. export calibrated ONNX graph (fast single-pass classification) | |
| ``` | |
| ### Evaluation | |
| On a weak-label 4,000-sentence benchmark the model reached **93.25% binary | |
| accuracy**: at the default 0.50 threshold it retained 730/1,000 Bashkir examples | |
| and rejected 3,000/3,000 non-Bashkir examples. A hand-authored adversarial suite | |
| passed 26/26 (Tatar, Kazakh, Chuvash, Russian with Bashkir inserts, code and | |
| noise), and an independent 100-example stress set reached 97/100. Labels are weak | |
| or manually constructed, not an expert-reviewed Gold benchmark; exact numbers are | |
| recorded in [`META.json`](META.json). | |
| ### Operating thresholds | |
| `predict()` selects `ba` when its probability is at least 0.50. Measured | |
| operating points (from [`META.json`](META.json)): | |
| | BA probability threshold | BA recall | Non-BA false acceptance | | |
| | ---: | ---: | ---: | | |
| | 0.50 | 73.0% (730/1,000) | 0/3,000 | | |
| | 0.40 | 78.3% (783/1,000) | 2/3,000 | | |
| | 0.35 | 80.8% (808/1,000) | 4/3,000 | | |
| Use `0.50` for the strictest rejection of non-Bashkir text. Use `0.35–0.40` when | |
| preserving more Bashkir text matters, preferably with dictionary and quality | |
| checks. The price of the strict gate is false rejection of some valid Bashkir | |
| sentences, especially short or BA/Tatar-like sentences. | |
| ## Quality and Use | |
| This is a binary gate, not a general multilingual language detector. It does not | |
| identify which non-BA language was found. Tatar and other languages, short | |
| strings, names, OCR artifacts and mixed-language text can be ambiguous. | |
| ### Limitations | |
| - Binary decision only: `ba` vs `non_ba`, with no language label for rejects. | |
| - Weak-label benchmark; not an expert-reviewed Gold evaluation. | |
| - Confidence values are model scores, not calibrated human probabilities. | |
| - Empty strings, OCR fragments and technical noise should be handled with | |
| explicit pipeline rules. | |
| - For high-recall corpus construction, combine with the multiclass LID, | |
| dictionaries, quality rules and human review. | |
| ## Related Resources | |
| - [Bashkir Multiclass LID](https://huggingface.co/failed09/bashkir-lid-multiclass) — | |
| distinguishes `ba`, `tt`, `ru` and `other`; use it when you need the language | |
| label instead of a pass/reject gate. | |
| ## Usage | |
| ```bash | |
| pip install huggingface_hub onnxruntime numpy | |
| ``` | |
| ```python | |
| import sys | |
| from huggingface_hub import snapshot_download | |
| model_dir = snapshot_download( | |
| "failed09/bashkir-lid", | |
| allow_patterns=[ | |
| "lid.py", "config.json", "META.json", "model.onnx", "vectorizer.json" | |
| ], | |
| revision="main", # pin a reviewed commit for reproducible deployments | |
| ) | |
| sys.path.insert(0, model_dir) | |
| from lid import LanguageIdentifier | |
| lid = LanguageIdentifier(model_dir=model_dir) | |
| texts = ["Мин башҡорт телен яратам.", "Сегодня хорошая погода."] | |
| print(lid.predict(texts).tolist()) | |
| print(lid.predict_proba(texts)) | |
| print(lid.classes.tolist()) # probability-column order | |
| ``` | |
| After downloading, inference runs locally without network access. For a local | |
| checkout with the model files alongside `lid.py`, `LanguageIdentifier()` also | |
| works without extra arguments. | |
| ## License | |
| The code and model export are released under the | |
| [Apache-2.0 license](https://huggingface.co/failed09/bashkir-lid/blob/main/LICENSE). Training texts | |
| are not included in this repository. | |
| ## Citation | |
| ```bibtex | |
| @software{failed09_bashkir_lid_2026, | |
| title = {Bashkir LID}, | |
| author = {failed09}, | |
| year = {2026}, | |
| publisher = {Hugging Face}, | |
| url = {https://huggingface.co/failed09/bashkir-lid}, | |
| note = {Compact binary Bashkir language gate} | |
| } | |
| ``` | |
| ## Open Bashkir Data and Sources 🐝 | |
| This release is part of an open-source effort to support the development, | |
| preservation and practical use of the Bashkir language. Other related models, | |
| datasets and tools are available on the author's Hugging Face profile. | |
| The author does not claim ownership or authorship of the source texts or other | |
| materials used to derive this release; rights and licensing remain with the | |
| original authors, publishers and dataset providers. Source texts are not | |
| redistributed in this repository, so users should follow the licenses and | |
| attribution requirements of the relevant upstream resources. | |