Balochi Dependency Parser 🇧🇦
A from-scratch biaffine dependency parser for the Balochi language — an Iranian language spoken by ~8 million people. Word-level and character-level representations are learned jointly from a 27,000+ sentence treebank — no pretrained embeddings, no transfer learning — with joint POS tagging + dependency parsing, MST decoding, and an honest, calibrated confidence system.
This repo hosts the model weights in safetensors format (no pickle —
safe to load anywhere). The companion config.json + vocab.json let you
rebuild the exact architecture, and calibration.json provides the
Platt-scaled confidence coefficients.
Model details
| Property | Value |
|---|---|
| Architecture | Biaffine parser (Dozat & Manning 2017), from scratch |
| Encoder | 4-layer BiLSTM (hidden 300) over word (150-d) + char (64-d) embeddings |
| Vocab | 160,875 word types / 268 chars |
| Decoding | Chu-Liu-Edmonds MST with UD single-root constraint |
| Extras | 12-rule UD constraint layer, rule-based morphology, enhanced DEPS |
| Parameters | 33,811,073 (56 tensors) |
Accuracy (official test split, 2,788 sentences)
| Metric | Value |
|---|---|
| UAS | 94.31% |
| LAS | 91.98% (with UD constraints) |
| POS accuracy | 98.20% |
Confidence calibration: ECE 0.135 → 0.011 on test (Platt scaling). At a calibrated gate of ≥ 0.95, 99.74% of emitted parses are exact.
The treebank and model use Perso-Arabic script — Latin-script Balochi is not supported.
Usage
This is a custom architecture, not a transformers model — there is no
AutoModel. Load it with the project's own loader
(balochi_parser.checkpoint.load_checkpoint_safetensors), which reads
model.safetensors + config.json + vocab.json from this repo.
1. Clone the source repo and install
git clone <your-balochi-parser-repo-url> balochi-parser
cd balochi-parser/backend
pip install -e ".[dev,serving]"
pip install safetensors
2. Download the weights from this repo
cd balochi-parser/backend
# put the 4 files into a folder, e.g. hf_export/
# (or use huggingface_hub: huggingface-cli download <owner>/balochi-parser --local-dir new_model)
3. Load and parse
import torch
from balochi_parser.checkpoint import load_checkpoint_safetensors
from balochi_parser.predict import parse_sentence
from balochi_parser.tokenizer import pretokenize
model, vocab, config, _ = load_checkpoint_safetensors("new_model")
model.eval()
tokens = pretokenize("ائی مرد بازارءَ شت", vocab["word2idx"])
rows = parse_sentence(model, vocab, tokens, torch.device("cpu"))
for row in rows:
print(f"{row['id']}\t{row['form']}\t{row['upos']}\t{row['head']}\t{row['deprel']}")
Output (CoNLL-U-style):
1 ائی PRON 5 nsubj
2 مرد NOUN 5 obj
3 بازار NOUN 5 obl
4 ءَ ADP 3 case
5 شت VERB 0 root
Each token row also carries lemma, feats, deps, misc, and a per-token
confidence conf (P(arc) × P(label)). calibration.json sits next to the
weights, so confidence calibration is applied automatically.
CLI (from the source repo)
cd backend
python -m balochi_parser.predict --ckpt new_model/model.safetensors \
--text "ائی مرد بازارءَ شت" --show_conf
Expand the vocabulary without retraining
Dialect words absent from training get their own embedding at load time (vocab grows, existing embeddings stay bit-for-bit). Works on the safetensors export too:
model, vocab, config, _ = load_checkpoint_safetensors(
"new_model", extra_vocab="data/external_text/vocab_makrani.txt"
)
# or through the shared loader (accepts the file or the directory):
# load_checkpoint("new_model/model.safetensors", extra_vocab="my_words.txt")
# CLI — the --ckpt flag accepts the .safetensors file directly:
python -m balochi_parser.predict --ckpt new_model/model.safetensors \
--extra_vocab my_words.txt --text "تو کُجئیگ ئے" --show_conf
Files
| File | Description |
|---|---|
model.safetensors |
Model weights (56 tensors, 135 MB) |
config.json |
Architecture config + vocab sizes |
vocab.json |
word2idx / char2idx / upos2idx / deprel2idx |
calibration.json |
Platt-scaling coefficients (confidence calibration) |
README.md |
This file |
Why safetensors?
The original checkpoint (new_model/best_model.pt) is a torch.save
pickle. model.safetensors is the same weights in a pickle-free, size-safe
format — the exact export is verified bit-for-bit identical
(scripts/export_safetensors.py --verify).
Training data & citation
- Treebank: 27,824 gold sentences (22,250 train / 2,786 dev / 2,788 test), augmented to 91,170 for training (gold + pseudo-labeled + BNER).
- Sources: Balochi Academy texts (novels, folktales, proverbs, poetry, articles) + public Southern/Makrani sources, incl. 17,854 validated Makrani sentences and a 22,559-word Makrani wordlist.
If you use this model in research, please cite:
@software{balochi_parser2026,
title={Balochi Dependency Parser},
year={2026},
description={A from-scratch biaffine dependency parser for Balochi},
}
License
MIT — see the source repo's LICENSE.
Built with ❤️ for the Balochi language community.
- Downloads last month
- 9