Balochi Dependency Parser 🇧🇦

A from-scratch biaffine dependency parser for the Balochi language — an Iranian language spoken by ~8 million people. Word-level and character-level representations are learned jointly from a 27,000+ sentence treebank — no pretrained embeddings, no transfer learning — with joint POS tagging + dependency parsing, MST decoding, and an honest, calibrated confidence system.

This repo hosts the model weights in safetensors format (no pickle — safe to load anywhere). The companion config.json + vocab.json let you rebuild the exact architecture, and calibration.json provides the Platt-scaled confidence coefficients.

Model details

Property Value
Architecture Biaffine parser (Dozat & Manning 2017), from scratch
Encoder 4-layer BiLSTM (hidden 300) over word (150-d) + char (64-d) embeddings
Vocab 160,875 word types / 268 chars
Decoding Chu-Liu-Edmonds MST with UD single-root constraint
Extras 12-rule UD constraint layer, rule-based morphology, enhanced DEPS
Parameters 33,811,073 (56 tensors)

Accuracy (official test split, 2,788 sentences)

Metric Value
UAS 94.31%
LAS 91.98% (with UD constraints)
POS accuracy 98.20%

Confidence calibration: ECE 0.135 → 0.011 on test (Platt scaling). At a calibrated gate of ≥ 0.95, 99.74% of emitted parses are exact.

The treebank and model use Perso-Arabic script — Latin-script Balochi is not supported.

Usage

This is a custom architecture, not a transformers model — there is no AutoModel. Load it with the project's own loader (balochi_parser.checkpoint.load_checkpoint_safetensors), which reads model.safetensors + config.json + vocab.json from this repo.

1. Clone the source repo and install

git clone <your-balochi-parser-repo-url> balochi-parser
cd balochi-parser/backend
pip install -e ".[dev,serving]"
pip install safetensors

2. Download the weights from this repo

cd balochi-parser/backend
# put the 4 files into a folder, e.g. hf_export/
# (or use huggingface_hub: huggingface-cli download <owner>/balochi-parser --local-dir new_model)

3. Load and parse

import torch
from balochi_parser.checkpoint import load_checkpoint_safetensors
from balochi_parser.predict import parse_sentence
from balochi_parser.tokenizer import pretokenize

model, vocab, config, _ = load_checkpoint_safetensors("new_model")
model.eval()

tokens = pretokenize("ائی مرد بازارءَ شت", vocab["word2idx"])
rows = parse_sentence(model, vocab, tokens, torch.device("cpu"))
for row in rows:
    print(f"{row['id']}\t{row['form']}\t{row['upos']}\t{row['head']}\t{row['deprel']}")

Output (CoNLL-U-style):

1	ائی	PRON	5	nsubj
2	مرد	NOUN	5	obj
3	بازار	NOUN	5	obl
4	ءَ	ADP	3	case
5	شت	VERB	0	root

Each token row also carries lemma, feats, deps, misc, and a per-token confidence conf (P(arc) × P(label)). calibration.json sits next to the weights, so confidence calibration is applied automatically.

CLI (from the source repo)

cd backend
python -m balochi_parser.predict --ckpt new_model/model.safetensors \
    --text "ائی مرد بازارءَ شت" --show_conf

Expand the vocabulary without retraining

Dialect words absent from training get their own embedding at load time (vocab grows, existing embeddings stay bit-for-bit). Works on the safetensors export too:

model, vocab, config, _ = load_checkpoint_safetensors(
    "new_model", extra_vocab="data/external_text/vocab_makrani.txt"
)
# or through the shared loader (accepts the file or the directory):
#   load_checkpoint("new_model/model.safetensors", extra_vocab="my_words.txt")
# CLI — the --ckpt flag accepts the .safetensors file directly:
python -m balochi_parser.predict --ckpt new_model/model.safetensors \
    --extra_vocab my_words.txt --text "تو کُجئیگ ئے" --show_conf

Files

File Description
model.safetensors Model weights (56 tensors, 135 MB)
config.json Architecture config + vocab sizes
vocab.json word2idx / char2idx / upos2idx / deprel2idx
calibration.json Platt-scaling coefficients (confidence calibration)
README.md This file

Why safetensors?

The original checkpoint (new_model/best_model.pt) is a torch.save pickle. model.safetensors is the same weights in a pickle-free, size-safe format — the exact export is verified bit-for-bit identical (scripts/export_safetensors.py --verify).

Training data & citation

  • Treebank: 27,824 gold sentences (22,250 train / 2,786 dev / 2,788 test), augmented to 91,170 for training (gold + pseudo-labeled + BNER).
  • Sources: Balochi Academy texts (novels, folktales, proverbs, poetry, articles) + public Southern/Makrani sources, incl. 17,854 validated Makrani sentences and a 22,559-word Makrani wordlist.

If you use this model in research, please cite:

@software{balochi_parser2026,
  title={Balochi Dependency Parser},
  year={2026},
  description={A from-scratch biaffine dependency parser for Balochi},
}

License

MIT — see the source repo's LICENSE.


Built with ❤️ for the Balochi language community.

Downloads last month
9
Safetensors
Model size
33.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support