Pāli → Sinhala Transformer (from scratch)

An encoder–decoder Transformer implemented from scratch in PyTorch — no nn.Transformer, no pretrained weights — for Pāli → Sinhala translation. Both languages are written in Sinhala script in this corpus, so a single shared subword vocabulary is used and the embeddings are tied across the encoder, decoder and output projection.

Results

Split BLEU chrF
Test 22.96 46.44
Validation 22.86 —

Metrics are sacreBLEU (13a tokenizer) and chrF, single reference.

Architecture

Layers 5 encoder + 5 decoder (Pre-LN)
d_model / d_ff / heads 512 / 2048 / 8
Activation GELU
Parameters ~38.8M
Tokenizer Shared SentencePiece unigram, 4000 vocab
Embeddings Tied (encoder ⊗ decoder ⊗ output)

Training

Trained for up to 300 epochs on ~16.5k parallel pairs with:

  • label smoothing 0.2, residual dropout 0.3
  • source word-dropout 0.1
  • subword regularisation (unigram sampling, α=0.1) — a fresh segmentation every epoch, acting as data augmentation
  • token-bucketed batching (~3000 tokens × 2 accumulation)
  • AdamW, linear warmup (1000 steps) + inverse-sqrt decay, peak LR 0.0007
  • early stopping on validation BLEU, and top-5 checkpoint averaging

Decoding uses beam search (beam 8) with a GNMT length penalty α=0.6, both selected on the validation set.

Data preprocessing

NFC Unicode normalisation → zero-width character stripping → whitespace squeezing → exact duplicate-pair removal (which also eliminates train/test leakage) → length window 1–80 words → length-ratio cap 4.0. Split 80/10/10.

Files

  • model_weights.pt — state dict (averaged checkpoint)
  • spm.model / spm.vocab — SentencePiece tokenizer
  • metadata.json — architecture + decoding config
  • test_results.csv — every test prediction with its sentence BLEU
  • figures: training report, attention map, tokenization analysis, sample translations

Usage

The model classes are not packaged as a transformers architecture — see the accompanying Space for a self-contained app.py that rebuilds the model and loads this state dict: https://huggingface.co/spaces/NethranjaliSE/Pali_Sinhala_Translator_From_Scratch

Limitations

Trained on a small, domain-specific corpus of canonical Buddhist text. Output is frequently fluent but not always faithful — the model can produce a well-formed Sinhala sentence that does not correspond to the Pāli source. It is a research prototype and should not be relied on for doctrinal or scholarly interpretation. Many Pāli lines have several valid Sinhala renderings, so single-reference BLEU understates quality somewhat.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using NethranjaliSE/Pali_Sinhala_From_Scratch 1