NLLB-200 Sidaama LoRA Adapter

A LoRA fine-tuned adapter on top of facebook/nllb-200-distilled-600M for English → Sidaama (Sidamo) machine translation.

Sidaama (ISO 639-3: sid) is a Cushitic language spoken by over 10 million people in the Sidama Region of southern Ethiopia. It is a low-resource language with very limited NLP tooling, making this adapter a meaningful step toward broader language inclusion.


Model Details

  • Base model: facebook/nllb-200-distilled-600M
  • Fine-tuning method: LoRA (Low-Rank Adaptation) via PEFT
  • Task: Sequence-to-sequence machine translation (English → Sidaama)
  • Language pair: eng_Latn → gaz_Latn (Oromo language code used as a Cushitic proxy, as Sidaama is not natively supported in NLLB-200)
  • LoRA rank: 32
  • LoRA alpha: 64
  • LoRA dropout: 0.05
  • Target modules: q_proj, v_proj, k_proj, out_proj, fc1, fc2

Training Details

Dataset

  • Primary source: michsethowusu/english-sidaama_sentence-pairs_mt560
  • Augmentation: Additional sentence pairs were extracted and augmented from a Sidaama linguistic thesis, parsed and cleaned using custom scripts.
  • Final training data: ~49,600 sentence pairs after deduplication and quality filtering.

Data Quality Filtering

  • Removed duplicate pairs
  • Removed sentences with 0 or >100 tokens
  • Enforced English/Sidaama length ratio between 0.15 and 6.0

Data Splits

Split Proportion
Train 90%
Validation 5%
Test 5%

Hyperparameters

Parameter Value
Learning rate 3e-4
Batch size (per device) 4
Gradient accumulation steps 8
Effective batch size 32
Epochs 5
Weight decay 0.01
Precision fp16
Best model metric BLEU
Max sequence length 128 tokens

Hardware

  • Trained on an 8GB VRAM GPU with fp16 mixed precision.

Usage

Installation

pip install transformers peft torch sentencepiece

Inference

import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
from peft import PeftModel

BASE_MODEL = "facebook/nllb-200-distilled-600M"
ADAPTER    = "Anley-dev/nllb-sidaama-lora"

SRC_LANG = "eng_Latn"
TGT_LANG = "gaz_Latn"   # Cushitic proxy for Sidaama

device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer  = AutoTokenizer.from_pretrained(ADAPTER, src_lang=SRC_LANG, tgt_lang=TGT_LANG)
base_model = AutoModelForSeq2SeqLM.from_pretrained(
    BASE_MODEL,
    torch_dtype=torch.float16 if device == "cuda" else torch.float32
)
model = PeftModel.from_pretrained(base_model, ADAPTER).to(device)
model.eval()

def translate(text: str) -> str:
    inputs = tokenizer(text, return_tensors="pt").to(device)
    forced_bos_token_id = tokenizer.convert_tokens_to_ids(TGT_LANG)
    with torch.no_grad():
        tokens = model.generate(
            **inputs,
            forced_bos_token_id=forced_bos_token_id,
            max_length=128,
            num_beams=4,
            early_stopping=True
        )
    return tokenizer.decode(tokens[0], skip_special_tokens=True)

# Example
print(translate("Education is important for everyone."))
# → "Barsiissi hundaaf barbaachisaadha."

Example Translations

English Sidaama
Welcome to our community. Yannano ninkera babbaxitino.
Education is important for everyone. Barsiissi hundaaf barbaachisaadha.
We need to protect natural resources for future generations. Seera umamiinsama kalaqamanno minna-ilaminse daafira haaˈlate hasiiˈneemmo.

Evaluation Results

Evaluated on the held-out validation set (2,336 sentence pairs) using beam search (4 beams):

Metric Score
BLEU 16.98
chrF++ 43.97

These scores are competitive for a low-resource language with limited parallel data. chrF++ is the more reliable metric here as it operates at character level, which better captures Sidaama's rich morphology.


Limitations & Bias

  • Language code proxy: NLLB-200 does not natively support Sidaama (sid). This adapter uses the Oromo (gaz_Latn) language code as a Cushitic-family proxy. This may introduce subtle biases from Oromo linguistic patterns.
  • Low-resource setting: Despite augmentation, the training data is limited compared to high-resource languages. Output quality is best for short to medium sentences (≤30 words).
  • Domain bias: Training data is primarily sourced from religious and educational texts. Performance on other domains (legal, medical, casual speech) may vary.
  • Script: Sidaama uses the Latin script with special characters (e.g., ˈ, ˊ). Ensure your environment renders these correctly.

Citation

If you use this model in your work, please cite:

@misc{nllb-sidaama-lora,
  author    = {Anley},
  title     = {NLLB-200 Sidaama LoRA Adapter},
  year      = {2024},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/Anley-dev/nllb-sidaama-lora}
}

Framework Versions

  • PEFT: 0.20.0
  • Transformers: latest
  • PyTorch: latest
Downloads last month
28
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Anley-dev/nllb-sidaama-lora

Adapter
(163)
this model

Evaluation results

  • BLEU on english-sidaama_sentence-pairs_mt560
    self-reported
    16.980
  • chrF++ on english-sidaama_sentence-pairs_mt560
    self-reported
    43.970