Instructions to use Anley-dev/nllb-sidaama-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Anley-dev/nllb-sidaama-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForSeq2SeqLM base_model = AutoModelForSeq2SeqLM.from_pretrained("facebook/nllb-200-distilled-600M") model = PeftModel.from_pretrained(base_model, "Anley-dev/nllb-sidaama-lora") - Transformers
How to use Anley-dev/nllb-sidaama-lora with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="Anley-dev/nllb-sidaama-lora")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Anley-dev/nllb-sidaama-lora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
NLLB-200 Sidaama LoRA Adapter
A LoRA fine-tuned adapter on top of facebook/nllb-200-distilled-600M for English → Sidaama (Sidamo) machine translation.
Sidaama (ISO 639-3: sid) is a Cushitic language spoken by over 10 million people in the Sidama Region of southern Ethiopia. It is a low-resource language with very limited NLP tooling, making this adapter a meaningful step toward broader language inclusion.
Model Details
- Base model:
facebook/nllb-200-distilled-600M - Fine-tuning method: LoRA (Low-Rank Adaptation) via PEFT
- Task: Sequence-to-sequence machine translation (English → Sidaama)
- Language pair:
eng_Latn→gaz_Latn(Oromo language code used as a Cushitic proxy, as Sidaama is not natively supported in NLLB-200) - LoRA rank: 32
- LoRA alpha: 64
- LoRA dropout: 0.05
- Target modules:
q_proj,v_proj,k_proj,out_proj,fc1,fc2
Training Details
Dataset
- Primary source:
michsethowusu/english-sidaama_sentence-pairs_mt560 - Augmentation: Additional sentence pairs were extracted and augmented from a Sidaama linguistic thesis, parsed and cleaned using custom scripts.
- Final training data: ~49,600 sentence pairs after deduplication and quality filtering.
Data Quality Filtering
- Removed duplicate pairs
- Removed sentences with 0 or >100 tokens
- Enforced English/Sidaama length ratio between 0.15 and 6.0
Data Splits
| Split | Proportion |
|---|---|
| Train | 90% |
| Validation | 5% |
| Test | 5% |
Hyperparameters
| Parameter | Value |
|---|---|
| Learning rate | 3e-4 |
| Batch size (per device) | 4 |
| Gradient accumulation steps | 8 |
| Effective batch size | 32 |
| Epochs | 5 |
| Weight decay | 0.01 |
| Precision | fp16 |
| Best model metric | BLEU |
| Max sequence length | 128 tokens |
Hardware
- Trained on an 8GB VRAM GPU with fp16 mixed precision.
Usage
Installation
pip install transformers peft torch sentencepiece
Inference
import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
from peft import PeftModel
BASE_MODEL = "facebook/nllb-200-distilled-600M"
ADAPTER = "Anley-dev/nllb-sidaama-lora"
SRC_LANG = "eng_Latn"
TGT_LANG = "gaz_Latn" # Cushitic proxy for Sidaama
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = AutoTokenizer.from_pretrained(ADAPTER, src_lang=SRC_LANG, tgt_lang=TGT_LANG)
base_model = AutoModelForSeq2SeqLM.from_pretrained(
BASE_MODEL,
torch_dtype=torch.float16 if device == "cuda" else torch.float32
)
model = PeftModel.from_pretrained(base_model, ADAPTER).to(device)
model.eval()
def translate(text: str) -> str:
inputs = tokenizer(text, return_tensors="pt").to(device)
forced_bos_token_id = tokenizer.convert_tokens_to_ids(TGT_LANG)
with torch.no_grad():
tokens = model.generate(
**inputs,
forced_bos_token_id=forced_bos_token_id,
max_length=128,
num_beams=4,
early_stopping=True
)
return tokenizer.decode(tokens[0], skip_special_tokens=True)
# Example
print(translate("Education is important for everyone."))
# → "Barsiissi hundaaf barbaachisaadha."
Example Translations
| English | Sidaama |
|---|---|
| Welcome to our community. | Yannano ninkera babbaxitino. |
| Education is important for everyone. | Barsiissi hundaaf barbaachisaadha. |
| We need to protect natural resources for future generations. | Seera umamiinsama kalaqamanno minna-ilaminse daafira haaˈlate hasiiˈneemmo. |
Evaluation Results
Evaluated on the held-out validation set (2,336 sentence pairs) using beam search (4 beams):
| Metric | Score |
|---|---|
| BLEU | 16.98 |
| chrF++ | 43.97 |
These scores are competitive for a low-resource language with limited parallel data. chrF++ is the more reliable metric here as it operates at character level, which better captures Sidaama's rich morphology.
Limitations & Bias
- Language code proxy: NLLB-200 does not natively support Sidaama (
sid). This adapter uses the Oromo (gaz_Latn) language code as a Cushitic-family proxy. This may introduce subtle biases from Oromo linguistic patterns. - Low-resource setting: Despite augmentation, the training data is limited compared to high-resource languages. Output quality is best for short to medium sentences (≤30 words).
- Domain bias: Training data is primarily sourced from religious and educational texts. Performance on other domains (legal, medical, casual speech) may vary.
- Script: Sidaama uses the Latin script with special characters (e.g.,
ˈ,ˊ). Ensure your environment renders these correctly.
Citation
If you use this model in your work, please cite:
@misc{nllb-sidaama-lora,
author = {Anley},
title = {NLLB-200 Sidaama LoRA Adapter},
year = {2024},
publisher = {Hugging Face},
url = {https://huggingface.co/Anley-dev/nllb-sidaama-lora}
}
Framework Versions
- PEFT: 0.20.0
- Transformers: latest
- PyTorch: latest
- Downloads last month
- 28
Model tree for Anley-dev/nllb-sidaama-lora
Base model
facebook/nllb-200-distilled-600MEvaluation results
- BLEU on english-sidaama_sentence-pairs_mt560self-reported16.980
- chrF++ on english-sidaama_sentence-pairs_mt560self-reported43.970