masakhane/masakhanews
Viewer • Updated • 31.1k • 2.47k • 16
How to use Heroine2/lingala-news-topic-classifier with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-classification", model="Heroine2/lingala-news-topic-classifier") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("Heroine2/lingala-news-topic-classifier")
model = AutoModelForSequenceClassification.from_pretrained("Heroine2/lingala-news-topic-classifier", device_map="auto")Classifies Lingala news articles into four topics: politics, health, sports and business.
It is Davlan/afro-xlmr-base (Alabi et al., 2022) fine-tuned on the Lingala part of
MasakhaNEWS (Adelani et al., 2023).
Built for the ALU NLP and Language Technologies summative project.
| Model | Macro-F1 | Weighted-F1 | Accuracy |
|---|---|---|---|
| Majority class | 0.182 | 0.416 | 0.571 |
| TF-IDF (word + char) + linear SVM | 0.822 | 0.852 | 0.857 |
| BiLSTM + attention (Word2Vec init) | 0.859 ± 0.026 | 0.888 | 0.893 |
| Transformer trained from scratch | 0.824 ± 0.041 | 0.876 | 0.880 |
| This model (fine-tuned AfroXLMR) | 0.911 ± 0.014 | 0.929 | 0.930 |
Mean ± standard deviation over 3 training seeds. Per-class F1: politics 0.94, health 0.93, sports 0.98, business 0.79.
from transformers import pipeline
classifier = pipeline("text-classification", model="Heroine2/lingala-news-topic-classifier")
classifier("Ba Léopards babetami lisusu fimbu na CAN 2019")
Input format used in training: headline + ". " + article text, truncated to the first 512 sub-word tokens.
Base model
Davlan/afro-xlmr-base