abdoelsayed/ArabicaQA
Updated โข 165 โข 6
How to use Khaaaleed5/arabert-qa with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("question-answering", model="Khaaaleed5/arabert-qa") # Load model directly
from transformers import AutoTokenizer, AutoModelForQuestionAnswering
tokenizer = AutoTokenizer.from_pretrained("Khaaaleed5/arabert-qa")
model = AutoModelForQuestionAnswering.from_pretrained("Khaaaleed5/arabert-qa", device_map="auto")arabert-qa)
Khaaaleed5/arabert-qa is an extractive question answering model for Arabic. It is a fine-tuned version of AraBERT. Given a question and a context passage, the model predicts the span of text in the passage that answers the question.
| Developed by | Khaleed (Khaaaleed5) |
| Model type | Transformer encoder (BERT) with a span-prediction head (BertForQuestionAnswering) |
| Language | Arabic (Modern Standard Arabic) |
| Base model | aubmindlab/bert-base-arabertv02 |
| Task | Extractive Question Answering |
| License | mit |
Direct use
Out-of-scope use
from transformers import pipeline
qa = pipeline(
"question-answering",
model="Khaaaleed5/arabert-qa",
tokenizer="Khaaaleed5/arabert-qa",
)
context = "ุงููุงูุฑุฉ ูู ุนุงุตู
ุฉ ุฌู
ููุฑูุฉ ู
ุตุฑ ุงูุนุฑุจูุฉ ูุฃูุจุฑ ู
ุฏููุงุ ูุชูุน ุนูู ุถูุงู ููุฑ ุงูููู."
question = "ู
ุง ูู ุนุงุตู
ุฉ ู
ุตุฑุ"
result = qa(question=question, context=context)
print(result)
# {'score': ..., 'start': ..., 'end': ..., 'answer': 'ุงููุงูุฑุฉ'}
import torch
from transformers import AutoTokenizer, AutoModelForQuestionAnswering
model_id = "Khaaaleed5/arabert-qa"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForQuestionAnswering.from_pretrained(model_id)
inputs = tokenizer(question, context, return_tensors="pt", truncation=True, max_length=384)
with torch.no_grad():
outputs = model(**inputs)
start = outputs.start_logits.argmax()
end = outputs.end_logits.argmax() + 1
answer = tokenizer.decode(inputs["input_ids"][0][start:end], skip_special_tokens=True)
print(answer)
| Hyperparameter | Value |
|---|---|
| Epochs | 3 |
| Learning rate | 3e-05 |
| train_batch_size | 8 |
| eval_batch_size | 8 |
| Max sequence length | 384 |
| Optimizer | AdamW |
| lr_scheduler_type | linear |
| mixed_precision_training | Native AMP |
| Hardware | 1ร T4 GPU on Colab |
| Training Loss | Epoch | Step | Validation Loss | Start Accuracy | End Accuracy | Span Accuracy |
|---|---|---|---|---|---|---|
| 1.2649 | 1.0 | 9664 | 1.3269 | 0.6315 | 0.6529 | 0.5322 |
| 1.0544 | 2.0 | 19328 | 1.3649 | 0.6349 | 0.6514 | 0.5332 |
| 0.7597 | 3.0 | 28992 | 1.5586 | 0.6310 | 0.6471 | 0.5283 |
Base model
aubmindlab/bert-base-arabertv02