AraBERT for Arabic Question Answering (arabert-qa)

Khaaaleed5/arabert-qa is an extractive question answering model for Arabic. It is a fine-tuned version of AraBERT. Given a question and a context passage, the model predicts the span of text in the passage that answers the question.

Model Details

Developed by Khaleed (Khaaaleed5)
Model type Transformer encoder (BERT) with a span-prediction head (BertForQuestionAnswering)
Language Arabic (Modern Standard Arabic)
Base model aubmindlab/bert-base-arabertv02
Task Extractive Question Answering
License mit

Intended Uses

Direct use

  • Answering factual questions from a supplied Arabic passage.
  • A reader component in Arabic retrieval-augmented QA pipelines.
  • Research and educational demos of Arabic NLP.

Out-of-scope use

  • Generating answers that are not present in the context. This is an extractive model and cannot synthesize or reason beyond the passage.
  • Open-domain QA without a retriever.
  • High-stakes decisions (medical, legal, financial) without human verification.

How to Use

Pipeline

from transformers import pipeline

qa = pipeline(
    "question-answering",
    model="Khaaaleed5/arabert-qa",
    tokenizer="Khaaaleed5/arabert-qa",
)

context = "ุงู„ู‚ุงู‡ุฑุฉ ู‡ูŠ ุนุงุตู…ุฉ ุฌู…ู‡ูˆุฑูŠุฉ ู…ุตุฑ ุงู„ุนุฑุจูŠุฉ ูˆุฃูƒุจุฑ ู…ุฏู†ู‡ุงุŒ ูˆุชู‚ุน ุนู„ู‰ ุถูุงู ู†ู‡ุฑ ุงู„ู†ูŠู„."
question = "ู…ุง ู‡ูŠ ุนุงุตู…ุฉ ู…ุตุฑุŸ"

result = qa(question=question, context=context)
print(result)
# {'score': ..., 'start': ..., 'end': ..., 'answer': 'ุงู„ู‚ุงู‡ุฑุฉ'}

Manual inference

import torch
from transformers import AutoTokenizer, AutoModelForQuestionAnswering

model_id = "Khaaaleed5/arabert-qa"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForQuestionAnswering.from_pretrained(model_id)

inputs = tokenizer(question, context, return_tensors="pt", truncation=True, max_length=384)
with torch.no_grad():
    outputs = model(**inputs)

start = outputs.start_logits.argmax()
end = outputs.end_logits.argmax() + 1
answer = tokenizer.decode(inputs["input_ids"][0][start:end], skip_special_tokens=True)
print(answer)

Training Details

Training data

Training hyperparameters

Hyperparameter Value
Epochs 3
Learning rate 3e-05
train_batch_size 8
eval_batch_size 8
Max sequence length 384
Optimizer AdamW
lr_scheduler_type linear
mixed_precision_training Native AMP
Hardware 1ร— T4 GPU on Colab

Training results

Training Loss Epoch Step Validation Loss Start Accuracy End Accuracy Span Accuracy
1.2649 1.0 9664 1.3269 0.6315 0.6529 0.5322
1.0544 2.0 19328 1.3649 0.6349 0.6514 0.5332
0.7597 3.0 28992 1.5586 0.6310 0.6471 0.5283

Framework versions

  • Transformers 5.16.1
  • Pytorch 2.11.0+cu128
  • Datasets 4.8.5
  • Tokenizers 0.23.1
Downloads last month
43
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Khaaaleed5/arabert-qa

Finetuned
(4043)
this model

Dataset used to train Khaaaleed5/arabert-qa

Space using Khaaaleed5/arabert-qa 1