Islamic3 / REFERENCES.md
Ghada-99-Ragab's picture
Upload 31 files
c08b36a verified
|
Raw History Blame Contribute Delete
9.36 kB

References & Data Provenance

Arabic version: REFERENCES_AR.md

This file documents what the project actually uses. Where the supplied materials do not record a fact (for example the upstream origin or licence of a data file), that is stated instead of guessed.

Project Specification

IslamicAIChallengePresentation.pdf: the project proposal for the challenge تحدي الذكاء الاصطناعي في خدمة المحتوى الإسلامي 2026 (12 slides, Arabic). It defines the workflow (detect, retrieve, verify with evidence, correct or refer to human review), the safety principles (source first, traceable evidence, abstain when unsure) and the expected deliverables. It also reports the previous IslamicEval 2025 research baselines (1A macro F1 0.9091 with CAMeLBERT-MSA; 1B accuracy 93.12%; 1C accuracy 70.39%). Those are historical figures, not results of this application.

IslamicEval 2025

  • Official source: Mubarak, H. et al. (2025). IslamicEval 2025: The First Shared Task of Capturing LLMs Hallucination in Islamic Content. Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, pp. 480–493, Suzhou, China. ACL Anthology ID 2025.arabicnlp-sharedtasks.67, DOI 10.18653/v1/2025.arabicnlp-sharedtasks.67. The task has two subtasks; this project uses only Subtask 1 (hallucination detection and correction of quoted ayahs and Hadith). Subtask 2 (question answering) is not used.
  • Subtasks represented: 1A (quotation span detection), 1B (verification: Correct / Incorrect), 1C (correction of incorrect quotations; خطأ marks "no source").
  • Files actually used: data/islamiceval_dev_subset.jsonl, 150 real development responses (50 each for 1A, 1B, 1C), with 194 / 247 / 179 annotated spans. It was taken from the project's unified dataset file (islamic_unified_dataset.jsonl), which the research notebook builds from the development files dev_SubtaskA/B/C (.xml + .tsv). Only the records with source = real_A, real_B, real_C were kept. The 1,500 synthetic records of the unified file are not in this evaluation subset: they are training data generated by the notebook from the Quran / Hadith corpora, not part of IslamicEval.
  • How it is used: evaluate.py reads the subset for the prototype evaluation; the optional training script reads the unified file. The running application reads neither.
  • Not included: the rest of the IslamicEval 2025 data (including training and test data).

Dataset Provenance

File Content What is verifiable
data/quran.json 6,236 ayahs, 114 surahs; fields surah_id, surah_name, ayah_id, ayah_text (with diacritics) Supplied by the project team as quranic_verses.json. Used unchanged.
data/hadith.json 34,994 records; fields hadithID, BookID (1–6), title, hadithTxt, Matn (present in 31,811 records) Supplied by the project team as six_hadith_books.json; only whitespace outside strings was removed (JSON re-serialised compactly), content unchanged.
data/islamiceval_dev_subset.jsonl See above Derived from the project's unified dataset file.
data/islamic_unified_dataset.jsonl 1,650 records: 150 real IslamicEval development responses + 1,500 synthetic ones generated by the research notebook Supplied by the project team; used only by the optional research/train_detector.py.
index/quran.idx.gz, index/hadith.idx.gz Pre-tokenised search indexes derived from data/quran.json and data/hadith.json by index_builder.py Generated; rebuild with python index_builder.py.

The research notebook uses quranic_verses.json and six_hadith_books.json as the reference database of its Subtask 1C scoring routine (copied there from the organisers' scoring.py). The supplied materials do not record where these two files were originally obtained or under which licence. The team should record the upstream source and licence here before redistributing the repository publicly.

Research References

  • Mubarak et al. (2025), cited above. This is the only paper the project relies on directly (task definition and subtask structure).

Models

  • The application uses no machine-learning model and loads no model weights. Detection is rule-based plus a corpus scan; retrieval is BM25 over pre-built indexes; verification is word / character alignment.
  • research/train_detector.py is optional research code (fine-tuning aubmindlab/bert-base-arabertv2 or CAMeLBERT-MSA for Subtask 1A). It is not used by the application and no trained weights exist in this repository.
  • Mode B calls a third-party language model chosen by the user (Google Gemini generateContent, OpenAI chat completions, or the Hugging Face router) with the user's own key. Default model names (gemini-2.5-flash, gpt-4o-mini, Qwen/Qwen2.5-72B-Instruct) are editable placeholders; only the Gemini name was checked against the provider's documentation. Provider request formats were tested against a local mock server, not against the live services.
  • The 1A baseline of 0.9091 in the proposal comes from earlier work with CAMeLBERT-MSA, not from this code.

Libraries

Library Purpose Source Version Licence
Python standard library (difflib, re, unicodedata, json, …) Matching, normalisation, data loading https://www.python.org Developed and tested with 3.12.3 Python Software Foundation License
Gradio Web interface (local / server runs only) https://github.com/gradio-app/gradio >=5.0,<6; not installed in the authoring environment (see below) Apache-2.0 (as published by the project)
RapidFuzz Optional faster Levenshtein similarity (pure-Python fallback exists) https://github.com/rapidfuzz/RapidFuzz >=3.0; not installed in the authoring environment MIT (as published by the project)
NumPy Character-level F1 in evaluate.py https://numpy.org 2.4.4 (tested) BSD-3-Clause (as published by the project)
Pyodide Runs the Python code inside a browser Web Worker for the Static Space demo (loaded from the jsDelivr CDN) https://pyodide.org 0.26.4 MPL-2.0 (as published by the project)

Licences above are those declared by the respective projects; confirm them at the source before relying on them.

Deployment

  • GitHub for the source code.
  • Hugging Face Static Space for the public demo: index.html starts a Web Worker that loads Pyodide 0.26.4 from the jsDelivr CDN and runs the project's Python modules in the visitor's browser (build_static_space.py assembles the folder). Gradio is used only for local runs, because Gradio and Docker Spaces require a paid Hugging Face plan.
  • No deployment has been performed from this repository yet, so no public URL exists in the documentation.

Licensing & Attribution

  • Source code: MIT (see LICENSE), applied to the code only.
  • Data files in data/ are third-party resources and are not covered by that licence. Their licences are not recorded in the supplied materials (see Dataset Provenance).
  • IslamicEval 2025 data belongs to the shared-task organisers; follow their terms and cite the overview paper.

Reproducibility

  • Python: 3.10 or newer; developed and tested with 3.12.3.
  • Requirements: requirements.txt. The core logic was tested with the Python standard library and NumPy only. The Gradio interface, RapidFuzz and Pyodide could not be installed or downloaded in the authoring environment (no network access). The rendering functions were tested natively, and the static page was exercised in a browser with a stubbed Pyodide; the real Gradio server and the real Pyodide runtime were not run there.
  • Data: the three files in data/, as described above.
  • Preprocessing: Arabic normalisation in retrieval.py (diacritic and tatweel removal, alef unification, punctuation stripping, and for matching ta marbuta / alef maqsura folding); the index is built in memory at start-up.
  • Execution: python index_builder.py (only after changing data/); python app.py; python evaluate.py; python -m unittest discover -s tests; python build_static_space.py for the browser demo.

Competition and benchmark links (not verified automatically)

The following were provided with the task and are cited as references. Automated access to the GitHub repository was blocked (robots rules) and the organiser pages and PDF were not read while preparing this version, so no rule, deadline or file format is quoted from them here:

Model references

  • CAMeLBERT-MSA (CAMeL Lab, CAMeL-Lab/bert-base-arabic-camelbert-msa) is the base model for the optional fine-tuned token classifier described in the README. The browser page runs it as a simulation unless HF_MODEL_ID is configured.
  • "Ask then Verify" uses the OpenAI Chat Completions API (gpt-4o-mini by default, set in web/config.js).