# References & Data Provenance Arabic version: [REFERENCES_AR.md](REFERENCES_AR.md) This file documents what the project actually uses. Where the supplied materials do not record a fact (for example the upstream origin or licence of a data file), that is stated instead of guessed. ## Project Specification `IslamicAIChallengePresentation.pdf`: the project proposal for the challenge *تحدي الذكاء الاصطناعي في خدمة المحتوى الإسلامي 2026* (12 slides, Arabic). It defines the workflow (detect, retrieve, verify with evidence, correct or refer to human review), the safety principles (source first, traceable evidence, abstain when unsure) and the expected deliverables. It also reports the previous IslamicEval 2025 research baselines (1A macro F1 0.9091 with CAMeLBERT-MSA; 1B accuracy 93.12%; 1C accuracy 70.39%). Those are historical figures, not results of this application. ## IslamicEval 2025 * **Official source:** Mubarak, H. et al. (2025). *IslamicEval 2025: The First Shared Task of Capturing LLMs Hallucination in Islamic Content.* Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, pp. 480–493, Suzhou, China. ACL Anthology ID `2025.arabicnlp-sharedtasks.67`, DOI `10.18653/v1/2025.arabicnlp-sharedtasks.67`. The task has two subtasks; this project uses only Subtask 1 (hallucination detection and correction of quoted ayahs and Hadith). Subtask 2 (question answering) is not used. * **Subtasks represented:** 1A (quotation span detection), 1B (verification: Correct / Incorrect), 1C (correction of incorrect quotations; `خطأ` marks "no source"). * **Files actually used:** `data/islamiceval_dev_subset.jsonl`, 150 real development responses (50 each for 1A, 1B, 1C), with 194 / 247 / 179 annotated spans. It was taken from the project's unified dataset file (`islamic_unified_dataset.jsonl`), which the research notebook builds from the development files `dev_SubtaskA/B/C` (`.xml` + `.tsv`). Only the records with `source` = `real_A`, `real_B`, `real_C` were kept. The 1,500 synthetic records of the unified file are not in this evaluation subset: they are training data generated by the notebook from the Quran / Hadith corpora, not part of IslamicEval. * **How it is used:** `evaluate.py` reads the subset for the prototype evaluation; the optional training script reads the unified file. The running application reads neither. * **Not included:** the rest of the IslamicEval 2025 data (including training and test data). ## Dataset Provenance | File | Content | What is verifiable | |---|---|---| | `data/quran.json` | 6,236 ayahs, 114 surahs; fields `surah_id`, `surah_name`, `ayah_id`, `ayah_text` (with diacritics) | Supplied by the project team as `quranic_verses.json`. Used unchanged. | | `data/hadith.json` | 34,994 records; fields `hadithID`, `BookID` (1–6), `title`, `hadithTxt`, `Matn` (present in 31,811 records) | Supplied by the project team as `six_hadith_books.json`; only whitespace outside strings was removed (JSON re-serialised compactly), content unchanged. | | `data/islamiceval_dev_subset.jsonl` | See above | Derived from the project's unified dataset file. | | `data/islamic_unified_dataset.jsonl` | 1,650 records: 150 real IslamicEval development responses + 1,500 synthetic ones generated by the research notebook | Supplied by the project team; used only by the optional `research/train_detector.py`. | | `index/quran.idx.gz`, `index/hadith.idx.gz` | Pre-tokenised search indexes derived from `data/quran.json` and `data/hadith.json` by `index_builder.py` | Generated; rebuild with `python index_builder.py`. | The research notebook uses `quranic_verses.json` and `six_hadith_books.json` as the reference database of its Subtask 1C scoring routine (copied there from the organisers' `scoring.py`). The supplied materials do **not** record where these two files were originally obtained or under which licence. The team should record the upstream source and licence here before redistributing the repository publicly. ## Research References * Mubarak et al. (2025), cited above. This is the only paper the project relies on directly (task definition and subtask structure). ## Models * **The application uses no machine-learning model and loads no model weights.** Detection is rule-based plus a corpus scan; retrieval is BM25 over pre-built indexes; verification is word / character alignment. * `research/train_detector.py` is optional research code (fine-tuning `aubmindlab/bert-base-arabertv2` or CAMeLBERT-MSA for Subtask 1A). It is not used by the application and no trained weights exist in this repository. * **Mode B** calls a third-party language model chosen by the user (Google Gemini `generateContent`, OpenAI chat completions, or the Hugging Face router) with the user's own key. Default model names (`gemini-2.5-flash`, `gpt-4o-mini`, `Qwen/Qwen2.5-72B-Instruct`) are editable placeholders; only the Gemini name was checked against the provider's documentation. Provider request formats were tested against a local mock server, not against the live services. * The 1A baseline of 0.9091 in the proposal comes from earlier work with CAMeLBERT-MSA, not from this code. ## Libraries | Library | Purpose | Source | Version | Licence | |---|---|---|---|---| | Python standard library (`difflib`, `re`, `unicodedata`, `json`, …) | Matching, normalisation, data loading | https://www.python.org | Developed and tested with 3.12.3 | Python Software Foundation License | | Gradio | Web interface (local / server runs only) | https://github.com/gradio-app/gradio | `>=5.0,<6`; not installed in the authoring environment (see below) | Apache-2.0 (as published by the project) | | RapidFuzz | Optional faster Levenshtein similarity (pure-Python fallback exists) | https://github.com/rapidfuzz/RapidFuzz | `>=3.0`; not installed in the authoring environment | MIT (as published by the project) | | NumPy | Character-level F1 in `evaluate.py` | https://numpy.org | 2.4.4 (tested) | BSD-3-Clause (as published by the project) | | Pyodide | Runs the Python code inside a browser Web Worker for the Static Space demo (loaded from the jsDelivr CDN) | https://pyodide.org | 0.26.4 | MPL-2.0 (as published by the project) | Licences above are those declared by the respective projects; confirm them at the source before relying on them. ## Deployment * GitHub for the source code. * Hugging Face **Static** Space for the public demo: `index.html` starts a Web Worker that loads Pyodide 0.26.4 from the jsDelivr CDN and runs the project's Python modules in the visitor's browser (`build_static_space.py` assembles the folder). Gradio is used only for local runs, because Gradio and Docker Spaces require a paid Hugging Face plan. * No deployment has been performed from this repository yet, so no public URL exists in the documentation. ## Licensing & Attribution * Source code: MIT (see `LICENSE`), applied to the code only. * Data files in `data/` are third-party resources and are not covered by that licence. Their licences are not recorded in the supplied materials (see Dataset Provenance). * IslamicEval 2025 data belongs to the shared-task organisers; follow their terms and cite the overview paper. ## Reproducibility * **Python:** 3.10 or newer; developed and tested with 3.12.3. * **Requirements:** `requirements.txt`. The core logic was tested with the Python standard library and NumPy only. The Gradio interface, RapidFuzz and Pyodide could not be installed or downloaded in the authoring environment (no network access). The rendering functions were tested natively, and the static page was exercised in a browser with a stubbed Pyodide; the real Gradio server and the real Pyodide runtime were not run there. * **Data:** the three files in `data/`, as described above. * **Preprocessing:** Arabic normalisation in `retrieval.py` (diacritic and tatweel removal, alef unification, punctuation stripping, and for matching ta marbuta / alef maqsura folding); the index is built in memory at start-up. * **Execution:** `python index_builder.py` (only after changing `data/`); `python app.py`; `python evaluate.py`; `python -m unittest discover -s tests`; `python build_static_space.py` for the browser demo. ## Competition and benchmark links (not verified automatically) The following were provided with the task and are cited as references. Automated access to the GitHub repository was blocked (robots rules) and the organiser pages and PDF were not read while preparing this version, so no rule, deadline or file format is quoted from them here: * IslamicEval 2025 task site: * Subtask 1 code repository: * Challenge site and timeline: * Hackathon rules (PDF): ## Model references * CAMeLBERT-MSA (CAMeL Lab, `CAMeL-Lab/bert-base-arabic-camelbert-msa`) is the base model for the optional fine-tuned token classifier described in the README. The browser page runs it as a simulation unless `HF_MODEL_ID` is configured. * "Ask then Verify" uses the OpenAI Chat Completions API (`gpt-4o-mini` by default, set in `web/config.js`).