Islamic4 / README.md
Ghada-99-Ragab's picture
Upload 31 files
0390c03 verified
|
Raw History Blame Contribute Delete
10.6 kB
---
title: التحقق من هلوسة القرآن والحديث وتصحيحها
emoji: 🕌
colorFrom: green
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: التحقق من هلوسة القرآن والحديث وتصحيحها
---
# التحقق من هلوسة القرآن والحديث وتصحيحها · Quran & Hadith Hallucination Verifier and Corrector
Arabic version: [README_AR.md](README_AR.md)
Detects Quran and Hadith quotations in any text (for example answers written by language models), verifies each one
against the sources, shows the evidence, and proposes a correction **taken verbatim from the source text**. When the
evidence is not sufficient the tool says so and refers the case to human review. The user interface is Arabic.
## Live Demo
**[LIVE_DEPLOYMENT_URL]** <!-- replace with the public Hugging Face Space URL after deployment -->
## Features
* **Direct verification:** paste any paragraph; every Quran / Hadith quotation is found, compared word by word and
highlighted, with the exact source text offered as the correction.
* **Ask then Verify:** type a question; a pre-configured ChatGPT (OpenAI) client answers and the same engine verifies the
answer and returns a corrected version. There is no model selection in this tab.
* **No trigger phrases required:** quotations are recognised by matching the text with the Quran and Hadith corpora, with or
without quotation marks and with or without openers such as "قال الله تعالى" or "قال رسول الله ﷺ". Such phrases are only a
weak fallback for altered quotations that the corpus cannot confirm.
* **Detection model selector (English):** *Standard · Rules + Corpus*, or *CAMeLBERT-MSA · Fine-tuned*. See below.
* **Benchmark export:** JSON and TSV in the IslamicEval 1A / 1B / 1C structure (see [Benchmark alignment](#benchmark-alignment)).
* **Copy corrected text:** a working clipboard button with a toast message ("تم نسخ النص بنجاح"), with a fallback for iframes.
* **Fast, informative loading:** the Quran index starts downloading before the scripts run, the Hadith index loads in the
background, files are cached by the browser (Cache API), verification results and demo examples are cached as JSON in
`localStorage`, and animated skeleton cards plus a three-step progress bar show what is happening.
## CAMeLBERT-MSA support
`web/camelbert.js` exposes one function, `detectSpans`, that returns BIO-tagged tokens (`B-Ayah`, `I-Ayah`, `B-Hadith`,
`I-Hadith`, `O`) and spans:
* **Live mode:** set `HF_MODEL_ID` (and `HF_TOKEN` for a private model) in `web/config.js`. The page calls the Hugging Face
Inference API for a fine-tuned token-classification model. The spans are then verified and corrected by the local pipeline.
* **Simulation mode (default):** with no model ID set, **no model weights are loaded**. The spans come from the standard
detector and are shown as BIO tags with deterministic pseudo-confidence values. The interface labels this clearly as
*Simulation*. Use it to demonstrate the flow, not to report model accuracy.
## Workflow
```
text -> detection (quotation marks + corpus match, then corpus scan for unannounced quotes) -> BM25 retrieval
-> span-local alignment (phonetic skeleton, sliding window, LCS, dynamic gap)
-> exact: verified | altered Quran passage: source correction | uncertain: human review | nothing similar: no source
```
| Decision | Colour | Meaning |
|---|---|---|
| موثّق (verified) | green | Equals its source after ignoring spelling and diacritic conventions. |
| غير مطابق (mismatch) | red | Differs from the source (Quran: the exact source text is offered) or no source resembles it. |
| يحتاج مراجعة بشرية (human review) | amber | Evidence is insufficient or ambiguous. No correction is produced. |
## Benchmark alignment
`benchmark.py` converts a pipeline result to the structure of the IslamicEval 2025 subtasks, mirroring the bundled
development subset (`data/islamiceval_dev_subset.jsonl`):
| Subtask | Field | Values |
|---|---|---|
| 1A detection | `Span_Start`, `Span_End`, `Span_Type` | character offsets (end exclusive, quotation marks excluded); `Ayah` / `Hadith` |
| 1B verification | `Label` | `Correct` / `Incorrect` (anything not verified is reported as `Incorrect`) |
| 1C correction | `Correction` | exact source text for incorrect spans, or `خطأ` when no authentic source exists |
TSV columns: `Response_ID, Span_Start, Span_End, Span_Type, Label, Correction`. **Check the column names against the
organisers' current submission page before an official submission**; the official sites could not be fetched automatically
while preparing this version, so the exact official column names are not verified here.
## Run and deploy
```bash
python index_builder.py # only after changing data/
python build_static_space.py # writes ./static_space/
```
Upload the **contents** of `static_space/` to a Hugging Face Space with SDK = *Static* (free). Local Gradio version:
`pip install -r requirements.txt && OPENAI_API_KEY=... python app.py`. Tests: `python -m unittest discover -s tests`.
## API key (important)
A Static Space has no server, so a key placed in `web/config.js` is **readable by every visitor**. This build ships the
ChatGPT key in that file for the live demo. Keep a small monthly spend limit on the key in the OpenAI dashboard, revoke it
after the evaluation, and never commit it to a public repository (GitHub scans for and revokes exposed OpenAI keys).
A safer production design is a small proxy (for example a Cloudflare Worker) that holds the key as a secret.
## Project structure
| Path | Role |
|---|---|
| `web/` | Front end: `index.html`, `styles.css` (glassmorphism, RTL), `app.js` (UI), `worker.js` (Pyodide worker), `cache.js`, `camelbert.js`, `config.js` |
| `normalization.py`, `idgham.py` | Phonetic-aware Arabic normalisation; mushaf idgham rendering for corrections |
| `index_builder.py`, `retrieval.py` | Pre-built BM25 / n-gram indexes and the lazy-loading retriever |
| `alignment.py`, `similarity.py` | Word alignment and similarity indicators |
| `detector.py`, `scanner.py` | Quotation detection (corpus-first; trigger phrases optional) and corpus scanning |
| `verifier.py` | Verification, source-backed correction, decisions |
| `benchmark.py` | IslamicEval-style JSON / TSV export |
| `llm_client.py`, `app.py`, `ui.py` | LLM client, Gradio app and Python entry points used by the page, HTML rendering |
| `research/train_camelbert.py` | Improved CAMeLBERT-MSA trainer (GPU, Colab/Kaggle) |
| `evaluate.py`, `tests/`, `research/` | Prototype evaluation, tests, original notebook and optional training code |
## Training CAMeLBERT-MSA (Subtask 1A)
`research/train_camelbert.py` is the improved trainer (generated LLM-style data with and without introductory phrases,
up-weighted real data, diacritic-free input with exact offset mapping, class-weighted loss, layer-wise learning-rate decay,
model selection on the official character-level macro-F1, clean BIO decoding, optional hybrid with the rule + corpus detector).
Training needs a GPU, which this project does not include: run it on a free Colab or Kaggle T4.
```bash
pip install torch transformers
python research/train_camelbert.py --holdout 10 --out models/camelbert_check # honest estimate; prints model / hybrid / rules F1
python research/train_camelbert.py --holdout 0 --epochs 6 --push-to-hub YOUR_USER/camelbert-msa-islamiceval-1a --out models/camelbert_1a
```
Then put `HF_MODEL_ID` (or `HF_ENDPOINT_URL` for a dedicated Inference Endpoint) in `web/config.js`. No score is promised;
compare the printed holdout F1 with the rule + corpus detector (`python evaluate.py`). The live call to a hosted model from
the browser was not tested in this build; the simulation is the default.
## Try these inputs
| Input | Expected result |
|---|---|
| `قال الله تعالى: "فَاسْتَقِمْ كَمَا أُمِرْتَ وَمَنْ تَابَ مَعَكَ".` | Ayah, verified |
| `قال الله تعالى: "إن الله لا يغفر أن يشرك به ويغفر ما دون ذلك لمن يريد".` | Ayah, mismatch with the exact source text |
| `من أعظم ما يثبّت القلب قوله إِنَّ مَعَ الْعُسْرِ يُسْرًا فلا تيأس.` (no phrase, no quotes) | Ayah, verified |
| `إنما الأعمال بالنيات وإنما لكل امرئ ما نوى.` (no phrase, no quotes) | Hadith, verified |
| `قال رسول الله ﷺ: "من قرأ سورة الإخلاص ألف مرة دخل الجنة بغير حساب ولا عقاب".` | Hadith, no matching source |
| `قال الله تعالى: "إن الله مع الصابرين والمحسنين في كل حين".` | Ayah, human review (invented verse) |
The same samples are buttons in the page.
## Troubleshooting "Ask then Verify"
The page shows the exact reason reported by OpenAI. "Rate limit / balance exceeded" (HTTP 429, code `insufficient_quota`) means the
OpenAI account behind the key has no credit left or billing is not active: add credit under Billing, or put another key in
`web/config.js`. Code cannot fix this. While it is unresolved the page offers a built-in sample answer so the verification
flow can still be demonstrated.
## Evaluation (development subset)
Computed with `python evaluate.py` on the bundled IslamicEval 2025 development subset after decoupling detection from
trigger phrases. These numbers are **not** the earlier research baseline (1A macro F1 0.9091 with CAMeLBERT-MSA; 1B 93.12%; 1C 70.39%).
| Component | Result |
|---|---|
| 1A detection, character-level macro F1, hybrid detector | 0.8263 |
| 1B verdict accuracy on gold spans (247 spans; always-"Correct" baseline 59.5%) | 93.12% |
| 1C correction accuracy (179 spans; always-"no source" baseline 62.0%) | 74.30% |
## Limitations
* Unannounced-quotation detection is corpus matching, not understanding; text far from the bundled corpora is missed.
* The tool aids textual checking; it is not a fatwa and does not replace specialist review.
* The first visit downloads the Pyodide runtime and about 16 MB of data; later visits use the browser cache.
## Competition requirements
See [REFERENCES.md](REFERENCES.md) for the sources used. Fill in the rules and deadlines from the organisers' documents
(<https://islamicaich.org/>) in your submission; they could not be read automatically while preparing this version.
License: MIT.