File size: 10,601 Bytes
01d4aab
0390c03
 
 
 
01d4aab
0390c03
01d4aab
0390c03
 
01d4aab
 
0390c03
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
---
title: التحقق من هلوسة القرآن والحديث وتصحيحها
emoji: 🕌
colorFrom: green
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: التحقق من هلوسة القرآن والحديث وتصحيحها
---


# التحقق من هلوسة القرآن والحديث وتصحيحها · Quran & Hadith Hallucination Verifier and Corrector

Arabic version: [README_AR.md](README_AR.md)

Detects Quran and Hadith quotations in any text (for example answers written by language models), verifies each one
against the sources, shows the evidence, and proposes a correction **taken verbatim from the source text**. When the
evidence is not sufficient the tool says so and refers the case to human review. The user interface is Arabic.

## Live Demo

**[LIVE_DEPLOYMENT_URL]**  <!-- replace with the public Hugging Face Space URL after deployment -->

## Features

* **Direct verification:** paste any paragraph; every Quran / Hadith quotation is found, compared word by word and
  highlighted, with the exact source text offered as the correction.
* **Ask then Verify:** type a question; a pre-configured ChatGPT (OpenAI) client answers and the same engine verifies the
  answer and returns a corrected version. There is no model selection in this tab.
* **No trigger phrases required:** quotations are recognised by matching the text with the Quran and Hadith corpora, with or
  without quotation marks and with or without openers such as "قال الله تعالى" or "قال رسول الله ﷺ". Such phrases are only a
  weak fallback for altered quotations that the corpus cannot confirm.
* **Detection model selector (English):** *Standard · Rules + Corpus*, or *CAMeLBERT-MSA · Fine-tuned*. See below.
* **Benchmark export:** JSON and TSV in the IslamicEval 1A / 1B / 1C structure (see [Benchmark alignment](#benchmark-alignment)).
* **Copy corrected text:** a working clipboard button with a toast message ("تم نسخ النص بنجاح"), with a fallback for iframes.
* **Fast, informative loading:** the Quran index starts downloading before the scripts run, the Hadith index loads in the
  background, files are cached by the browser (Cache API), verification results and demo examples are cached as JSON in
  `localStorage`, and animated skeleton cards plus a three-step progress bar show what is happening.

## CAMeLBERT-MSA support

`web/camelbert.js` exposes one function, `detectSpans`, that returns BIO-tagged tokens (`B-Ayah`, `I-Ayah`, `B-Hadith`,
`I-Hadith`, `O`) and spans:

* **Live mode:** set `HF_MODEL_ID` (and `HF_TOKEN` for a private model) in `web/config.js`. The page calls the Hugging Face
  Inference API for a fine-tuned token-classification model. The spans are then verified and corrected by the local pipeline.
* **Simulation mode (default):** with no model ID set, **no model weights are loaded**. The spans come from the standard
  detector and are shown as BIO tags with deterministic pseudo-confidence values. The interface labels this clearly as
  *Simulation*. Use it to demonstrate the flow, not to report model accuracy.

## Workflow

```
text -> detection (quotation marks + corpus match, then corpus scan for unannounced quotes) -> BM25 retrieval
     -> span-local alignment (phonetic skeleton, sliding window, LCS, dynamic gap)
     -> exact: verified | altered Quran passage: source correction | uncertain: human review | nothing similar: no source
```

| Decision | Colour | Meaning |
|---|---|---|
| موثّق (verified) | green | Equals its source after ignoring spelling and diacritic conventions. |
| غير مطابق (mismatch) | red | Differs from the source (Quran: the exact source text is offered) or no source resembles it. |
| يحتاج مراجعة بشرية (human review) | amber | Evidence is insufficient or ambiguous. No correction is produced. |

## Benchmark alignment

`benchmark.py` converts a pipeline result to the structure of the IslamicEval 2025 subtasks, mirroring the bundled
development subset (`data/islamiceval_dev_subset.jsonl`):

| Subtask | Field | Values |
|---|---|---|
| 1A detection | `Span_Start`, `Span_End`, `Span_Type` | character offsets (end exclusive, quotation marks excluded); `Ayah` / `Hadith` |
| 1B verification | `Label` | `Correct` / `Incorrect` (anything not verified is reported as `Incorrect`) |
| 1C correction | `Correction` | exact source text for incorrect spans, or `خطأ` when no authentic source exists |

TSV columns: `Response_ID, Span_Start, Span_End, Span_Type, Label, Correction`. **Check the column names against the
organisers' current submission page before an official submission**; the official sites could not be fetched automatically
while preparing this version, so the exact official column names are not verified here.

## Run and deploy

```bash
python index_builder.py          # only after changing data/
python build_static_space.py     # writes ./static_space/
```

Upload the **contents** of `static_space/` to a Hugging Face Space with SDK = *Static* (free). Local Gradio version:
`pip install -r requirements.txt && OPENAI_API_KEY=... python app.py`. Tests: `python -m unittest discover -s tests`.

## API key (important)

A Static Space has no server, so a key placed in `web/config.js` is **readable by every visitor**. This build ships the
ChatGPT key in that file for the live demo. Keep a small monthly spend limit on the key in the OpenAI dashboard, revoke it
after the evaluation, and never commit it to a public repository (GitHub scans for and revokes exposed OpenAI keys).
A safer production design is a small proxy (for example a Cloudflare Worker) that holds the key as a secret.

## Project structure

| Path | Role |
|---|---|
| `web/` | Front end: `index.html`, `styles.css` (glassmorphism, RTL), `app.js` (UI), `worker.js` (Pyodide worker), `cache.js`, `camelbert.js`, `config.js` |
| `normalization.py`, `idgham.py` | Phonetic-aware Arabic normalisation; mushaf idgham rendering for corrections |
| `index_builder.py`, `retrieval.py` | Pre-built BM25 / n-gram indexes and the lazy-loading retriever |
| `alignment.py`, `similarity.py` | Word alignment and similarity indicators |
| `detector.py`, `scanner.py` | Quotation detection (corpus-first; trigger phrases optional) and corpus scanning |
| `verifier.py` | Verification, source-backed correction, decisions |
| `benchmark.py` | IslamicEval-style JSON / TSV export |
| `llm_client.py`, `app.py`, `ui.py` | LLM client, Gradio app and Python entry points used by the page, HTML rendering |
| `research/train_camelbert.py` | Improved CAMeLBERT-MSA trainer (GPU, Colab/Kaggle) |
| `evaluate.py`, `tests/`, `research/` | Prototype evaluation, tests, original notebook and optional training code |

## Training CAMeLBERT-MSA (Subtask 1A)

`research/train_camelbert.py` is the improved trainer (generated LLM-style data with and without introductory phrases,
up-weighted real data, diacritic-free input with exact offset mapping, class-weighted loss, layer-wise learning-rate decay,
model selection on the official character-level macro-F1, clean BIO decoding, optional hybrid with the rule + corpus detector).
Training needs a GPU, which this project does not include: run it on a free Colab or Kaggle T4.

```bash
pip install torch transformers
python research/train_camelbert.py --holdout 10 --out models/camelbert_check        # honest estimate; prints model / hybrid / rules F1
python research/train_camelbert.py --holdout 0 --epochs 6 --push-to-hub YOUR_USER/camelbert-msa-islamiceval-1a --out models/camelbert_1a
```

Then put `HF_MODEL_ID` (or `HF_ENDPOINT_URL` for a dedicated Inference Endpoint) in `web/config.js`. No score is promised;
compare the printed holdout F1 with the rule + corpus detector (`python evaluate.py`). The live call to a hosted model from
the browser was not tested in this build; the simulation is the default.

## Try these inputs

| Input | Expected result |
|---|---|
| `قال الله تعالى: "فَاسْتَقِمْ كَمَا أُمِرْتَ وَمَنْ تَابَ مَعَكَ".` | Ayah, verified |
| `قال الله تعالى: "إن الله لا يغفر أن يشرك به ويغفر ما دون ذلك لمن يريد".` | Ayah, mismatch with the exact source text |
| `من أعظم ما يثبّت القلب قوله إِنَّ مَعَ الْعُسْرِ يُسْرًا فلا تيأس.` (no phrase, no quotes) | Ayah, verified |
| `إنما الأعمال بالنيات وإنما لكل امرئ ما نوى.` (no phrase, no quotes) | Hadith, verified |
| `قال رسول الله ﷺ: "من قرأ سورة الإخلاص ألف مرة دخل الجنة بغير حساب ولا عقاب".` | Hadith, no matching source |
| `قال الله تعالى: "إن الله مع الصابرين والمحسنين في كل حين".` | Ayah, human review (invented verse) |

The same samples are buttons in the page.

## Troubleshooting "Ask then Verify"

The page shows the exact reason reported by OpenAI. "Rate limit / balance exceeded" (HTTP 429, code `insufficient_quota`) means the
OpenAI account behind the key has no credit left or billing is not active: add credit under Billing, or put another key in
`web/config.js`. Code cannot fix this. While it is unresolved the page offers a built-in sample answer so the verification
flow can still be demonstrated.

## Evaluation (development subset)

Computed with `python evaluate.py` on the bundled IslamicEval 2025 development subset after decoupling detection from
trigger phrases. These numbers are **not** the earlier research baseline (1A macro F1 0.9091 with CAMeLBERT-MSA; 1B 93.12%; 1C 70.39%).

| Component | Result |
|---|---|
| 1A detection, character-level macro F1, hybrid detector | 0.8263 |
| 1B verdict accuracy on gold spans (247 spans; always-"Correct" baseline 59.5%) | 93.12% |
| 1C correction accuracy (179 spans; always-"no source" baseline 62.0%) | 74.30% |

## Limitations

* Unannounced-quotation detection is corpus matching, not understanding; text far from the bundled corpora is missed.
* The tool aids textual checking; it is not a fatwa and does not replace specialist review.
* The first visit downloads the Pyodide runtime and about 16 MB of data; later visits use the browser cache.

## Competition requirements

See [REFERENCES.md](REFERENCES.md) for the sources used. Fill in the rules and deadlines from the organisers' documents
(<https://islamicaich.org/>) in your submission; they could not be read automatically while preparing this version.

License: MIT.