Spaces:
Running
Running
File size: 7,889 Bytes
806b04d c08b36a 806b04d c08b36a 806b04d c08b36a 806b04d c08b36a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | ---
title: التحقق من هلوسة القرآن والحديث وتصحيحها
emoji: 🕌
colorFrom: green
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: التحقق من هلوسة القرآن والحديث وتصحيحها
---
# التحقق من هلوسة القرآن والحديث وتصحيحها · Quran & Hadith Hallucination Verifier and Corrector
Arabic version: [README_AR.md](README_AR.md)
Detects Quran and Hadith quotations in any text (for example answers written by language models), verifies each one
against the sources, shows the evidence, and proposes a correction **taken verbatim from the source text**. When the
evidence is not sufficient the tool says so and refers the case to human review. The user interface is Arabic.
## Live Demo
**[LIVE_DEPLOYMENT_URL]** <!-- replace with the public Hugging Face Space URL after deployment -->
## Features
* **Direct verification:** paste any paragraph; every Quran / Hadith quotation is found, compared word by word and
highlighted, with the exact source text offered as the correction.
* **Ask then Verify:** type a question; a pre-configured ChatGPT (OpenAI) client answers and the same engine verifies the
answer and returns a corrected version. There is no model selection in this tab.
* **No trigger phrases required:** quotations are recognised by matching the text with the Quran and Hadith corpora, with or
without quotation marks and with or without openers such as "قال الله تعالى" or "قال رسول الله ﷺ". Such phrases are only a
weak fallback for altered quotations that the corpus cannot confirm.
* **Detection model selector (English):** *Standard · Rules + Corpus*, or *CAMeLBERT-MSA · Fine-tuned*. See below.
* **Benchmark export:** JSON and TSV in the IslamicEval 1A / 1B / 1C structure (see [Benchmark alignment](#benchmark-alignment)).
* **Copy corrected text:** a working clipboard button with a toast message ("تم نسخ النص بنجاح"), with a fallback for iframes.
* **Fast, informative loading:** the Quran index starts downloading before the scripts run, the Hadith index loads in the
background, files are cached by the browser (Cache API), verification results and demo examples are cached as JSON in
`localStorage`, and animated skeleton cards plus a three-step progress bar show what is happening.
## CAMeLBERT-MSA support
`web/camelbert.js` exposes one function, `detectSpans`, that returns BIO-tagged tokens (`B-Ayah`, `I-Ayah`, `B-Hadith`,
`I-Hadith`, `O`) and spans:
* **Live mode:** set `HF_MODEL_ID` (and `HF_TOKEN` for a private model) in `web/config.js`. The page calls the Hugging Face
Inference API for a fine-tuned token-classification model. The spans are then verified and corrected by the local pipeline.
* **Simulation mode (default):** with no model ID set, **no model weights are loaded**. The spans come from the standard
detector and are shown as BIO tags with deterministic pseudo-confidence values. The interface labels this clearly as
*Simulation*. Use it to demonstrate the flow, not to report model accuracy.
## Workflow
```
text -> detection (quotation marks + corpus match, then corpus scan for unannounced quotes) -> BM25 retrieval
-> span-local alignment (phonetic skeleton, sliding window, LCS, dynamic gap)
-> exact: verified | altered Quran passage: source correction | uncertain: human review | nothing similar: no source
```
| Decision | Colour | Meaning |
|---|---|---|
| موثّق (verified) | green | Equals its source after ignoring spelling and diacritic conventions. |
| غير مطابق (mismatch) | red | Differs from the source (Quran: the exact source text is offered) or no source resembles it. |
| يحتاج مراجعة بشرية (human review) | amber | Evidence is insufficient or ambiguous. No correction is produced. |
## Benchmark alignment
`benchmark.py` converts a pipeline result to the structure of the IslamicEval 2025 subtasks, mirroring the bundled
development subset (`data/islamiceval_dev_subset.jsonl`):
| Subtask | Field | Values |
|---|---|---|
| 1A detection | `Span_Start`, `Span_End`, `Span_Type` | character offsets (end exclusive, quotation marks excluded); `Ayah` / `Hadith` |
| 1B verification | `Label` | `Correct` / `Incorrect` (anything not verified is reported as `Incorrect`) |
| 1C correction | `Correction` | exact source text for incorrect spans, or `خطأ` when no authentic source exists |
TSV columns: `Response_ID, Span_Start, Span_End, Span_Type, Label, Correction`. **Check the column names against the
organisers' current submission page before an official submission**; the official sites could not be fetched automatically
while preparing this version, so the exact official column names are not verified here.
## Run and deploy
```bash
python index_builder.py # only after changing data/
python build_static_space.py # writes ./static_space/
```
Upload the **contents** of `static_space/` to a Hugging Face Space with SDK = *Static* (free). Local Gradio version:
`pip install -r requirements.txt && OPENAI_API_KEY=... python app.py`. Tests: `python -m unittest discover -s tests`.
## API key (important)
A Static Space has no server, so a key placed in `web/config.js` is **readable by every visitor**. This build ships the
ChatGPT key in that file for the live demo. Keep a small monthly spend limit on the key in the OpenAI dashboard, revoke it
after the evaluation, and never commit it to a public repository (GitHub scans for and revokes exposed OpenAI keys).
A safer production design is a small proxy (for example a Cloudflare Worker) that holds the key as a secret.
## Project structure
| Path | Role |
|---|---|
| `web/` | Front end: `index.html`, `styles.css` (glassmorphism, RTL), `app.js` (UI), `worker.js` (Pyodide worker), `cache.js`, `camelbert.js`, `config.js` |
| `normalization.py`, `idgham.py` | Phonetic-aware Arabic normalisation; mushaf idgham rendering for corrections |
| `index_builder.py`, `retrieval.py` | Pre-built BM25 / n-gram indexes and the lazy-loading retriever |
| `alignment.py`, `similarity.py` | Word alignment and similarity indicators |
| `detector.py`, `scanner.py` | Quotation detection (corpus-first; trigger phrases optional) and corpus scanning |
| `verifier.py` | Verification, source-backed correction, decisions |
| `benchmark.py` | IslamicEval-style JSON / TSV export |
| `llm_client.py`, `app.py`, `ui.py` | LLM client, Gradio app and Python entry points used by the page, HTML rendering |
| `evaluate.py`, `tests/`, `research/` | Prototype evaluation, tests, original notebook and optional training code |
## Evaluation (development subset)
Computed with `python evaluate.py` on the bundled IslamicEval 2025 development subset after decoupling detection from
trigger phrases. These numbers are **not** the earlier research baseline (1A macro F1 0.9091 with CAMeLBERT-MSA; 1B 93.12%; 1C 70.39%).
| Component | Result |
|---|---|
| 1A detection, character-level macro F1, hybrid detector | 0.8247 |
| 1B verdict accuracy on gold spans (247 spans; always-"Correct" baseline 59.5%) | 93.12% |
| 1C correction accuracy (179 spans; always-"no source" baseline 62.0%) | 74.30% |
## Limitations
* Unannounced-quotation detection is corpus matching, not understanding; text far from the bundled corpora is missed.
* The tool aids textual checking; it is not a fatwa and does not replace specialist review.
* The first visit downloads the Pyodide runtime and about 16 MB of data; later visits use the browser cache.
## Competition requirements
See [REFERENCES.md](REFERENCES.md) for the sources used. Fill in the rules and deadlines from the organisers' documents
(<https://islamicaich.org/>) in your submission; they could not be read automatically while preparing this version.
License: MIT.
|