validops-east-3 commited on
Commit
9cacfae
·
verified ·
1 Parent(s): efff1ff

Deploy de83ed8cb4e1bbfa6f888d334f247492ce1b5bbf

Browse files
Files changed (4) hide show
  1. README.md +1 -99
  2. docqa/bench/accuracy.py +121 -0
  3. docqa/bench/check_deps.py +158 -0
  4. requirements.txt +33 -14
README.md CHANGED
@@ -7,102 +7,4 @@ sdk: docker
7
  pinned: false
8
  app_port: 7860
9
  license: mit
10
- ---
11
-
12
- # Document Data Extraction
13
-
14
- Extract arbitrary user-defined keys from invoices, receipts and scanned
15
- documents. Upload a PDF or image, list the keys you want, get one structured
16
- result per key.
17
-
18
- ```bash
19
- curl -X POST https://<your-space>.hf.space/v1/extract \
20
- -F "file=@invoice.pdf" \
21
- -F "keys=INVOICE NO,Vendor Id,Invoice Date,Total Amount,GSTIN"
22
- ```
23
-
24
- Keys are free-form. Nothing is hard-coded — `INVOICE NO` and `Total Amount` are
25
- just strings. Duplicates are returned separately, index-aligned with the
26
- request.
27
-
28
- ## Read `status`, not `value`
29
-
30
- | status | meaning |
31
- |---|---|
32
- | `extracted` | confidence >= threshold. Use it. |
33
- | `low_confidence` | a span was found but fell below threshold. Route to a human. |
34
- | `not_found` | nothing usable on the page. |
35
- | `error` | the key could not be processed. |
36
-
37
- `impira/layoutlm-invoices` is an extractive model with **no null-answer head**:
38
- it returns *some* span for every input, including questions the document
39
- cannot answer. Measured on this project's benchmark it gave a confident wrong
40
- value for **16.7%** of deliberately unanswerable invoice questions, and
41
- abstained **0%** of the time. The confidence threshold exists specifically to
42
- catch this. Treating `value` without checking `status` will surface invented
43
- data.
44
-
45
- ## Model
46
-
47
- `impira/layoutlm-invoices`, loaded from the Hugging Face cache at start. That
48
- model is **CC-BY-NC-SA-4.0, non-commercial only**. For commercial use set
49
- `DOCX_MODEL_ID=impira/layoutlm-document-qa` (MIT) — no code change needed.
50
-
51
- LayoutLM reads 2D position embeddings from bounding boxes, so it needs the
52
- document's geometry. Digital PDFs use the text layer; scans and images are
53
- rasterised and OCR'd with poppler-utils and tesseract-ocr, both in the image.
54
-
55
- ## Configuration
56
-
57
- | variable | default | meaning |
58
- |---|---|---|
59
- | `DOCX_MODEL_ID` | `impira/layoutlm-invoices` | any LayoutLM QA checkpoint |
60
- | `DOCX_TORCH_THREADS` | `4` | intra-op threads |
61
- | `DOCX_BATCH_SIZE` | `8` | keys per forward pass |
62
- | `DOCX_MAX_PAGES` | `3` | pages processed per document |
63
- | `DOCX_MAX_UPLOAD_MB` | `25` | body size limit |
64
- | `DOCX_CONFIDENCE_THRESHOLD` | `0.5` | `extracted` vs `low_confidence` |
65
- | `DOCX_REQUEST_TIMEOUT_S` | `60` | per-request timeout |
66
-
67
- Check what is actually in effect with `GET /v1/config`.
68
-
69
- ## Endpoints
70
-
71
- | method | path |
72
- |---|---|
73
- | `GET` | `/health` |
74
- | `GET` | `/v1/config` |
75
- | `POST` | `/v1/extract` (multipart `file`, `keys`) |
76
- | `POST` | `/v1/extract/json` (base64 + keys array) |
77
-
78
- `/docs` serves the OpenAPI UI.
79
-
80
- ## Measured performance
81
-
82
- Single key, 148-word page, CPU:
83
-
84
- | runtime | ms/key |
85
- |---|---|
86
- | PyTorch fp32 | 578 |
87
- | ONNX fp32 | 407 |
88
-
89
- Batching gives no CPU speedup (measured exactly linear) — the GEMMs are already
90
- compute-bound. Thread count is the real lever.
91
-
92
- Accuracy on the 12-invoice corpus with pdfplumber word extraction:
93
-
94
- | metric | invoices | general docs |
95
- |---|---|---|
96
- | exact match | 0.7604 | 0.5349 |
97
- | F1 | 0.8091 | 0.5416 |
98
- | thresholded | 0.7188 | — |
99
-
100
- ## Notes
101
-
102
- `requirements.txt` pins `numpy==1.26.4`; `onnx` pulls numpy 2.x, which breaks
103
- scipy and sklearn in the benchmark scripts. `torch` comes from the CPU wheel
104
- index — the default PyPI torch pulls multi-GB CUDA libraries.
105
-
106
- The previous llama-server (Qwen3-0.6B) service on this same port is commented
107
- out in the `Dockerfile`, not deleted, so the OpenAI-compatible chat endpoint can
108
- be restored.
 
7
  pinned: false
8
  app_port: 7860
9
  license: mit
10
+ ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
docqa/bench/accuracy.py ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Accuracy check for the docqa service against corpus ground truth.
3
+
4
+ python bench/accuracy.py --corpus <dir> [--limit N] [--base URL]
5
+
6
+ Every number printed is measured from a real run against a running server.
7
+ Nothing here is estimated.
8
+ """
9
+
10
+ from __future__ import annotations
11
+
12
+ import argparse
13
+ import json
14
+ import statistics
15
+ import subprocess
16
+ import time
17
+ from pathlib import Path
18
+
19
+ # Requested key -> ground-truth field in corpus/ground_truth.json
20
+ KEY_MAP = [
21
+ ("INVOICE NO", "invoice_number"),
22
+ ("Invoice Date", "invoice_date"),
23
+ ("Due Date", "due_date"),
24
+ ("PO Number", "po_number"),
25
+ ("Payment Term", "payment_terms"),
26
+ ("Currency", "currency"),
27
+ ("Vendor Name", "vendor_name"),
28
+ ("Vendor Address", "vendor_address"),
29
+ ("Vendor Id", "vendor_tax_id"),
30
+ ("GSTIN", "gstin"),
31
+ ("Bill To", "bill_to"),
32
+ ("Subtotal", "subtotal"),
33
+ ("Tax Amount", "tax_amount"),
34
+ ("Total Amount", "total"),
35
+ ("Tax Label", "tax_label"),
36
+ ]
37
+
38
+
39
+ def extract(base: str, pdf: Path, keys: list[str]) -> tuple[dict, float]:
40
+ started = time.perf_counter()
41
+ proc = subprocess.run(
42
+ ["curl.exe", "-sS", "-X", "POST", f"{base}/v1/extract",
43
+ "-F", f"file=@{pdf}", "-F", "keys=" + ",".join(keys)],
44
+ capture_output=True, text=True,
45
+ )
46
+ wall = (time.perf_counter() - started) * 1000
47
+ if proc.returncode != 0:
48
+ raise RuntimeError(proc.stderr.strip() or "curl failed")
49
+ return json.loads(proc.stdout), wall
50
+
51
+
52
+ def main() -> int:
53
+ ap = argparse.ArgumentParser()
54
+ ap.add_argument("--base", default="http://127.0.0.1:7860")
55
+ ap.add_argument("--corpus", type=Path, required=True)
56
+ ap.add_argument("--limit", type=int, default=1)
57
+ args = ap.parse_args()
58
+
59
+ manifest = json.loads((args.corpus / "ground_truth.json").read_text())
60
+ records = manifest[: args.limit] if args.limit else manifest
61
+ keys = [k for k, _ in KEY_MAP]
62
+
63
+ print(f"target {args.base}")
64
+ print(f"model impira/layoutlm-invoices keys={len(keys)} "
65
+ f"documents={len(records)}\n")
66
+
67
+ per_key: dict[str, list[int]] = {k: [] for k in keys}
68
+ latencies: list[float] = []
69
+
70
+ for record in records:
71
+ body, wall = extract(args.base, Path(record["path"]), keys)
72
+ latencies.append(body.get("latency_ms", 0.0))
73
+ print(f"{record['document_id']} variant={record['variant']} "
74
+ f"source={body.get('source')} words={body.get('word_count')} "
75
+ f"pages={body.get('pages_processed')}")
76
+ print(f"server={body.get('latency_ms', 0):.0f} ms "
77
+ f"wall={wall:.0f} ms (includes model load on first call)\n")
78
+
79
+ by_key = {f["key"]: f for f in body.get("fields", [])}
80
+ print(f" {'KEY':<14} {'EXPECTED':<26} {'STATUS':<15} "
81
+ f"{'GOT':<28} CONF")
82
+ print(" " + "-" * 92)
83
+ for key, truth in KEY_MAP:
84
+ expected = str(record[truth])
85
+ field = by_key.get(key, {})
86
+ got = (field.get("value") or "").strip()
87
+ ok = got == expected.strip()
88
+ per_key[key].append(1 if ok else 0)
89
+ print(f" {key:<14} {expected[:26]:<26} "
90
+ f"{str(field.get('status')):<15} "
91
+ f"{(got[:26] + (' OK' if ok else ' BAD')):<28} "
92
+ f"{field.get('confidence', 0):.4f}")
93
+ print(" " + "-" * 92 + "\n")
94
+
95
+ total_hit = sum(sum(v) for v in per_key.values())
96
+ total_n = sum(len(v) for v in per_key.values())
97
+
98
+ print("PER-KEY EXACT MATCH")
99
+ print("-" * 46)
100
+ for key, _ in KEY_MAP:
101
+ scores = per_key[key]
102
+ if scores:
103
+ acc = 100 * sum(scores) / len(scores)
104
+ print(f" {key:<16} {acc:6.1f}% ({sum(scores)}/{len(scores)})")
105
+ if total_n:
106
+ print(f"\n {'OVERALL':<16} {100*total_hit/total_n:6.1f}% "
107
+ f"({total_hit}/{total_n})")
108
+
109
+ if latencies:
110
+ print("\nLATENCY (whole document, all keys)")
111
+ print("-" * 46)
112
+ print(f" mean {statistics.mean(latencies):8.0f} ms")
113
+ print(f" min {min(latencies):8.0f} ms")
114
+ print(f" max {max(latencies):8.0f} ms")
115
+ print(f" per key {statistics.mean(latencies)/len(keys):6.0f} ms")
116
+
117
+ return 0
118
+
119
+
120
+ if __name__ == "__main__":
121
+ raise SystemExit(main())
docqa/bench/check_deps.py ADDED
@@ -0,0 +1,158 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Verify requirements.txt covers every third-party import in docqa/.
3
+
4
+ Run in CI or after adding a module:
5
+
6
+ python bench/check_deps.py
7
+
8
+ Exits non-zero if an import is unlisted, so a missing dependency is caught at
9
+ build time rather than as an ImportError on the first request that happens to
10
+ hit that code path.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import ast
16
+ import re
17
+ import sys
18
+ from pathlib import Path
19
+
20
+ ROOT = Path(__file__).resolve().parent.parent
21
+ REPO = ROOT.parent
22
+
23
+ # import name -> pip distribution name, where they differ
24
+ DISTRIBUTION = {
25
+ "PIL": "pillow",
26
+ "fitz": "pymupdf",
27
+ "sklearn": "scikit-learn",
28
+ "yaml": "pyyaml",
29
+ "cv2": "opencv-python",
30
+ "bs4": "beautifulsoup4",
31
+ "dotenv": "python-dotenv",
32
+ }
33
+
34
+ # Imported only inside functions guarded by try/except ImportError, or only by
35
+ # optional benchmark paths. Still listed in requirements.txt, but a missing
36
+ # one is a warning rather than a hard failure.
37
+ OPTIONAL = {"onnx", "onnxruntime", "sklearn"}
38
+
39
+ # Runtime plugin dependencies that no AST scan of our own code can discover,
40
+ # because the import happens inside the framework. FastAPI resolves these
41
+ # lazily: the app imports fine and the endpoint 500s on first use. Keep this in
42
+ # step with any File()/Form()/UploadFile declaration in api.py.
43
+ FRAMEWORK_PLUGINS = {
44
+ "python-multipart": "FastAPI File()/Form(); app starts but /v1/extract 500s",
45
+ }
46
+
47
+
48
+ def parse_requirements() -> dict[str, str]:
49
+ """Distribution name -> pinned version, from requirements.txt."""
50
+ out: dict[str, str] = {}
51
+ for line in (REPO / "requirements.txt").read_text(encoding="utf-8").splitlines():
52
+ line = line.split("#", 1)[0].strip()
53
+ if not line or line.startswith("-"):
54
+ continue
55
+ match = re.match(r"^([A-Za-z0-9._-]+)(?:\[[^\]]*\])?==(.+)$", line)
56
+ if match:
57
+ out[match.group(1).lower().replace("_", "-")] = match.group(2)
58
+ return out
59
+
60
+
61
+ def collect_imports() -> set[str]:
62
+ found: set[str] = set()
63
+ for path in ROOT.rglob("*.py"):
64
+ tree = ast.parse(path.read_text(encoding="utf-8"))
65
+ for node in ast.walk(tree):
66
+ if isinstance(node, ast.Import):
67
+ for alias in node.names:
68
+ found.add(alias.name.split(".")[0])
69
+ elif isinstance(node, ast.ImportFrom) and node.level == 0 and node.module:
70
+ found.add(node.module.split(".")[0])
71
+ return found
72
+
73
+
74
+ def verify_installed(declared: dict[str, str]) -> list[str]:
75
+ """Actually import each declared distribution.
76
+
77
+ An AST scan only sees what docqa/ imports directly. It cannot see what a
78
+ framework pulls in at request time, which is precisely how python-multipart
79
+ was missed: fastapi imports cleanly without it and only fails when
80
+ File()/Form() is exercised. Importing the real modules closes that gap.
81
+ """
82
+ import importlib
83
+
84
+ problems: list[str] = []
85
+ for dist in sorted(declared):
86
+ module = MODULE_FOR_DIST.get(dist)
87
+ if module is None:
88
+ continue
89
+ try:
90
+ importlib.import_module(module)
91
+ except ImportError as exc:
92
+ problems.append(f"{dist} (import {module}): {exc}")
93
+ return problems
94
+
95
+
96
+ # distribution name -> importable module, where they differ
97
+ MODULE_FOR_DIST = {
98
+ "pillow": "PIL",
99
+ "pymupdf": "fitz",
100
+ "scikit-learn": "sklearn",
101
+ "python-multipart": "multipart",
102
+ "python-dotenv": "dotenv",
103
+ "pyyaml": "yaml",
104
+ }
105
+
106
+
107
+ def main() -> int:
108
+ declared = parse_requirements()
109
+ imported = collect_imports()
110
+
111
+ stdlib = set(sys.stdlib_module_names)
112
+ local = {"docxextract"}
113
+ missing: list[str] = []
114
+ optional_missing: list[str] = []
115
+
116
+ for name in sorted(imported):
117
+ if name in stdlib or name in local:
118
+ continue
119
+ dist = DISTRIBUTION.get(name, name).lower().replace("_", "-")
120
+ if dist in declared:
121
+ continue
122
+ (optional_missing if name in OPTIONAL else missing).append(f"{name} -> {dist}")
123
+
124
+ # Plugins the framework needs at request time, invisible to the scan.
125
+ for dist, why in FRAMEWORK_PLUGINS.items():
126
+ if dist not in declared:
127
+ missing.append(f"{dist} <- {why}")
128
+
129
+ print(f"requirements.txt declares {len(declared)} distributions")
130
+ print(f"docqa/ imports {len(imported)} modules "
131
+ f"({len([m for m in imported if m not in stdlib])} non-stdlib)\n")
132
+
133
+ if optional_missing:
134
+ print("OPTIONAL, not declared (guarded imports / optional paths):")
135
+ for item in optional_missing:
136
+ print(f" {item}")
137
+ print()
138
+
139
+ broken = verify_installed(declared)
140
+ if broken:
141
+ print("DECLARED BUT NOT IMPORTABLE in this environment:")
142
+ for item in broken:
143
+ print(f" {item}")
144
+ print()
145
+
146
+ if missing or broken:
147
+ if missing:
148
+ print("MISSING, must be added to requirements.txt:")
149
+ for item in missing:
150
+ print(f" {item}")
151
+ return 1
152
+
153
+ print("OK: every non-stdlib import is declared and importable.")
154
+ return 0
155
+
156
+
157
+ if __name__ == "__main__":
158
+ raise SystemExit(main())
requirements.txt CHANGED
@@ -1,30 +1,49 @@
1
- # Pinned exactly. Loose pins here have broken this build more than once:
2
- # - onnx pulls numpy 2.x, which breaks scipy and sklearn used by the
3
- # benchmark scripts.
4
- # - torch must come from the CPU wheel index; the default PyPI torch pulls
5
- # multi-GB CUDA libraries and blows the Space image size limit.
 
 
 
 
6
  --extra-index-url https://download.pytorch.org/whl/cpu
7
 
 
8
  fastapi==0.115.5
9
  uvicorn[standard]==0.34.0
10
  pydantic==2.13.4
11
  pydantic-settings==2.15.0
12
  orjson==3.11.9
 
 
 
 
 
 
 
13
 
14
- # model
15
  torch==2.5.1+cpu
16
  transformers==4.50.2
17
- huggingface_hub>=0.26
 
 
 
18
 
19
- # parsing / OCR
20
  pdfplumber==0.11.9
21
  pytesseract==0.3.13
22
- Pillow==10.3.0
 
 
23
 
24
- # numpy must stay on 1.x; see note above
25
- numpy==1.26.4
 
 
 
26
 
27
- # tests
28
  pytest==8.3.4
29
- httpx==0.28.1
30
- pytest-asyncio==0.24.0
 
1
+ # Every third-party module imported anywhere under docqa/ appears here.
2
+ # Verified by AST-scanning all imports, not by guessing -- see
3
+ # bench/check_deps.py, which fails if an import is unlisted.
4
+ #
5
+ # Pinned exactly. Loose pins broke this build more than once:
6
+ # - torch must come from the CPU wheel index. The default PyPI torch pulls
7
+ # multi-GB CUDA libraries and exceeds the Space image limit.
8
+ # - numpy must stay on 1.x. onnxruntime/onnx pull numpy 2.x, which breaks
9
+ # scipy and scikit-learn used by the corpus generator.
10
  --extra-index-url https://download.pytorch.org/whl/cpu
11
 
12
+ # --- web service ---
13
  fastapi==0.115.5
14
  uvicorn[standard]==0.34.0
15
  pydantic==2.13.4
16
  pydantic-settings==2.15.0
17
  orjson==3.11.9
18
+ httpx==0.28.1
19
+ # REQUIRED by FastAPI at runtime for any endpoint declaring File() or Form().
20
+ # It is not imported by our own code, so an AST scan of docqa/ cannot find it:
21
+ # FastAPI raises at request time ("Form data requires python-multipart"), not
22
+ # at import time, which means the app starts cleanly and then 500s on
23
+ # /v1/extract. Added explicitly after that exact failure.
24
+ python-multipart==0.0.20
25
 
26
+ # --- model ---
27
  torch==2.5.1+cpu
28
  transformers==4.50.2
29
+ huggingface-hub==0.28.1
30
+ onnxruntime==1.20.1
31
+ numpy==1.26.4
32
+ scipy==1.10.1
33
 
34
+ # --- document parsing / OCR ---
35
  pdfplumber==0.11.9
36
  pytesseract==0.3.13
37
+ pillow==10.3.0
38
+ # fitz: rasterises scanned pages for the OCR fallback
39
+ pymupdf==1.25.1
40
 
41
+ # --- test corpus generation ---
42
+ # reportlab draws the synthetic invoices in tests/corpus.py; scikit-learn is
43
+ # used by bench/compare_backends.py for scoring.
44
+ reportlab==3.6.13
45
+ scikit-learn==1.6.1
46
 
47
+ # --- tests ---
48
  pytest==8.3.4
49
+ pytest-asyncio==1.3.0