Spaces:
Running
Running
Deploy de83ed8cb4e1bbfa6f888d334f247492ce1b5bbf
Browse files- README.md +1 -99
- docqa/bench/accuracy.py +121 -0
- docqa/bench/check_deps.py +158 -0
- requirements.txt +33 -14
README.md
CHANGED
|
@@ -7,102 +7,4 @@ sdk: docker
|
|
| 7 |
pinned: false
|
| 8 |
app_port: 7860
|
| 9 |
license: mit
|
| 10 |
-
---
|
| 11 |
-
|
| 12 |
-
# Document Data Extraction
|
| 13 |
-
|
| 14 |
-
Extract arbitrary user-defined keys from invoices, receipts and scanned
|
| 15 |
-
documents. Upload a PDF or image, list the keys you want, get one structured
|
| 16 |
-
result per key.
|
| 17 |
-
|
| 18 |
-
```bash
|
| 19 |
-
curl -X POST https://<your-space>.hf.space/v1/extract \
|
| 20 |
-
-F "file=@invoice.pdf" \
|
| 21 |
-
-F "keys=INVOICE NO,Vendor Id,Invoice Date,Total Amount,GSTIN"
|
| 22 |
-
```
|
| 23 |
-
|
| 24 |
-
Keys are free-form. Nothing is hard-coded — `INVOICE NO` and `Total Amount` are
|
| 25 |
-
just strings. Duplicates are returned separately, index-aligned with the
|
| 26 |
-
request.
|
| 27 |
-
|
| 28 |
-
## Read `status`, not `value`
|
| 29 |
-
|
| 30 |
-
| status | meaning |
|
| 31 |
-
|---|---|
|
| 32 |
-
| `extracted` | confidence >= threshold. Use it. |
|
| 33 |
-
| `low_confidence` | a span was found but fell below threshold. Route to a human. |
|
| 34 |
-
| `not_found` | nothing usable on the page. |
|
| 35 |
-
| `error` | the key could not be processed. |
|
| 36 |
-
|
| 37 |
-
`impira/layoutlm-invoices` is an extractive model with **no null-answer head**:
|
| 38 |
-
it returns *some* span for every input, including questions the document
|
| 39 |
-
cannot answer. Measured on this project's benchmark it gave a confident wrong
|
| 40 |
-
value for **16.7%** of deliberately unanswerable invoice questions, and
|
| 41 |
-
abstained **0%** of the time. The confidence threshold exists specifically to
|
| 42 |
-
catch this. Treating `value` without checking `status` will surface invented
|
| 43 |
-
data.
|
| 44 |
-
|
| 45 |
-
## Model
|
| 46 |
-
|
| 47 |
-
`impira/layoutlm-invoices`, loaded from the Hugging Face cache at start. That
|
| 48 |
-
model is **CC-BY-NC-SA-4.0, non-commercial only**. For commercial use set
|
| 49 |
-
`DOCX_MODEL_ID=impira/layoutlm-document-qa` (MIT) — no code change needed.
|
| 50 |
-
|
| 51 |
-
LayoutLM reads 2D position embeddings from bounding boxes, so it needs the
|
| 52 |
-
document's geometry. Digital PDFs use the text layer; scans and images are
|
| 53 |
-
rasterised and OCR'd with poppler-utils and tesseract-ocr, both in the image.
|
| 54 |
-
|
| 55 |
-
## Configuration
|
| 56 |
-
|
| 57 |
-
| variable | default | meaning |
|
| 58 |
-
|---|---|---|
|
| 59 |
-
| `DOCX_MODEL_ID` | `impira/layoutlm-invoices` | any LayoutLM QA checkpoint |
|
| 60 |
-
| `DOCX_TORCH_THREADS` | `4` | intra-op threads |
|
| 61 |
-
| `DOCX_BATCH_SIZE` | `8` | keys per forward pass |
|
| 62 |
-
| `DOCX_MAX_PAGES` | `3` | pages processed per document |
|
| 63 |
-
| `DOCX_MAX_UPLOAD_MB` | `25` | body size limit |
|
| 64 |
-
| `DOCX_CONFIDENCE_THRESHOLD` | `0.5` | `extracted` vs `low_confidence` |
|
| 65 |
-
| `DOCX_REQUEST_TIMEOUT_S` | `60` | per-request timeout |
|
| 66 |
-
|
| 67 |
-
Check what is actually in effect with `GET /v1/config`.
|
| 68 |
-
|
| 69 |
-
## Endpoints
|
| 70 |
-
|
| 71 |
-
| method | path |
|
| 72 |
-
|---|---|
|
| 73 |
-
| `GET` | `/health` |
|
| 74 |
-
| `GET` | `/v1/config` |
|
| 75 |
-
| `POST` | `/v1/extract` (multipart `file`, `keys`) |
|
| 76 |
-
| `POST` | `/v1/extract/json` (base64 + keys array) |
|
| 77 |
-
|
| 78 |
-
`/docs` serves the OpenAPI UI.
|
| 79 |
-
|
| 80 |
-
## Measured performance
|
| 81 |
-
|
| 82 |
-
Single key, 148-word page, CPU:
|
| 83 |
-
|
| 84 |
-
| runtime | ms/key |
|
| 85 |
-
|---|---|
|
| 86 |
-
| PyTorch fp32 | 578 |
|
| 87 |
-
| ONNX fp32 | 407 |
|
| 88 |
-
|
| 89 |
-
Batching gives no CPU speedup (measured exactly linear) — the GEMMs are already
|
| 90 |
-
compute-bound. Thread count is the real lever.
|
| 91 |
-
|
| 92 |
-
Accuracy on the 12-invoice corpus with pdfplumber word extraction:
|
| 93 |
-
|
| 94 |
-
| metric | invoices | general docs |
|
| 95 |
-
|---|---|---|
|
| 96 |
-
| exact match | 0.7604 | 0.5349 |
|
| 97 |
-
| F1 | 0.8091 | 0.5416 |
|
| 98 |
-
| thresholded | 0.7188 | — |
|
| 99 |
-
|
| 100 |
-
## Notes
|
| 101 |
-
|
| 102 |
-
`requirements.txt` pins `numpy==1.26.4`; `onnx` pulls numpy 2.x, which breaks
|
| 103 |
-
scipy and sklearn in the benchmark scripts. `torch` comes from the CPU wheel
|
| 104 |
-
index — the default PyPI torch pulls multi-GB CUDA libraries.
|
| 105 |
-
|
| 106 |
-
The previous llama-server (Qwen3-0.6B) service on this same port is commented
|
| 107 |
-
out in the `Dockerfile`, not deleted, so the OpenAI-compatible chat endpoint can
|
| 108 |
-
be restored.
|
|
|
|
| 7 |
pinned: false
|
| 8 |
app_port: 7860
|
| 9 |
license: mit
|
| 10 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
docqa/bench/accuracy.py
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Accuracy check for the docqa service against corpus ground truth.
|
| 3 |
+
|
| 4 |
+
python bench/accuracy.py --corpus <dir> [--limit N] [--base URL]
|
| 5 |
+
|
| 6 |
+
Every number printed is measured from a real run against a running server.
|
| 7 |
+
Nothing here is estimated.
|
| 8 |
+
"""
|
| 9 |
+
|
| 10 |
+
from __future__ import annotations
|
| 11 |
+
|
| 12 |
+
import argparse
|
| 13 |
+
import json
|
| 14 |
+
import statistics
|
| 15 |
+
import subprocess
|
| 16 |
+
import time
|
| 17 |
+
from pathlib import Path
|
| 18 |
+
|
| 19 |
+
# Requested key -> ground-truth field in corpus/ground_truth.json
|
| 20 |
+
KEY_MAP = [
|
| 21 |
+
("INVOICE NO", "invoice_number"),
|
| 22 |
+
("Invoice Date", "invoice_date"),
|
| 23 |
+
("Due Date", "due_date"),
|
| 24 |
+
("PO Number", "po_number"),
|
| 25 |
+
("Payment Term", "payment_terms"),
|
| 26 |
+
("Currency", "currency"),
|
| 27 |
+
("Vendor Name", "vendor_name"),
|
| 28 |
+
("Vendor Address", "vendor_address"),
|
| 29 |
+
("Vendor Id", "vendor_tax_id"),
|
| 30 |
+
("GSTIN", "gstin"),
|
| 31 |
+
("Bill To", "bill_to"),
|
| 32 |
+
("Subtotal", "subtotal"),
|
| 33 |
+
("Tax Amount", "tax_amount"),
|
| 34 |
+
("Total Amount", "total"),
|
| 35 |
+
("Tax Label", "tax_label"),
|
| 36 |
+
]
|
| 37 |
+
|
| 38 |
+
|
| 39 |
+
def extract(base: str, pdf: Path, keys: list[str]) -> tuple[dict, float]:
|
| 40 |
+
started = time.perf_counter()
|
| 41 |
+
proc = subprocess.run(
|
| 42 |
+
["curl.exe", "-sS", "-X", "POST", f"{base}/v1/extract",
|
| 43 |
+
"-F", f"file=@{pdf}", "-F", "keys=" + ",".join(keys)],
|
| 44 |
+
capture_output=True, text=True,
|
| 45 |
+
)
|
| 46 |
+
wall = (time.perf_counter() - started) * 1000
|
| 47 |
+
if proc.returncode != 0:
|
| 48 |
+
raise RuntimeError(proc.stderr.strip() or "curl failed")
|
| 49 |
+
return json.loads(proc.stdout), wall
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
def main() -> int:
|
| 53 |
+
ap = argparse.ArgumentParser()
|
| 54 |
+
ap.add_argument("--base", default="http://127.0.0.1:7860")
|
| 55 |
+
ap.add_argument("--corpus", type=Path, required=True)
|
| 56 |
+
ap.add_argument("--limit", type=int, default=1)
|
| 57 |
+
args = ap.parse_args()
|
| 58 |
+
|
| 59 |
+
manifest = json.loads((args.corpus / "ground_truth.json").read_text())
|
| 60 |
+
records = manifest[: args.limit] if args.limit else manifest
|
| 61 |
+
keys = [k for k, _ in KEY_MAP]
|
| 62 |
+
|
| 63 |
+
print(f"target {args.base}")
|
| 64 |
+
print(f"model impira/layoutlm-invoices keys={len(keys)} "
|
| 65 |
+
f"documents={len(records)}\n")
|
| 66 |
+
|
| 67 |
+
per_key: dict[str, list[int]] = {k: [] for k in keys}
|
| 68 |
+
latencies: list[float] = []
|
| 69 |
+
|
| 70 |
+
for record in records:
|
| 71 |
+
body, wall = extract(args.base, Path(record["path"]), keys)
|
| 72 |
+
latencies.append(body.get("latency_ms", 0.0))
|
| 73 |
+
print(f"{record['document_id']} variant={record['variant']} "
|
| 74 |
+
f"source={body.get('source')} words={body.get('word_count')} "
|
| 75 |
+
f"pages={body.get('pages_processed')}")
|
| 76 |
+
print(f"server={body.get('latency_ms', 0):.0f} ms "
|
| 77 |
+
f"wall={wall:.0f} ms (includes model load on first call)\n")
|
| 78 |
+
|
| 79 |
+
by_key = {f["key"]: f for f in body.get("fields", [])}
|
| 80 |
+
print(f" {'KEY':<14} {'EXPECTED':<26} {'STATUS':<15} "
|
| 81 |
+
f"{'GOT':<28} CONF")
|
| 82 |
+
print(" " + "-" * 92)
|
| 83 |
+
for key, truth in KEY_MAP:
|
| 84 |
+
expected = str(record[truth])
|
| 85 |
+
field = by_key.get(key, {})
|
| 86 |
+
got = (field.get("value") or "").strip()
|
| 87 |
+
ok = got == expected.strip()
|
| 88 |
+
per_key[key].append(1 if ok else 0)
|
| 89 |
+
print(f" {key:<14} {expected[:26]:<26} "
|
| 90 |
+
f"{str(field.get('status')):<15} "
|
| 91 |
+
f"{(got[:26] + (' OK' if ok else ' BAD')):<28} "
|
| 92 |
+
f"{field.get('confidence', 0):.4f}")
|
| 93 |
+
print(" " + "-" * 92 + "\n")
|
| 94 |
+
|
| 95 |
+
total_hit = sum(sum(v) for v in per_key.values())
|
| 96 |
+
total_n = sum(len(v) for v in per_key.values())
|
| 97 |
+
|
| 98 |
+
print("PER-KEY EXACT MATCH")
|
| 99 |
+
print("-" * 46)
|
| 100 |
+
for key, _ in KEY_MAP:
|
| 101 |
+
scores = per_key[key]
|
| 102 |
+
if scores:
|
| 103 |
+
acc = 100 * sum(scores) / len(scores)
|
| 104 |
+
print(f" {key:<16} {acc:6.1f}% ({sum(scores)}/{len(scores)})")
|
| 105 |
+
if total_n:
|
| 106 |
+
print(f"\n {'OVERALL':<16} {100*total_hit/total_n:6.1f}% "
|
| 107 |
+
f"({total_hit}/{total_n})")
|
| 108 |
+
|
| 109 |
+
if latencies:
|
| 110 |
+
print("\nLATENCY (whole document, all keys)")
|
| 111 |
+
print("-" * 46)
|
| 112 |
+
print(f" mean {statistics.mean(latencies):8.0f} ms")
|
| 113 |
+
print(f" min {min(latencies):8.0f} ms")
|
| 114 |
+
print(f" max {max(latencies):8.0f} ms")
|
| 115 |
+
print(f" per key {statistics.mean(latencies)/len(keys):6.0f} ms")
|
| 116 |
+
|
| 117 |
+
return 0
|
| 118 |
+
|
| 119 |
+
|
| 120 |
+
if __name__ == "__main__":
|
| 121 |
+
raise SystemExit(main())
|
docqa/bench/check_deps.py
ADDED
|
@@ -0,0 +1,158 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Verify requirements.txt covers every third-party import in docqa/.
|
| 3 |
+
|
| 4 |
+
Run in CI or after adding a module:
|
| 5 |
+
|
| 6 |
+
python bench/check_deps.py
|
| 7 |
+
|
| 8 |
+
Exits non-zero if an import is unlisted, so a missing dependency is caught at
|
| 9 |
+
build time rather than as an ImportError on the first request that happens to
|
| 10 |
+
hit that code path.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
import ast
|
| 16 |
+
import re
|
| 17 |
+
import sys
|
| 18 |
+
from pathlib import Path
|
| 19 |
+
|
| 20 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 21 |
+
REPO = ROOT.parent
|
| 22 |
+
|
| 23 |
+
# import name -> pip distribution name, where they differ
|
| 24 |
+
DISTRIBUTION = {
|
| 25 |
+
"PIL": "pillow",
|
| 26 |
+
"fitz": "pymupdf",
|
| 27 |
+
"sklearn": "scikit-learn",
|
| 28 |
+
"yaml": "pyyaml",
|
| 29 |
+
"cv2": "opencv-python",
|
| 30 |
+
"bs4": "beautifulsoup4",
|
| 31 |
+
"dotenv": "python-dotenv",
|
| 32 |
+
}
|
| 33 |
+
|
| 34 |
+
# Imported only inside functions guarded by try/except ImportError, or only by
|
| 35 |
+
# optional benchmark paths. Still listed in requirements.txt, but a missing
|
| 36 |
+
# one is a warning rather than a hard failure.
|
| 37 |
+
OPTIONAL = {"onnx", "onnxruntime", "sklearn"}
|
| 38 |
+
|
| 39 |
+
# Runtime plugin dependencies that no AST scan of our own code can discover,
|
| 40 |
+
# because the import happens inside the framework. FastAPI resolves these
|
| 41 |
+
# lazily: the app imports fine and the endpoint 500s on first use. Keep this in
|
| 42 |
+
# step with any File()/Form()/UploadFile declaration in api.py.
|
| 43 |
+
FRAMEWORK_PLUGINS = {
|
| 44 |
+
"python-multipart": "FastAPI File()/Form(); app starts but /v1/extract 500s",
|
| 45 |
+
}
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
def parse_requirements() -> dict[str, str]:
|
| 49 |
+
"""Distribution name -> pinned version, from requirements.txt."""
|
| 50 |
+
out: dict[str, str] = {}
|
| 51 |
+
for line in (REPO / "requirements.txt").read_text(encoding="utf-8").splitlines():
|
| 52 |
+
line = line.split("#", 1)[0].strip()
|
| 53 |
+
if not line or line.startswith("-"):
|
| 54 |
+
continue
|
| 55 |
+
match = re.match(r"^([A-Za-z0-9._-]+)(?:\[[^\]]*\])?==(.+)$", line)
|
| 56 |
+
if match:
|
| 57 |
+
out[match.group(1).lower().replace("_", "-")] = match.group(2)
|
| 58 |
+
return out
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
def collect_imports() -> set[str]:
|
| 62 |
+
found: set[str] = set()
|
| 63 |
+
for path in ROOT.rglob("*.py"):
|
| 64 |
+
tree = ast.parse(path.read_text(encoding="utf-8"))
|
| 65 |
+
for node in ast.walk(tree):
|
| 66 |
+
if isinstance(node, ast.Import):
|
| 67 |
+
for alias in node.names:
|
| 68 |
+
found.add(alias.name.split(".")[0])
|
| 69 |
+
elif isinstance(node, ast.ImportFrom) and node.level == 0 and node.module:
|
| 70 |
+
found.add(node.module.split(".")[0])
|
| 71 |
+
return found
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def verify_installed(declared: dict[str, str]) -> list[str]:
|
| 75 |
+
"""Actually import each declared distribution.
|
| 76 |
+
|
| 77 |
+
An AST scan only sees what docqa/ imports directly. It cannot see what a
|
| 78 |
+
framework pulls in at request time, which is precisely how python-multipart
|
| 79 |
+
was missed: fastapi imports cleanly without it and only fails when
|
| 80 |
+
File()/Form() is exercised. Importing the real modules closes that gap.
|
| 81 |
+
"""
|
| 82 |
+
import importlib
|
| 83 |
+
|
| 84 |
+
problems: list[str] = []
|
| 85 |
+
for dist in sorted(declared):
|
| 86 |
+
module = MODULE_FOR_DIST.get(dist)
|
| 87 |
+
if module is None:
|
| 88 |
+
continue
|
| 89 |
+
try:
|
| 90 |
+
importlib.import_module(module)
|
| 91 |
+
except ImportError as exc:
|
| 92 |
+
problems.append(f"{dist} (import {module}): {exc}")
|
| 93 |
+
return problems
|
| 94 |
+
|
| 95 |
+
|
| 96 |
+
# distribution name -> importable module, where they differ
|
| 97 |
+
MODULE_FOR_DIST = {
|
| 98 |
+
"pillow": "PIL",
|
| 99 |
+
"pymupdf": "fitz",
|
| 100 |
+
"scikit-learn": "sklearn",
|
| 101 |
+
"python-multipart": "multipart",
|
| 102 |
+
"python-dotenv": "dotenv",
|
| 103 |
+
"pyyaml": "yaml",
|
| 104 |
+
}
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def main() -> int:
|
| 108 |
+
declared = parse_requirements()
|
| 109 |
+
imported = collect_imports()
|
| 110 |
+
|
| 111 |
+
stdlib = set(sys.stdlib_module_names)
|
| 112 |
+
local = {"docxextract"}
|
| 113 |
+
missing: list[str] = []
|
| 114 |
+
optional_missing: list[str] = []
|
| 115 |
+
|
| 116 |
+
for name in sorted(imported):
|
| 117 |
+
if name in stdlib or name in local:
|
| 118 |
+
continue
|
| 119 |
+
dist = DISTRIBUTION.get(name, name).lower().replace("_", "-")
|
| 120 |
+
if dist in declared:
|
| 121 |
+
continue
|
| 122 |
+
(optional_missing if name in OPTIONAL else missing).append(f"{name} -> {dist}")
|
| 123 |
+
|
| 124 |
+
# Plugins the framework needs at request time, invisible to the scan.
|
| 125 |
+
for dist, why in FRAMEWORK_PLUGINS.items():
|
| 126 |
+
if dist not in declared:
|
| 127 |
+
missing.append(f"{dist} <- {why}")
|
| 128 |
+
|
| 129 |
+
print(f"requirements.txt declares {len(declared)} distributions")
|
| 130 |
+
print(f"docqa/ imports {len(imported)} modules "
|
| 131 |
+
f"({len([m for m in imported if m not in stdlib])} non-stdlib)\n")
|
| 132 |
+
|
| 133 |
+
if optional_missing:
|
| 134 |
+
print("OPTIONAL, not declared (guarded imports / optional paths):")
|
| 135 |
+
for item in optional_missing:
|
| 136 |
+
print(f" {item}")
|
| 137 |
+
print()
|
| 138 |
+
|
| 139 |
+
broken = verify_installed(declared)
|
| 140 |
+
if broken:
|
| 141 |
+
print("DECLARED BUT NOT IMPORTABLE in this environment:")
|
| 142 |
+
for item in broken:
|
| 143 |
+
print(f" {item}")
|
| 144 |
+
print()
|
| 145 |
+
|
| 146 |
+
if missing or broken:
|
| 147 |
+
if missing:
|
| 148 |
+
print("MISSING, must be added to requirements.txt:")
|
| 149 |
+
for item in missing:
|
| 150 |
+
print(f" {item}")
|
| 151 |
+
return 1
|
| 152 |
+
|
| 153 |
+
print("OK: every non-stdlib import is declared and importable.")
|
| 154 |
+
return 0
|
| 155 |
+
|
| 156 |
+
|
| 157 |
+
if __name__ == "__main__":
|
| 158 |
+
raise SystemExit(main())
|
requirements.txt
CHANGED
|
@@ -1,30 +1,49 @@
|
|
| 1 |
-
#
|
| 2 |
-
#
|
| 3 |
-
#
|
| 4 |
-
#
|
| 5 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
--extra-index-url https://download.pytorch.org/whl/cpu
|
| 7 |
|
|
|
|
| 8 |
fastapi==0.115.5
|
| 9 |
uvicorn[standard]==0.34.0
|
| 10 |
pydantic==2.13.4
|
| 11 |
pydantic-settings==2.15.0
|
| 12 |
orjson==3.11.9
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
|
| 14 |
-
# model
|
| 15 |
torch==2.5.1+cpu
|
| 16 |
transformers==4.50.2
|
| 17 |
-
|
|
|
|
|
|
|
|
|
|
| 18 |
|
| 19 |
-
# parsing / OCR
|
| 20 |
pdfplumber==0.11.9
|
| 21 |
pytesseract==0.3.13
|
| 22 |
-
|
|
|
|
|
|
|
| 23 |
|
| 24 |
-
#
|
| 25 |
-
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
-
# tests
|
| 28 |
pytest==8.3.4
|
| 29 |
-
|
| 30 |
-
pytest-asyncio==0.24.0
|
|
|
|
| 1 |
+
# Every third-party module imported anywhere under docqa/ appears here.
|
| 2 |
+
# Verified by AST-scanning all imports, not by guessing -- see
|
| 3 |
+
# bench/check_deps.py, which fails if an import is unlisted.
|
| 4 |
+
#
|
| 5 |
+
# Pinned exactly. Loose pins broke this build more than once:
|
| 6 |
+
# - torch must come from the CPU wheel index. The default PyPI torch pulls
|
| 7 |
+
# multi-GB CUDA libraries and exceeds the Space image limit.
|
| 8 |
+
# - numpy must stay on 1.x. onnxruntime/onnx pull numpy 2.x, which breaks
|
| 9 |
+
# scipy and scikit-learn used by the corpus generator.
|
| 10 |
--extra-index-url https://download.pytorch.org/whl/cpu
|
| 11 |
|
| 12 |
+
# --- web service ---
|
| 13 |
fastapi==0.115.5
|
| 14 |
uvicorn[standard]==0.34.0
|
| 15 |
pydantic==2.13.4
|
| 16 |
pydantic-settings==2.15.0
|
| 17 |
orjson==3.11.9
|
| 18 |
+
httpx==0.28.1
|
| 19 |
+
# REQUIRED by FastAPI at runtime for any endpoint declaring File() or Form().
|
| 20 |
+
# It is not imported by our own code, so an AST scan of docqa/ cannot find it:
|
| 21 |
+
# FastAPI raises at request time ("Form data requires python-multipart"), not
|
| 22 |
+
# at import time, which means the app starts cleanly and then 500s on
|
| 23 |
+
# /v1/extract. Added explicitly after that exact failure.
|
| 24 |
+
python-multipart==0.0.20
|
| 25 |
|
| 26 |
+
# --- model ---
|
| 27 |
torch==2.5.1+cpu
|
| 28 |
transformers==4.50.2
|
| 29 |
+
huggingface-hub==0.28.1
|
| 30 |
+
onnxruntime==1.20.1
|
| 31 |
+
numpy==1.26.4
|
| 32 |
+
scipy==1.10.1
|
| 33 |
|
| 34 |
+
# --- document parsing / OCR ---
|
| 35 |
pdfplumber==0.11.9
|
| 36 |
pytesseract==0.3.13
|
| 37 |
+
pillow==10.3.0
|
| 38 |
+
# fitz: rasterises scanned pages for the OCR fallback
|
| 39 |
+
pymupdf==1.25.1
|
| 40 |
|
| 41 |
+
# --- test corpus generation ---
|
| 42 |
+
# reportlab draws the synthetic invoices in tests/corpus.py; scikit-learn is
|
| 43 |
+
# used by bench/compare_backends.py for scoring.
|
| 44 |
+
reportlab==3.6.13
|
| 45 |
+
scikit-learn==1.6.1
|
| 46 |
|
| 47 |
+
# --- tests ---
|
| 48 |
pytest==8.3.4
|
| 49 |
+
pytest-asyncio==1.3.0
|
|
|