--- license: apache-2.0 library_name: onnx pipeline_tag: image-classification tags: - medical-imaging - modality-classification - dicom - onnx - int8 --- # ModalityScan Single-label imaging-modality classification for medical images. One shared ConvNet backbone, two softmax heads (`family`, `modality`), shipped as a static-INT8 ONNX graph. Predicts **what kind of image this is**, not what is in it. Anatomy is a different model. ## What it outputs ``` input pixel_values float32 [batch, 3, 224, 224] NCHW, ImageNet-normalized, letterboxed outputs family_logits float32 [batch, 6] modality_logits float32 [batch, 20] ``` Softmax, temperature and thresholds live **outside** the graph so calibration and policy stay configuration. `modality-taxonomy.json` travels with the weights - a 20-element logits vector is 20 meaningless numbers without it. ## Classes | class | family | tier | train frames | sources | P | R | F1 | support | |---|---|---|---|---|---|---|---|---| | `CT` | cross_sectional | model_trained | 5,996 | 3 | 0.8469 | 0.8059 | 0.8259 | 2,684 | | `MR` | cross_sectional | experimental | 404 | 1 | 0.2476 | 1.0 | 0.3969 | 52 | | `US` | ultrasound | model_trained | 3,779 | 2 | 0.861 | 0.9558 | 0.9059 | 3,123 | | `CR` | projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `DX` | projection_radiography | model_trained | 6,000 | 2 | 0.9956 | 0.9814 | 0.9884 | 3,442 | | `MG` | projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `RF` | projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 | | `XA` | projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 | | `PT` | cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `NM` | cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `OCT` | optical | experimental | 6,000 | 1 | 0.9977 | 1.0 | 0.9989 | 881 | | `FUNDUS` | optical | experimental | 1,600 | 1 | 1.0 | 1.0 | 1.0 | 246 | | `DERMOSCOPY` | optical | experimental | 6,000 | 1 | 0.5773 | 0.9943 | 0.7305 | 879 | | `ENDOSCOPY` | optical | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `HISTOPATHOLOGY` | microscopy | experimental | 6,000 | 1 | 0.2853 | 1.0 | 0.4439 | 926 | | `MICROSCOPY` | microscopy | not_trained | 6,000 | 2 | n/a | n/a | n/a | 3,479 | | `PHOTO` | non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `RENDER_3D` | non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 | | `DOCUMENT` | non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 | | `WAVEFORM` | non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 | **Tier is a statement about evidence, not difficulty.** `model_trained` means the corpus had at least two independent sources and enough patients to make the number falsifiable. `experimental` means it did not. `not_trained` classes are never emitted by the decoder, so no metrics are reported for them. Some `not_trained` classes *do* have training frames. They were withdrawn after measurement rather than never attempted, and the decoder never emits them: * **`MICROSCOPY` is not claimed.** Withdrawn from the claim after measurement. Held-out-source recall was 97/3,479 = 2.8%, with 60% of held-out frames predicted HISTOPATHOLOGY. Its two sources are not interchangeable samples of one modality: TissueMNIST (trained on) is grayscale kidney cortex, BloodMNIST (held out) is colour-stained blood smears, so training never saw colour microscopy. The slot stays in the 20-wide output because that is a published contract; the decoder never emits it. ## Results Headline: **macro-F1 0.9067** over the `model_trained` classes (`CT`, `DX`, `US`), balanced policy, on the held-out test split. Accuracy on frames whose true class is one the model claims: **0.9218**. Accuracy over *every* labelled frame including classes the model does not claim: **0.7446** - the difference is almost entirely withdrawn-class frames, which are counted as errors here by design. | benchmark tier | rows | accuracy | macro-F1 | |---|---|---|---| | held_out_source | 11,000 | 0.6618 | 0.9021 | | in_distribution | 4,712 | 0.9378 | 0.9109 | In-distribution macro-F1 is 0.9109 and held-out-source macro-F1 is 0.9021. The gap is small, which is the evidence that the model reads the image rather than the dataset it came from. Coarse and fine accuracy are identical here: no collapse group has two populated classes in this corpus, so the group step of the cascade never fires. The cascade is still live - it is simply untested by this evaluation. Open-set rejection: **not measured**. The corpus contains no `open_set` rows, so the energy cut was fitted for false-rejection rate only and its true detection rate is unknown. `out_of_distribution: true` is provisional - do not rely on it in production until it has been measured against real non-diagnostic material. Expected calibration error on the test split: **0.0991** (calibrated). Temperature was fitted on validation; the test split is harder because it carries every held-out source, so this is the pessimistic figure. ## The abstention cascade When no single class clears its threshold the decoder falls back, and always reports how far it fell: | `granularity` | meaning | |---|---| | `modality` | a specific modality, e.g. `CT` | | `group` | a confusable group, e.g. `XR_PLAIN` when CR and DX are not separable | | `family` | an acquisition family, e.g. `projection_radiography` | | `none` | abstained, or rejected as out-of-distribution | A caller routing to a CT pipeline should require `granularity == "modality"`. A caller deciding whether an image is worth processing at all is satisfied by `family`. ## Limitations * **MR sequence, contrast phase and anatomy are out of scope.** This model says `MR`; it does not say `T2 FLAIR`, and it does not say `brain`. * **Secondary Capture is not a modality.** An SC object is a screenshot of something else. The model sees pixels and will answer for what was on screen. * **Compound figures have no single answer.** A literature figure with CT and MR panels side by side violates the single-label assumption this model is built on. * **Digitised screen-film mammography is over-represented** in the `MG` training data relative to modern full-field digital mammography. * Scores are calibrated. Read the measured precision and recall above, not the score, when choosing a policy. ## Quick start (Python + onnxruntime) No clone, no framework - the artifact and its config files are all you need. ```bash pip install onnxruntime pillow numpy huggingface_hub ``` ```python import json, numpy as np, onnxruntime as ort from PIL import Image from huggingface_hub import hf_hub_download REPO = "vectorsense/modalityscan" sess = ort.InferenceSession(hf_hub_download(REPO, "model.int8.onnx"), providers=["CPUExecutionProvider"]) tax = json.load(open(hf_hub_download(REPO, "modality-taxonomy.json"), encoding="utf-8")) book = json.load(open(hf_hub_download(REPO, "modality-thresholds.json"), encoding="utf-8")) pre = json.load(open(hf_hub_download(REPO, "preprocessor.json"), encoding="utf-8")) MODALITY = tax["levels"]["modality"]["index"] FAMILY = tax["levels"]["family"]["index"] TIER = {label: tier for tier in ("model_trained", "experimental", "not_trained") for label in tax["tiers"][tier]["types"]} mean = np.array(pre["image_mean"], np.float32)[:, None, None] std = np.array(pre["image_std"], np.float32)[:, None, None] def pixel_values(path, size=224): # Letterbox, never centre-crop. The field-of-view outline - ultrasound sector, circular # fundus aperture, endoscope vignette - is a top-tier modality cue; cropping discards it. image = Image.open(path).convert("RGB") scale = size / max(image.size) small = image.resize((max(1, round(image.width * scale)), max(1, round(image.height * scale))), Image.BILINEAR) canvas = Image.new("RGB", (size, size), (0, 0, 0)) canvas.paste(small, ((size - small.width) // 2, (size - small.height) // 2)) array = np.asarray(canvas, np.float32) / 255.0 return ((array.transpose(2, 0, 1) - mean) / std)[None] def softmax(x): e = np.exp(x - x.max()) return e / e.sum() family_logits, modality_logits = sess.run( ["family_logits", "modality_logits"], {"pixel_values": pixel_values("frame.png")} ) # Temperature lives outside the graph so calibration stays configuration, not a re-export. T = book["calibration"]["temperature"] p_mod = softmax(modality_logits[0] / T["modality"]) p_fam = softmax(family_logits[0] / T["family"]) policy = book["default_policy"] floor = book["policies"][policy]["modality"] over = book["per_class_overrides"].get(policy, {}) # Energy over logits, not softmax: softmax is confident on garbage by construction. energy = -float(np.log(np.exp(modality_logits[0] / T["modality"]).sum())) if energy > book["ood"]["energy_threshold"][policy]: print("out of distribution ->", book["ood"]["fallback_label"]) else: # `not_trained` classes keep their slot because the output width is a published contract, # but they are not claims. A raw argmax will eventually return one of them. best = next(i for i in np.argsort(-p_mod) if TIER.get(MODALITY[i]) != "not_trained") label, score = MODALITY[best], float(p_mod[best]) if score >= over.get(label, floor): print(f"modality: {label} {score:.3f} tier={TIER[label]}") else: print(f"family: {FAMILY[int(p_fam.argmax())]} {p_fam.max():.3f} (cascade fell back)") ``` Two lines carry most of the correctness. **Letterbox rather than centre-crop** - the field-of-view outline is signal, not padding. **Skip `not_trained` classes** - a raw `argmax` will eventually hand you a label this model does not stand behind. ## Run it as a REST API The companion repo ships a FastAPI service with the cascade, the calibrated thresholds, the energy rejection and series aggregation already wired up. ```bash pip install fastapi "uvicorn[standard]" onnxruntime pillow numpy uvicorn service.app:app --port 5003 # auto-loads model.int8.onnx from ./models ``` ```bash curl -s http://127.0.0.1:5003/classify_modality \ -H "Content-Type: application/json" \ -d '{"frames":[{"id":"series-1/f000","image_base64":"iVBORw0KGgo..."}], "options":{"policy":"balanced"}}' ``` The same call from Python - send a PNG or JPEG, get parsed JSON back: ```python import base64, requests URL = "http://127.0.0.1:5003/classify_modality" def classify(path, policy="balanced"): payload = { "frames": [{"id": "series-1/f000", "image_base64": base64.b64encode(open(path, "rb").read()).decode()}], "options": {"policy": policy, "top_k": 3}, } response = requests.post(URL, json=payload, timeout=30) response.raise_for_status() # unknown request fields return 422, never a silent default return response.json() frame = classify("frame.png")["results"][0] print(frame["modality"]["label"], frame["modality"]["granularity"], frame["modality"]["score"]) print("out of distribution:", frame["out_of_distribution"]) ``` A whole DICOM series in one request, which is the cheapest accuracy in the serving path: ```python payload = { "frames": [{"id": f"series-1/f{i:03d}", "image_base64": base64.b64encode(open(p, "rb").read()).decode()} for i, p in enumerate(frame_paths)], "options": {"policy": "balanced", "aggregate_by": "series"}, } print(requests.post(URL, json=payload, timeout=60).json()["aggregate"]) ``` Endpoints: `GET /`, `POST /classify_modality`, `GET /health/ready`, `GET /metadata`. `GET /metadata` reports the tier of every class at runtime, so a caller can tell a claimed class from an experimental one without reading this card. Prefer the service over the raw-ONNX snippet unless you have a reason not to: it already applies the calibrated temperature, the per-class thresholds, the abstention cascade and the rejection path. The snippet reimplements those in miniature and will drift from the config. ## Usage via the library ```python from modalityscan import ModalityClassifier classifier = ModalityClassifier(model_dir="models") results, series = classifier.classify( [{"id": "series-1/f000", "image_path": "frame.png"}], policy="balanced", aggregate_by="series", ) print(results[0].modality.label, results[0].modality.granularity) ``` Modality is constant across a DICOM series, so passing several frames with a shared `series_id` and `aggregate_by="series"` is the cheapest accuracy available in the serving path. ## Attribution Training data is used under the licences recorded in `ATTRIBUTION.md`. Repository: `vectorsense/modalityscan`.