modalityscan / README.md
vectorsense's picture
Upload README.md with huggingface_hub
2a85d5a verified
|
Raw History Blame Contribute Delete
13.1 kB
metadata
license: apache-2.0
library_name: onnx
pipeline_tag: image-classification
tags:
  - medical-imaging
  - modality-classification
  - dicom
  - onnx
  - int8

ModalityScan

Single-label imaging-modality classification for medical images. One shared ConvNet backbone, two softmax heads (family, modality), shipped as a static-INT8 ONNX graph.

Predicts what kind of image this is, not what is in it. Anatomy is a different model.

What it outputs

input   pixel_values     float32  [batch, 3, 224, 224]   NCHW, ImageNet-normalized, letterboxed
outputs family_logits    float32  [batch, 6]
        modality_logits  float32  [batch, 20]

Softmax, temperature and thresholds live outside the graph so calibration and policy stay configuration. modality-taxonomy.json travels with the weights - a 20-element logits vector is 20 meaningless numbers without it.

Classes

class family tier train frames sources P R F1 support
CT cross_sectional model_trained 5,996 3 0.8469 0.8059 0.8259 2,684
MR cross_sectional experimental 404 1 0.2476 1.0 0.3969 52
US ultrasound model_trained 3,779 2 0.861 0.9558 0.9059 3,123
CR projection_radiography experimental 0 0 0.0 0.0 0.0 0
DX projection_radiography model_trained 6,000 2 0.9956 0.9814 0.9884 3,442
MG projection_radiography experimental 0 0 0.0 0.0 0.0 0
RF projection_radiography not_trained 0 0 n/a n/a n/a 0
XA projection_radiography not_trained 0 0 n/a n/a n/a 0
PT cross_sectional experimental 0 0 0.0 0.0 0.0 0
NM cross_sectional experimental 0 0 0.0 0.0 0.0 0
OCT optical experimental 6,000 1 0.9977 1.0 0.9989 881
FUNDUS optical experimental 1,600 1 1.0 1.0 1.0 246
DERMOSCOPY optical experimental 6,000 1 0.5773 0.9943 0.7305 879
ENDOSCOPY optical experimental 0 0 0.0 0.0 0.0 0
HISTOPATHOLOGY microscopy experimental 6,000 1 0.2853 1.0 0.4439 926
MICROSCOPY microscopy not_trained 6,000 2 n/a n/a n/a 3,479
PHOTO non_diagnostic experimental 0 0 0.0 0.0 0.0 0
RENDER_3D non_diagnostic experimental 0 0 0.0 0.0 0.0 0
DOCUMENT non_diagnostic not_trained 0 0 n/a n/a n/a 0
WAVEFORM non_diagnostic not_trained 0 0 n/a n/a n/a 0

Tier is a statement about evidence, not difficulty. model_trained means the corpus had at least two independent sources and enough patients to make the number falsifiable. experimental means it did not. not_trained classes are never emitted by the decoder, so no metrics are reported for them.

Some not_trained classes do have training frames. They were withdrawn after measurement rather than never attempted, and the decoder never emits them:

  • MICROSCOPY is not claimed. Withdrawn from the claim after measurement. Held-out-source recall was 97/3,479 = 2.8%, with 60% of held-out frames predicted HISTOPATHOLOGY. Its two sources are not interchangeable samples of one modality: TissueMNIST (trained on) is grayscale kidney cortex, BloodMNIST (held out) is colour-stained blood smears, so training never saw colour microscopy. The slot stays in the 20-wide output because that is a published contract; the decoder never emits it.

Results

Headline: macro-F1 0.9067 over the model_trained classes (CT, DX, US), balanced policy, on the held-out test split.

Accuracy on frames whose true class is one the model claims: 0.9218. Accuracy over every labelled frame including classes the model does not claim: 0.7446 - the difference is almost entirely withdrawn-class frames, which are counted as errors here by design.

benchmark tier rows accuracy macro-F1
held_out_source 11,000 0.6618 0.9021
in_distribution 4,712 0.9378 0.9109

In-distribution macro-F1 is 0.9109 and held-out-source macro-F1 is 0.9021. The gap is small, which is the evidence that the model reads the image rather than the dataset it came from.

Coarse and fine accuracy are identical here: no collapse group has two populated classes in this corpus, so the group step of the cascade never fires. The cascade is still live - it is simply untested by this evaluation.

Open-set rejection: not measured. The corpus contains no open_set rows, so the energy cut was fitted for false-rejection rate only and its true detection rate is unknown. out_of_distribution: true is provisional - do not rely on it in production until it has been measured against real non-diagnostic material.

Expected calibration error on the test split: 0.0991 (calibrated). Temperature was fitted on validation; the test split is harder because it carries every held-out source, so this is the pessimistic figure.

The abstention cascade

When no single class clears its threshold the decoder falls back, and always reports how far it fell:

granularity meaning
modality a specific modality, e.g. CT
group a confusable group, e.g. XR_PLAIN when CR and DX are not separable
family an acquisition family, e.g. projection_radiography
none abstained, or rejected as out-of-distribution

A caller routing to a CT pipeline should require granularity == "modality". A caller deciding whether an image is worth processing at all is satisfied by family.

Limitations

  • MR sequence, contrast phase and anatomy are out of scope. This model says MR; it does not say T2 FLAIR, and it does not say brain.
  • Secondary Capture is not a modality. An SC object is a screenshot of something else. The model sees pixels and will answer for what was on screen.
  • Compound figures have no single answer. A literature figure with CT and MR panels side by side violates the single-label assumption this model is built on.
  • Digitised screen-film mammography is over-represented in the MG training data relative to modern full-field digital mammography.
  • Scores are calibrated. Read the measured precision and recall above, not the score, when choosing a policy.

Quick start (Python + onnxruntime)

No clone, no framework - the artifact and its config files are all you need.

pip install onnxruntime pillow numpy huggingface_hub
import json, numpy as np, onnxruntime as ort
from PIL import Image
from huggingface_hub import hf_hub_download

REPO = "vectorsense/modalityscan"
sess = ort.InferenceSession(hf_hub_download(REPO, "model.int8.onnx"),
                            providers=["CPUExecutionProvider"])
tax  = json.load(open(hf_hub_download(REPO, "modality-taxonomy.json"), encoding="utf-8"))
book = json.load(open(hf_hub_download(REPO, "modality-thresholds.json"), encoding="utf-8"))
pre  = json.load(open(hf_hub_download(REPO, "preprocessor.json"), encoding="utf-8"))

MODALITY = tax["levels"]["modality"]["index"]
FAMILY   = tax["levels"]["family"]["index"]
TIER     = {label: tier
            for tier in ("model_trained", "experimental", "not_trained")
            for label in tax["tiers"][tier]["types"]}

mean = np.array(pre["image_mean"], np.float32)[:, None, None]
std  = np.array(pre["image_std"],  np.float32)[:, None, None]

def pixel_values(path, size=224):
    # Letterbox, never centre-crop. The field-of-view outline - ultrasound sector, circular
    # fundus aperture, endoscope vignette - is a top-tier modality cue; cropping discards it.
    image = Image.open(path).convert("RGB")
    scale = size / max(image.size)
    small = image.resize((max(1, round(image.width * scale)),
                          max(1, round(image.height * scale))), Image.BILINEAR)
    canvas = Image.new("RGB", (size, size), (0, 0, 0))
    canvas.paste(small, ((size - small.width) // 2, (size - small.height) // 2))
    array = np.asarray(canvas, np.float32) / 255.0
    return ((array.transpose(2, 0, 1) - mean) / std)[None]

def softmax(x):
    e = np.exp(x - x.max())
    return e / e.sum()

family_logits, modality_logits = sess.run(
    ["family_logits", "modality_logits"], {"pixel_values": pixel_values("frame.png")}
)

# Temperature lives outside the graph so calibration stays configuration, not a re-export.
T = book["calibration"]["temperature"]
p_mod = softmax(modality_logits[0] / T["modality"])
p_fam = softmax(family_logits[0] / T["family"])

policy = book["default_policy"]
floor  = book["policies"][policy]["modality"]
over   = book["per_class_overrides"].get(policy, {})

# Energy over logits, not softmax: softmax is confident on garbage by construction.
energy = -float(np.log(np.exp(modality_logits[0] / T["modality"]).sum()))
if energy > book["ood"]["energy_threshold"][policy]:
    print("out of distribution ->", book["ood"]["fallback_label"])
else:
    # `not_trained` classes keep their slot because the output width is a published contract,
    # but they are not claims. A raw argmax will eventually return one of them.
    best = next(i for i in np.argsort(-p_mod) if TIER.get(MODALITY[i]) != "not_trained")
    label, score = MODALITY[best], float(p_mod[best])
    if score >= over.get(label, floor):
        print(f"modality: {label} {score:.3f}  tier={TIER[label]}")
    else:
        print(f"family:   {FAMILY[int(p_fam.argmax())]} {p_fam.max():.3f}  (cascade fell back)")

Two lines carry most of the correctness. Letterbox rather than centre-crop - the field-of-view outline is signal, not padding. Skip not_trained classes - a raw argmax will eventually hand you a label this model does not stand behind.

Run it as a REST API

The companion repo ships a FastAPI service with the cascade, the calibrated thresholds, the energy rejection and series aggregation already wired up.

pip install fastapi "uvicorn[standard]" onnxruntime pillow numpy
uvicorn service.app:app --port 5003      # auto-loads model.int8.onnx from ./models
curl -s http://127.0.0.1:5003/classify_modality \
  -H "Content-Type: application/json" \
  -d '{"frames":[{"id":"series-1/f000","image_base64":"iVBORw0KGgo..."}],
       "options":{"policy":"balanced"}}'

The same call from Python - send a PNG or JPEG, get parsed JSON back:

import base64, requests

URL = "http://127.0.0.1:5003/classify_modality"

def classify(path, policy="balanced"):
    payload = {
        "frames": [{"id": "series-1/f000",
                    "image_base64": base64.b64encode(open(path, "rb").read()).decode()}],
        "options": {"policy": policy, "top_k": 3},
    }
    response = requests.post(URL, json=payload, timeout=30)
    response.raise_for_status()     # unknown request fields return 422, never a silent default
    return response.json()

frame = classify("frame.png")["results"][0]
print(frame["modality"]["label"], frame["modality"]["granularity"], frame["modality"]["score"])
print("out of distribution:", frame["out_of_distribution"])

A whole DICOM series in one request, which is the cheapest accuracy in the serving path:

payload = {
    "frames": [{"id": f"series-1/f{i:03d}",
                "image_base64": base64.b64encode(open(p, "rb").read()).decode()}
               for i, p in enumerate(frame_paths)],
    "options": {"policy": "balanced", "aggregate_by": "series"},
}
print(requests.post(URL, json=payload, timeout=60).json()["aggregate"])

Endpoints: GET /, POST /classify_modality, GET /health/ready, GET /metadata. GET /metadata reports the tier of every class at runtime, so a caller can tell a claimed class from an experimental one without reading this card.

Prefer the service over the raw-ONNX snippet unless you have a reason not to: it already applies the calibrated temperature, the per-class thresholds, the abstention cascade and the rejection path. The snippet reimplements those in miniature and will drift from the config.

Usage via the library

from modalityscan import ModalityClassifier

classifier = ModalityClassifier(model_dir="models")
results, series = classifier.classify(
    [{"id": "series-1/f000", "image_path": "frame.png"}],
    policy="balanced",
    aggregate_by="series",
)
print(results[0].modality.label, results[0].modality.granularity)

Modality is constant across a DICOM series, so passing several frames with a shared series_id and aggregate_by="series" is the cheapest accuracy available in the serving path.

Attribution

Training data is used under the licences recorded in ATTRIBUTION.md. Repository: vectorsense/modalityscan.