ModalityScan
Single-label imaging-modality classification for medical images. One shared ConvNet backbone, two
softmax heads (family, modality), shipped as a static-INT8 ONNX graph.
Predicts what kind of image this is, not what is in it. Anatomy is a different model.
What it outputs
input pixel_values float32 [batch, 3, 224, 224] NCHW, ImageNet-normalized, letterboxed
outputs family_logits float32 [batch, 6]
modality_logits float32 [batch, 20]
Softmax, temperature and thresholds live outside the graph so calibration and policy stay
configuration. modality-taxonomy.json travels with the weights - a 20-element
logits vector is 20 meaningless numbers without it.
Classes
| class | family | tier | train frames | sources | P | R | F1 | support |
|---|---|---|---|---|---|---|---|---|
CT |
cross_sectional | model_trained | 5,996 | 3 | 0.8469 | 0.8059 | 0.8259 | 2,684 |
MR |
cross_sectional | experimental | 404 | 1 | 0.2476 | 1.0 | 0.3969 | 52 |
US |
ultrasound | model_trained | 3,779 | 2 | 0.861 | 0.9558 | 0.9059 | 3,123 |
CR |
projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
DX |
projection_radiography | model_trained | 6,000 | 2 | 0.9956 | 0.9814 | 0.9884 | 3,442 |
MG |
projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
RF |
projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
XA |
projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
PT |
cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
NM |
cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
OCT |
optical | experimental | 6,000 | 1 | 0.9977 | 1.0 | 0.9989 | 881 |
FUNDUS |
optical | experimental | 1,600 | 1 | 1.0 | 1.0 | 1.0 | 246 |
DERMOSCOPY |
optical | experimental | 6,000 | 1 | 0.5773 | 0.9943 | 0.7305 | 879 |
ENDOSCOPY |
optical | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
HISTOPATHOLOGY |
microscopy | experimental | 6,000 | 1 | 0.2853 | 1.0 | 0.4439 | 926 |
MICROSCOPY |
microscopy | not_trained | 6,000 | 2 | n/a | n/a | n/a | 3,479 |
PHOTO |
non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
RENDER_3D |
non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
DOCUMENT |
non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
WAVEFORM |
non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
Tier is a statement about evidence, not difficulty. model_trained means the corpus had at
least two independent sources and enough patients to make the number falsifiable. experimental
means it did not. not_trained classes are never emitted by the decoder, so no metrics are
reported for them.
Some not_trained classes do have training frames. They were withdrawn after
measurement rather than never attempted, and the decoder never emits them:
MICROSCOPYis not claimed. Withdrawn from the claim after measurement. Held-out-source recall was 97/3,479 = 2.8%, with 60% of held-out frames predicted HISTOPATHOLOGY. Its two sources are not interchangeable samples of one modality: TissueMNIST (trained on) is grayscale kidney cortex, BloodMNIST (held out) is colour-stained blood smears, so training never saw colour microscopy. The slot stays in the 20-wide output because that is a published contract; the decoder never emits it.
Results
Headline: macro-F1 0.9067 over the model_trained classes
(CT, DX, US), balanced policy, on the held-out test split.
Accuracy on frames whose true class is one the model claims: 0.9218. Accuracy over every labelled frame including classes the model does not claim: 0.7446 - the difference is almost entirely withdrawn-class frames, which are counted as errors here by design.
| benchmark tier | rows | accuracy | macro-F1 |
|---|---|---|---|
| held_out_source | 11,000 | 0.6618 | 0.9021 |
| in_distribution | 4,712 | 0.9378 | 0.9109 |
In-distribution macro-F1 is 0.9109 and held-out-source macro-F1 is 0.9021. The gap is small, which is the evidence that the model reads the image rather than the dataset it came from.
Coarse and fine accuracy are identical here: no collapse group has two populated classes in this corpus, so the group step of the cascade never fires. The cascade is still live - it is simply untested by this evaluation.
Open-set rejection: not measured. The corpus contains no open_set rows, so the energy cut was fitted for false-rejection rate only and its true detection rate is unknown. out_of_distribution: true is provisional - do not rely on it in production until it has been measured against real non-diagnostic material.
Expected calibration error on the test split: 0.0991 (calibrated). Temperature was fitted on validation; the test split is harder because it carries every held-out source, so this is the pessimistic figure.
The abstention cascade
When no single class clears its threshold the decoder falls back, and always reports how far it fell:
granularity |
meaning |
|---|---|
modality |
a specific modality, e.g. CT |
group |
a confusable group, e.g. XR_PLAIN when CR and DX are not separable |
family |
an acquisition family, e.g. projection_radiography |
none |
abstained, or rejected as out-of-distribution |
A caller routing to a CT pipeline should require granularity == "modality". A caller deciding
whether an image is worth processing at all is satisfied by family.
Limitations
- MR sequence, contrast phase and anatomy are out of scope. This model says
MR; it does not sayT2 FLAIR, and it does not saybrain. - Secondary Capture is not a modality. An SC object is a screenshot of something else. The model sees pixels and will answer for what was on screen.
- Compound figures have no single answer. A literature figure with CT and MR panels side by side violates the single-label assumption this model is built on.
- Digitised screen-film mammography is over-represented in the
MGtraining data relative to modern full-field digital mammography. - Scores are calibrated. Read the measured precision and recall above, not the score, when choosing a policy.
Quick start (Python + onnxruntime)
No clone, no framework - the artifact and its config files are all you need.
pip install onnxruntime pillow numpy huggingface_hub
import json, numpy as np, onnxruntime as ort
from PIL import Image
from huggingface_hub import hf_hub_download
REPO = "vectorsense/modalityscan"
sess = ort.InferenceSession(hf_hub_download(REPO, "model.int8.onnx"),
providers=["CPUExecutionProvider"])
tax = json.load(open(hf_hub_download(REPO, "modality-taxonomy.json"), encoding="utf-8"))
book = json.load(open(hf_hub_download(REPO, "modality-thresholds.json"), encoding="utf-8"))
pre = json.load(open(hf_hub_download(REPO, "preprocessor.json"), encoding="utf-8"))
MODALITY = tax["levels"]["modality"]["index"]
FAMILY = tax["levels"]["family"]["index"]
TIER = {label: tier
for tier in ("model_trained", "experimental", "not_trained")
for label in tax["tiers"][tier]["types"]}
mean = np.array(pre["image_mean"], np.float32)[:, None, None]
std = np.array(pre["image_std"], np.float32)[:, None, None]
def pixel_values(path, size=224):
# Letterbox, never centre-crop. The field-of-view outline - ultrasound sector, circular
# fundus aperture, endoscope vignette - is a top-tier modality cue; cropping discards it.
image = Image.open(path).convert("RGB")
scale = size / max(image.size)
small = image.resize((max(1, round(image.width * scale)),
max(1, round(image.height * scale))), Image.BILINEAR)
canvas = Image.new("RGB", (size, size), (0, 0, 0))
canvas.paste(small, ((size - small.width) // 2, (size - small.height) // 2))
array = np.asarray(canvas, np.float32) / 255.0
return ((array.transpose(2, 0, 1) - mean) / std)[None]
def softmax(x):
e = np.exp(x - x.max())
return e / e.sum()
family_logits, modality_logits = sess.run(
["family_logits", "modality_logits"], {"pixel_values": pixel_values("frame.png")}
)
# Temperature lives outside the graph so calibration stays configuration, not a re-export.
T = book["calibration"]["temperature"]
p_mod = softmax(modality_logits[0] / T["modality"])
p_fam = softmax(family_logits[0] / T["family"])
policy = book["default_policy"]
floor = book["policies"][policy]["modality"]
over = book["per_class_overrides"].get(policy, {})
# Energy over logits, not softmax: softmax is confident on garbage by construction.
energy = -float(np.log(np.exp(modality_logits[0] / T["modality"]).sum()))
if energy > book["ood"]["energy_threshold"][policy]:
print("out of distribution ->", book["ood"]["fallback_label"])
else:
# `not_trained` classes keep their slot because the output width is a published contract,
# but they are not claims. A raw argmax will eventually return one of them.
best = next(i for i in np.argsort(-p_mod) if TIER.get(MODALITY[i]) != "not_trained")
label, score = MODALITY[best], float(p_mod[best])
if score >= over.get(label, floor):
print(f"modality: {label} {score:.3f} tier={TIER[label]}")
else:
print(f"family: {FAMILY[int(p_fam.argmax())]} {p_fam.max():.3f} (cascade fell back)")
Two lines carry most of the correctness. Letterbox rather than centre-crop - the field-of-view
outline is signal, not padding. Skip not_trained classes - a raw argmax will eventually
hand you a label this model does not stand behind.
Run it as a REST API
The companion repo ships a FastAPI service with the cascade, the calibrated thresholds, the energy rejection and series aggregation already wired up.
pip install fastapi "uvicorn[standard]" onnxruntime pillow numpy
uvicorn service.app:app --port 5003 # auto-loads model.int8.onnx from ./models
curl -s http://127.0.0.1:5003/classify_modality \
-H "Content-Type: application/json" \
-d '{"frames":[{"id":"series-1/f000","image_base64":"iVBORw0KGgo..."}],
"options":{"policy":"balanced"}}'
The same call from Python - send a PNG or JPEG, get parsed JSON back:
import base64, requests
URL = "http://127.0.0.1:5003/classify_modality"
def classify(path, policy="balanced"):
payload = {
"frames": [{"id": "series-1/f000",
"image_base64": base64.b64encode(open(path, "rb").read()).decode()}],
"options": {"policy": policy, "top_k": 3},
}
response = requests.post(URL, json=payload, timeout=30)
response.raise_for_status() # unknown request fields return 422, never a silent default
return response.json()
frame = classify("frame.png")["results"][0]
print(frame["modality"]["label"], frame["modality"]["granularity"], frame["modality"]["score"])
print("out of distribution:", frame["out_of_distribution"])
A whole DICOM series in one request, which is the cheapest accuracy in the serving path:
payload = {
"frames": [{"id": f"series-1/f{i:03d}",
"image_base64": base64.b64encode(open(p, "rb").read()).decode()}
for i, p in enumerate(frame_paths)],
"options": {"policy": "balanced", "aggregate_by": "series"},
}
print(requests.post(URL, json=payload, timeout=60).json()["aggregate"])
Endpoints: GET /, POST /classify_modality, GET /health/ready, GET /metadata.
GET /metadata reports the tier of every class at runtime, so a caller can tell a claimed class
from an experimental one without reading this card.
Prefer the service over the raw-ONNX snippet unless you have a reason not to: it already applies the calibrated temperature, the per-class thresholds, the abstention cascade and the rejection path. The snippet reimplements those in miniature and will drift from the config.
Usage via the library
from modalityscan import ModalityClassifier
classifier = ModalityClassifier(model_dir="models")
results, series = classifier.classify(
[{"id": "series-1/f000", "image_path": "frame.png"}],
policy="balanced",
aggregate_by="series",
)
print(results[0].modality.label, results[0].modality.granularity)
Modality is constant across a DICOM series, so passing several frames with a shared series_id
and aggregate_by="series" is the cheapest accuracy available in the serving path.
Attribution
Training data is used under the licences recorded in ATTRIBUTION.md. Repository: vectorsense/modalityscan.
- Downloads last month
- -