File size: 13,052 Bytes
edaf365 2a85d5a edaf365 2a85d5a edaf365 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 | ---
license: apache-2.0
library_name: onnx
pipeline_tag: image-classification
tags:
- medical-imaging
- modality-classification
- dicom
- onnx
- int8
---
# ModalityScan
Single-label imaging-modality classification for medical images. One shared ConvNet backbone, two
softmax heads (`family`, `modality`), shipped as a static-INT8 ONNX graph.
Predicts **what kind of image this is**, not what is in it. Anatomy is a different model.
## What it outputs
```
input pixel_values float32 [batch, 3, 224, 224] NCHW, ImageNet-normalized, letterboxed
outputs family_logits float32 [batch, 6]
modality_logits float32 [batch, 20]
```
Softmax, temperature and thresholds live **outside** the graph so calibration and policy stay
configuration. `modality-taxonomy.json` travels with the weights - a 20-element
logits vector is 20 meaningless numbers without it.
## Classes
| class | family | tier | train frames | sources | P | R | F1 | support |
|---|---|---|---|---|---|---|---|---|
| `CT` | cross_sectional | model_trained | 5,996 | 3 | 0.8469 | 0.8059 | 0.8259 | 2,684 |
| `MR` | cross_sectional | experimental | 404 | 1 | 0.2476 | 1.0 | 0.3969 | 52 |
| `US` | ultrasound | model_trained | 3,779 | 2 | 0.861 | 0.9558 | 0.9059 | 3,123 |
| `CR` | projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `DX` | projection_radiography | model_trained | 6,000 | 2 | 0.9956 | 0.9814 | 0.9884 | 3,442 |
| `MG` | projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `RF` | projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
| `XA` | projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
| `PT` | cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `NM` | cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `OCT` | optical | experimental | 6,000 | 1 | 0.9977 | 1.0 | 0.9989 | 881 |
| `FUNDUS` | optical | experimental | 1,600 | 1 | 1.0 | 1.0 | 1.0 | 246 |
| `DERMOSCOPY` | optical | experimental | 6,000 | 1 | 0.5773 | 0.9943 | 0.7305 | 879 |
| `ENDOSCOPY` | optical | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `HISTOPATHOLOGY` | microscopy | experimental | 6,000 | 1 | 0.2853 | 1.0 | 0.4439 | 926 |
| `MICROSCOPY` | microscopy | not_trained | 6,000 | 2 | n/a | n/a | n/a | 3,479 |
| `PHOTO` | non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `RENDER_3D` | non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `DOCUMENT` | non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
| `WAVEFORM` | non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
**Tier is a statement about evidence, not difficulty.** `model_trained` means the corpus had at
least two independent sources and enough patients to make the number falsifiable. `experimental`
means it did not. `not_trained` classes are never emitted by the decoder, so no metrics are
reported for them.
Some `not_trained` classes *do* have training frames. They were withdrawn after
measurement rather than never attempted, and the decoder never emits them:
* **`MICROSCOPY` is not claimed.** Withdrawn from the claim after measurement. Held-out-source recall was 97/3,479 = 2.8%, with 60% of held-out frames predicted HISTOPATHOLOGY. Its two sources are not interchangeable samples of one modality: TissueMNIST (trained on) is grayscale kidney cortex, BloodMNIST (held out) is colour-stained blood smears, so training never saw colour microscopy. The slot stays in the 20-wide output because that is a published contract; the decoder never emits it.
## Results
Headline: **macro-F1 0.9067** over the `model_trained` classes
(`CT`, `DX`, `US`), balanced policy, on the held-out test split.
Accuracy on frames whose true class is one the model claims: **0.9218**.
Accuracy over *every* labelled frame including classes the model does not claim:
**0.7446** - the difference is almost entirely withdrawn-class frames,
which are counted as errors here by design.
| benchmark tier | rows | accuracy | macro-F1 |
|---|---|---|---|
| held_out_source | 11,000 | 0.6618 | 0.9021 |
| in_distribution | 4,712 | 0.9378 | 0.9109 |
In-distribution macro-F1 is 0.9109 and held-out-source macro-F1 is 0.9021. The gap is small, which is the evidence that the model reads the image rather than the dataset it came from.
Coarse and fine accuracy are identical here: no collapse group has two populated classes in this corpus, so the group step of the cascade never fires. The cascade is still live - it is simply untested by this evaluation.
Open-set rejection: **not measured**. The corpus contains no `open_set` rows, so the energy cut was fitted for false-rejection rate only and its true detection rate is unknown. `out_of_distribution: true` is provisional - do not rely on it in production until it has been measured against real non-diagnostic material.
Expected calibration error on the test split: **0.0991**
(calibrated). Temperature was fitted on validation; the test
split is harder because it carries every held-out source, so this is the pessimistic figure.
## The abstention cascade
When no single class clears its threshold the decoder falls back, and always reports how far it
fell:
| `granularity` | meaning |
|---|---|
| `modality` | a specific modality, e.g. `CT` |
| `group` | a confusable group, e.g. `XR_PLAIN` when CR and DX are not separable |
| `family` | an acquisition family, e.g. `projection_radiography` |
| `none` | abstained, or rejected as out-of-distribution |
A caller routing to a CT pipeline should require `granularity == "modality"`. A caller deciding
whether an image is worth processing at all is satisfied by `family`.
## Limitations
* **MR sequence, contrast phase and anatomy are out of scope.** This model says `MR`; it does not
say `T2 FLAIR`, and it does not say `brain`.
* **Secondary Capture is not a modality.** An SC object is a screenshot of something else. The
model sees pixels and will answer for what was on screen.
* **Compound figures have no single answer.** A literature figure with CT and MR panels side by
side violates the single-label assumption this model is built on.
* **Digitised screen-film mammography is over-represented** in the `MG` training data relative to
modern full-field digital mammography.
* Scores are calibrated. Read the measured precision and
recall above, not the score, when choosing a policy.
## Quick start (Python + onnxruntime)
No clone, no framework - the artifact and its config files are all you need.
```bash
pip install onnxruntime pillow numpy huggingface_hub
```
```python
import json, numpy as np, onnxruntime as ort
from PIL import Image
from huggingface_hub import hf_hub_download
REPO = "vectorsense/modalityscan"
sess = ort.InferenceSession(hf_hub_download(REPO, "model.int8.onnx"),
providers=["CPUExecutionProvider"])
tax = json.load(open(hf_hub_download(REPO, "modality-taxonomy.json"), encoding="utf-8"))
book = json.load(open(hf_hub_download(REPO, "modality-thresholds.json"), encoding="utf-8"))
pre = json.load(open(hf_hub_download(REPO, "preprocessor.json"), encoding="utf-8"))
MODALITY = tax["levels"]["modality"]["index"]
FAMILY = tax["levels"]["family"]["index"]
TIER = {label: tier
for tier in ("model_trained", "experimental", "not_trained")
for label in tax["tiers"][tier]["types"]}
mean = np.array(pre["image_mean"], np.float32)[:, None, None]
std = np.array(pre["image_std"], np.float32)[:, None, None]
def pixel_values(path, size=224):
# Letterbox, never centre-crop. The field-of-view outline - ultrasound sector, circular
# fundus aperture, endoscope vignette - is a top-tier modality cue; cropping discards it.
image = Image.open(path).convert("RGB")
scale = size / max(image.size)
small = image.resize((max(1, round(image.width * scale)),
max(1, round(image.height * scale))), Image.BILINEAR)
canvas = Image.new("RGB", (size, size), (0, 0, 0))
canvas.paste(small, ((size - small.width) // 2, (size - small.height) // 2))
array = np.asarray(canvas, np.float32) / 255.0
return ((array.transpose(2, 0, 1) - mean) / std)[None]
def softmax(x):
e = np.exp(x - x.max())
return e / e.sum()
family_logits, modality_logits = sess.run(
["family_logits", "modality_logits"], {"pixel_values": pixel_values("frame.png")}
)
# Temperature lives outside the graph so calibration stays configuration, not a re-export.
T = book["calibration"]["temperature"]
p_mod = softmax(modality_logits[0] / T["modality"])
p_fam = softmax(family_logits[0] / T["family"])
policy = book["default_policy"]
floor = book["policies"][policy]["modality"]
over = book["per_class_overrides"].get(policy, {})
# Energy over logits, not softmax: softmax is confident on garbage by construction.
energy = -float(np.log(np.exp(modality_logits[0] / T["modality"]).sum()))
if energy > book["ood"]["energy_threshold"][policy]:
print("out of distribution ->", book["ood"]["fallback_label"])
else:
# `not_trained` classes keep their slot because the output width is a published contract,
# but they are not claims. A raw argmax will eventually return one of them.
best = next(i for i in np.argsort(-p_mod) if TIER.get(MODALITY[i]) != "not_trained")
label, score = MODALITY[best], float(p_mod[best])
if score >= over.get(label, floor):
print(f"modality: {label} {score:.3f} tier={TIER[label]}")
else:
print(f"family: {FAMILY[int(p_fam.argmax())]} {p_fam.max():.3f} (cascade fell back)")
```
Two lines carry most of the correctness. **Letterbox rather than centre-crop** - the field-of-view
outline is signal, not padding. **Skip `not_trained` classes** - a raw `argmax` will eventually
hand you a label this model does not stand behind.
## Run it as a REST API
The companion repo ships a FastAPI service with the cascade, the calibrated thresholds, the
energy rejection and series aggregation already wired up.
```bash
pip install fastapi "uvicorn[standard]" onnxruntime pillow numpy
uvicorn service.app:app --port 5003 # auto-loads model.int8.onnx from ./models
```
```bash
curl -s http://127.0.0.1:5003/classify_modality \
-H "Content-Type: application/json" \
-d '{"frames":[{"id":"series-1/f000","image_base64":"iVBORw0KGgo..."}],
"options":{"policy":"balanced"}}'
```
The same call from Python - send a PNG or JPEG, get parsed JSON back:
```python
import base64, requests
URL = "http://127.0.0.1:5003/classify_modality"
def classify(path, policy="balanced"):
payload = {
"frames": [{"id": "series-1/f000",
"image_base64": base64.b64encode(open(path, "rb").read()).decode()}],
"options": {"policy": policy, "top_k": 3},
}
response = requests.post(URL, json=payload, timeout=30)
response.raise_for_status() # unknown request fields return 422, never a silent default
return response.json()
frame = classify("frame.png")["results"][0]
print(frame["modality"]["label"], frame["modality"]["granularity"], frame["modality"]["score"])
print("out of distribution:", frame["out_of_distribution"])
```
A whole DICOM series in one request, which is the cheapest accuracy in the serving path:
```python
payload = {
"frames": [{"id": f"series-1/f{i:03d}",
"image_base64": base64.b64encode(open(p, "rb").read()).decode()}
for i, p in enumerate(frame_paths)],
"options": {"policy": "balanced", "aggregate_by": "series"},
}
print(requests.post(URL, json=payload, timeout=60).json()["aggregate"])
```
Endpoints: `GET /`, `POST /classify_modality`, `GET /health/ready`, `GET /metadata`.
`GET /metadata` reports the tier of every class at runtime, so a caller can tell a claimed class
from an experimental one without reading this card.
Prefer the service over the raw-ONNX snippet unless you have a reason not to: it already applies
the calibrated temperature, the per-class thresholds, the abstention cascade and the rejection
path. The snippet reimplements those in miniature and will drift from the config.
## Usage via the library
```python
from modalityscan import ModalityClassifier
classifier = ModalityClassifier(model_dir="models")
results, series = classifier.classify(
[{"id": "series-1/f000", "image_path": "frame.png"}],
policy="balanced",
aggregate_by="series",
)
print(results[0].modality.label, results[0].modality.granularity)
```
Modality is constant across a DICOM series, so passing several frames with a shared `series_id`
and `aggregate_by="series"` is the cheapest accuracy available in the serving path.
## Attribution
Training data is used under the licences recorded in `ATTRIBUTION.md`. Repository: `vectorsense/modalityscan`.
|