File size: 13,052 Bytes
edaf365
 
2a85d5a
edaf365
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2a85d5a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
edaf365
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
---

license: apache-2.0
library_name: onnx
pipeline_tag: image-classification
tags:
  - medical-imaging
  - modality-classification
  - dicom
  - onnx
  - int8
---


# ModalityScan

Single-label imaging-modality classification for medical images. One shared ConvNet backbone, two
softmax heads (`family`, `modality`), shipped as a static-INT8 ONNX graph.

Predicts **what kind of image this is**, not what is in it. Anatomy is a different model.

## What it outputs

```

input   pixel_values     float32  [batch, 3, 224, 224]   NCHW, ImageNet-normalized, letterboxed

outputs family_logits    float32  [batch, 6]

        modality_logits  float32  [batch, 20]

```

Softmax, temperature and thresholds live **outside** the graph so calibration and policy stay
configuration. `modality-taxonomy.json` travels with the weights - a 20-element
logits vector is 20 meaningless numbers without it.

## Classes

| class | family | tier | train frames | sources | P | R | F1 | support |
|---|---|---|---|---|---|---|---|---|
| `CT` | cross_sectional | model_trained | 5,996 | 3 | 0.8469 | 0.8059 | 0.8259 | 2,684 |
| `MR` | cross_sectional | experimental | 404 | 1 | 0.2476 | 1.0 | 0.3969 | 52 |

| `US` | ultrasound | model_trained | 3,779 | 2 | 0.861 | 0.9558 | 0.9059 | 3,123 |
| `CR` | projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |

| `DX` | projection_radiography | model_trained | 6,000 | 2 | 0.9956 | 0.9814 | 0.9884 | 3,442 |

| `MG` | projection_radiography | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `RF` | projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
| `XA` | projection_radiography | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |
| `PT` | cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |

| `NM` | cross_sectional | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `OCT` | optical | experimental | 6,000 | 1 | 0.9977 | 1.0 | 0.9989 | 881 |
| `FUNDUS` | optical | experimental | 1,600 | 1 | 1.0 | 1.0 | 1.0 | 246 |
| `DERMOSCOPY` | optical | experimental | 6,000 | 1 | 0.5773 | 0.9943 | 0.7305 | 879 |
| `ENDOSCOPY` | optical | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `HISTOPATHOLOGY` | microscopy | experimental | 6,000 | 1 | 0.2853 | 1.0 | 0.4439 | 926 |
| `MICROSCOPY` | microscopy | not_trained | 6,000 | 2 | n/a | n/a | n/a | 3,479 |

| `PHOTO` | non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |
| `RENDER_3D` | non_diagnostic | experimental | 0 | 0 | 0.0 | 0.0 | 0.0 | 0 |

| `DOCUMENT` | non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |

| `WAVEFORM` | non_diagnostic | not_trained | 0 | 0 | n/a | n/a | n/a | 0 |



**Tier is a statement about evidence, not difficulty.** `model_trained` means the corpus had at
least two independent sources and enough patients to make the number falsifiable. `experimental`
means it did not. `not_trained` classes are never emitted by the decoder, so no metrics are
reported for them.

Some `not_trained` classes *do* have training frames. They were withdrawn after
measurement rather than never attempted, and the decoder never emits them:

* **`MICROSCOPY` is not claimed.** Withdrawn from the claim after measurement. Held-out-source recall was 97/3,479 = 2.8%, with 60% of held-out frames predicted HISTOPATHOLOGY. Its two sources are not interchangeable samples of one modality: TissueMNIST (trained on) is grayscale kidney cortex, BloodMNIST (held out) is colour-stained blood smears, so training never saw colour microscopy. The slot stays in the 20-wide output because that is a published contract; the decoder never emits it.

## Results

Headline: **macro-F1 0.9067** over the `model_trained` classes
(`CT`, `DX`, `US`), balanced policy, on the held-out test split.

Accuracy on frames whose true class is one the model claims: **0.9218**.
Accuracy over *every* labelled frame including classes the model does not claim:
**0.7446** - the difference is almost entirely withdrawn-class frames,
which are counted as errors here by design.

| benchmark tier | rows | accuracy | macro-F1 |
|---|---|---|---|
| held_out_source | 11,000 | 0.6618 | 0.9021 |
| in_distribution | 4,712 | 0.9378 | 0.9109 |



In-distribution macro-F1 is 0.9109 and held-out-source macro-F1 is 0.9021. The gap is small, which is the evidence that the model reads the image rather than the dataset it came from.



Coarse and fine accuracy are identical here: no collapse group has two populated classes in this corpus, so the group step of the cascade never fires. The cascade is still live - it is simply untested by this evaluation.



Open-set rejection: **not measured**. The corpus contains no `open_set` rows, so the energy cut was fitted for false-rejection rate only and its true detection rate is unknown. `out_of_distribution: true` is provisional - do not rely on it in production until it has been measured against real non-diagnostic material.

Expected calibration error on the test split: **0.0991**
(calibrated). Temperature was fitted on validation; the test
split is harder because it carries every held-out source, so this is the pessimistic figure.

## The abstention cascade

When no single class clears its threshold the decoder falls back, and always reports how far it
fell:

| `granularity` | meaning |
|---|---|
| `modality` | a specific modality, e.g. `CT` |
| `group` | a confusable group, e.g. `XR_PLAIN` when CR and DX are not separable |
| `family` | an acquisition family, e.g. `projection_radiography` |
| `none` | abstained, or rejected as out-of-distribution |

A caller routing to a CT pipeline should require `granularity == "modality"`. A caller deciding
whether an image is worth processing at all is satisfied by `family`.

## Limitations

* **MR sequence, contrast phase and anatomy are out of scope.** This model says `MR`; it does not
  say `T2 FLAIR`, and it does not say `brain`.
* **Secondary Capture is not a modality.** An SC object is a screenshot of something else. The
  model sees pixels and will answer for what was on screen.
* **Compound figures have no single answer.** A literature figure with CT and MR panels side by
  side violates the single-label assumption this model is built on.
* **Digitised screen-film mammography is over-represented** in the `MG` training data relative to
  modern full-field digital mammography.
* Scores are calibrated. Read the measured precision and
  recall above, not the score, when choosing a policy.

## Quick start (Python + onnxruntime)

No clone, no framework - the artifact and its config files are all you need.

```bash

pip install onnxruntime pillow numpy huggingface_hub

```

```python

import json, numpy as np, onnxruntime as ort

from PIL import Image

from huggingface_hub import hf_hub_download



REPO = "vectorsense/modalityscan"

sess = ort.InferenceSession(hf_hub_download(REPO, "model.int8.onnx"),

                            providers=["CPUExecutionProvider"])

tax  = json.load(open(hf_hub_download(REPO, "modality-taxonomy.json"), encoding="utf-8"))

book = json.load(open(hf_hub_download(REPO, "modality-thresholds.json"), encoding="utf-8"))

pre  = json.load(open(hf_hub_download(REPO, "preprocessor.json"), encoding="utf-8"))



MODALITY = tax["levels"]["modality"]["index"]

FAMILY   = tax["levels"]["family"]["index"]

TIER     = {label: tier

            for tier in ("model_trained", "experimental", "not_trained")

            for label in tax["tiers"][tier]["types"]}



mean = np.array(pre["image_mean"], np.float32)[:, None, None]

std  = np.array(pre["image_std"],  np.float32)[:, None, None]



def pixel_values(path, size=224):

    # Letterbox, never centre-crop. The field-of-view outline - ultrasound sector, circular

    # fundus aperture, endoscope vignette - is a top-tier modality cue; cropping discards it.

    image = Image.open(path).convert("RGB")

    scale = size / max(image.size)

    small = image.resize((max(1, round(image.width * scale)),

                          max(1, round(image.height * scale))), Image.BILINEAR)

    canvas = Image.new("RGB", (size, size), (0, 0, 0))

    canvas.paste(small, ((size - small.width) // 2, (size - small.height) // 2))

    array = np.asarray(canvas, np.float32) / 255.0

    return ((array.transpose(2, 0, 1) - mean) / std)[None]



def softmax(x):

    e = np.exp(x - x.max())

    return e / e.sum()



family_logits, modality_logits = sess.run(

    ["family_logits", "modality_logits"], {"pixel_values": pixel_values("frame.png")}

)



# Temperature lives outside the graph so calibration stays configuration, not a re-export.

T = book["calibration"]["temperature"]

p_mod = softmax(modality_logits[0] / T["modality"])

p_fam = softmax(family_logits[0] / T["family"])



policy = book["default_policy"]

floor  = book["policies"][policy]["modality"]

over   = book["per_class_overrides"].get(policy, {})



# Energy over logits, not softmax: softmax is confident on garbage by construction.

energy = -float(np.log(np.exp(modality_logits[0] / T["modality"]).sum()))

if energy > book["ood"]["energy_threshold"][policy]:

    print("out of distribution ->", book["ood"]["fallback_label"])

else:

    # `not_trained` classes keep their slot because the output width is a published contract,

    # but they are not claims. A raw argmax will eventually return one of them.

    best = next(i for i in np.argsort(-p_mod) if TIER.get(MODALITY[i]) != "not_trained")

    label, score = MODALITY[best], float(p_mod[best])

    if score >= over.get(label, floor):

        print(f"modality: {label} {score:.3f}  tier={TIER[label]}")

    else:

        print(f"family:   {FAMILY[int(p_fam.argmax())]} {p_fam.max():.3f}  (cascade fell back)")

```

Two lines carry most of the correctness. **Letterbox rather than centre-crop** - the field-of-view
outline is signal, not padding. **Skip `not_trained` classes** - a raw `argmax` will eventually

hand you a label this model does not stand behind.



## Run it as a REST API



The companion repo ships a FastAPI service with the cascade, the calibrated thresholds, the

energy rejection and series aggregation already wired up.



```bash

pip install fastapi "uvicorn[standard]" onnxruntime pillow numpy

uvicorn service.app:app --port 5003      # auto-loads model.int8.onnx from ./models

```



```bash

curl -s http://127.0.0.1:5003/classify_modality \

  -H "Content-Type: application/json" \

  -d '{"frames":[{"id":"series-1/f000","image_base64":"iVBORw0KGgo..."}],

       "options":{"policy":"balanced"}}'

```



The same call from Python - send a PNG or JPEG, get parsed JSON back:



```python

import base64, requests



URL = "http://127.0.0.1:5003/classify_modality"



def classify(path, policy="balanced"):

    payload = {

        "frames": [{"id": "series-1/f000",

                    "image_base64": base64.b64encode(open(path, "rb").read()).decode()}],

        "options": {"policy": policy, "top_k": 3},

    }

    response = requests.post(URL, json=payload, timeout=30)

    response.raise_for_status()     # unknown request fields return 422, never a silent default

    return response.json()



frame = classify("frame.png")["results"][0]

print(frame["modality"]["label"], frame["modality"]["granularity"], frame["modality"]["score"])

print("out of distribution:", frame["out_of_distribution"])

```



A whole DICOM series in one request, which is the cheapest accuracy in the serving path:



```python

payload = {

    "frames": [{"id": f"series-1/f{i:03d}",

                "image_base64": base64.b64encode(open(p, "rb").read()).decode()}

               for i, p in enumerate(frame_paths)],

    "options": {"policy": "balanced", "aggregate_by": "series"},

}

print(requests.post(URL, json=payload, timeout=60).json()["aggregate"])

```



Endpoints: `GET /`, `POST /classify_modality`, `GET /health/ready`, `GET /metadata`.

`GET /metadata` reports the tier of every class at runtime, so a caller can tell a claimed class

from an experimental one without reading this card.



Prefer the service over the raw-ONNX snippet unless you have a reason not to: it already applies

the calibrated temperature, the per-class thresholds, the abstention cascade and the rejection

path. The snippet reimplements those in miniature and will drift from the config.



## Usage via the library



```python

from modalityscan import ModalityClassifier



classifier = ModalityClassifier(model_dir="models")

results, series = classifier.classify(

    [{"id": "series-1/f000", "image_path": "frame.png"}],

    policy="balanced",

    aggregate_by="series",

)

print(results[0].modality.label, results[0].modality.granularity)

```



Modality is constant across a DICOM series, so passing several frames with a shared `series_id`

and `aggregate_by="series"` is the cheapest accuracy available in the serving path.



## Attribution



Training data is used under the licences recorded in `ATTRIBUTION.md`. Repository: `vectorsense/modalityscan`.