platypus weights
Trained weights for pyplatypus and the platypus R
package - segmentation and object detection. Every weights file has a .json beside it
recording what it was trained for and on.
Weights are loaded by name, and a name is pinned to a commit of this repository, so it always means the same numbers:
from pyplatypus import Engine, from_dict
engine = Engine(from_dict({
"data": {"train_path": "images/", "validation_path": "images/",
"colormap": [[0, 0, 0], [255, 255, 255]]},
"models": [{"name": "nuclei", "input_shape": [256, 256], "channels": 3, "n_class": 2,
"blocks": 4, "filters": 16, "weights": "dsbowl-unet", "fit": False}],
}))
engine.fit() # loads the weights, trains nothing
masks = engine.predict("nuclei", split="validation")
library(platypus)
spec <- platypus_spec(
data = segmentation_data("images/", "images/", colormap = binary_colormap),
models = list(u_net("nuclei", input_shape = c(256, 256), blocks = 4, filters = 16,
weights = "dsbowl-unet", fit = FALSE))
)
masks <- predict(platypus_fit(spec), split = "validation")
Fetching by name needs pip install "pyplatypus[hub]". A local path to a downloaded file works
without it, which is the route for machines with no internet access.
dsbowl-unet
A plain U-Net that separates cell nuclei from background in light microscopy images.
| Architecture | U-Net, 4 blocks, 16 base filters |
| Input | 256 × 256, 3 channels |
| Output | 2 classes: background, nucleus |
| Parameters | 1,942,594 |
| File | 7.8 MB, safetensors |
| Loss | CCE-Dice |
What it was trained on
BBBC038v1, the 2018 Data Science Bowl image set, stage1_train, from the
Broad Bioimage Benchmark Collection. 670 images of
nuclei in varied cell types, magnifications and imaging modalities, split 536 / 134 by image with
seed 1.
The data is CC0: "the various contributors of the imagesets have waived all copyright and related or neighboring rights to BBBC038v1". Taken from the Broad rather than from the Kaggle mirror on purpose — the images are the same, but the Kaggle copy is governed by competition rules accepted at download, and weights are a derivative of the data behind them.
Citation, which CC0 does not require and which is owed anyway:
Caicedo et al., Nucleus segmentation across imaging experiments: the 2018 Data Science Bowl, Nature Methods, 2019. Image set BBBC038v1, Broad Bioimage Benchmark Collection.
How it scores
On the 134 held-out images, per image rather than pooled:
| metric | mean | sd | median | min | max |
|---|---|---|---|---|---|
| Dice | 0.9205 | 0.0532 | 0.9314 | 0.7240 | 0.9861 |
| IoU | 0.8570 | 0.0865 | 0.8716 | 0.5675 | 0.9725 |
The minimum is there deliberately. A mean of 0.92 with a worst case of 0.72 is a different model
from a mean of 0.92 with a worst case of 0.90, and only one of those numbers tells you which you
have. The five worst images are listed in dsbowl-unet.json.
Why a U-Net, and not one of the others
pyplatypus builds four architectures from one specification, so the question was settled by
running all four on this data - same split, same augmentation, same 60 epochs, one command:
| model | parameters | s/epoch | epochs | Dice mean | sd | median | min |
|---|---|---|---|---|---|---|---|
| U-Net | 1,942,594 | 15.5 | 60 | 0.9217 | 0.054 | 0.9353 | 0.7227 |
| U-Net++ | 2,263,730 | 18.8 | 44 | 0.9212 | 0.055 | 0.9323 | 0.7459 |
| LinkNet | 1,746,754 | 14.0 | 60 | 0.9203 | 0.053 | 0.9293 | 0.7205 |
| Res-U-Net | 2,029,682 | 14.2 | 55 | 0.9187 | 0.058 | 0.9293 | 0.6706 |
The spread across all four is 0.0030. The spread across images within any one of them is 0.054. The architectures are eighteen times closer to each other than the images are, which means they are not distinguishable on this data.
One number makes it plainer. These published weights score 0.9205; the U-Net in the table, same configuration and a different random seed, scored 0.9217. That 0.0012 of seed noise is 40% of the entire spread between the four architectures — switching from LinkNet to U-Net++ buys about what re-running the same model buys.
So only one set of weights is published, and the choice of U-Net is arbitrary but measured. If you are deciding where to spend effort on a problem like this one, the table says it is not on the architecture: more data, better augmentation or a longer run will all move the number further than swapping the decoder.
Reproduce it with python examples/compare_dsbowl_architectures.py.
What it does not do
It does not separate touching nuclei. This is semantic segmentation: every nucleus pixel gets the same class, so two nuclei in contact come back as one region. The 2018 Data Science Bowl was scored on instances, and the numbers above are not comparable to that leaderboard — they answer "which pixels are nuclei", not "how many nuclei are there". If you need counts, or per-nucleus measurements, you need an instance step on top (watershed on a distance transform is the usual starting point) and these weights are a reasonable input to it, not a replacement for it.
It has seen one collection. BBBC038v1 is varied by design, which is the point of the challenge, but generalisation was measured on a held-out split of the same collection and nothing else. On a microscope or a stain unlike anything in it, measure before trusting.
It is an alpha. pyplatypus is pre-1.0 and its API still moves. The weights will keep
working - a pinned commit does not change - but the code that loads them may.
Reproducing it
pip install "pyplatypus[hub]"
python examples/publish_dsbowl_weights.py --out weights/
The script downloads BBBC038v1 from the Broad, splits it with the same seed, trains, measures per image, and writes the weights with this provenance in their sidecar. Trained on a GTX 1070: 60 epochs, about 15 seconds each.
bccd-yolo3
A YOLOv3 that finds red blood cells, white blood cells and platelets in blood-smear photographs.
| Architecture | YOLOv3, Darknet-53 backbone, 3 anchors per grid |
| Input | 416 × 416, 3 channels, letterboxed |
| Output | 3 classes: RBC, WBC, Platelets |
| Parameters | 61,534,504 |
| File | 246 MB, safetensors |
| Anchors | fitted to this data, recorded in the sidecar |
The anchors are not optional. A detector's weights decode boxes relative to the anchors
they were trained with, so the same numbers read with other anchors produce boxes scaled by a
fixed factor - plausible boxes, plausible scores, wrong places, and nothing in the output to
say so. pyplatypus takes them from the sidecar when it loads, and refuses a specification
that names its own.
What it was trained on
BCCD: 364 blood-smear photographs with 4,888 boxes, in the dataset's own train/validation/test split of 205 / 87 / 72. MIT licence, copyright 2017 shenggan, with the original images and annotations from cosmicad and akshaylamba, re-organised into Pascal VOC format.
Taken from that repository rather than the Kaggle mirror. The images are the same, but the
Kaggle copy is governed by competition rules accepted at download, and weights are a
derivative of the data behind them — the same reasoning as dsbowl-unet above.
The split is the dataset's own, which is the only reason these numbers can be compared with anybody else's on BCCD.
How it scores
On the 72 held-out test images, at IoU 0.5 and averaged over 0.50–0.95, with precision and recall read at confidence 0.5:
| class | AP@0.5 | IoU of matches | truth | precision | recall |
|---|---|---|---|---|---|
| RBC | 0.7934 | 0.810 | 805 | 0.715 | 0.793 |
| WBC | 0.9587 | 0.849 | 71 | 0.922 | 1.000 |
| Platelets | 0.8191 | 0.715 | 69 | 0.543 | 0.913 |
| all | 0.8571 | 0.8058 | 945 |
mAP@[.50:.95] is 0.5003. The gap between that and 0.8571 is localisation: the objects are
found and the boxes fit them loosely, which mean_matched_iou says in one number instead of
through the difference between two.
These are the median seed, and that matters here
Five runs of this exact recipe, differing only in the seed:
| mean | sd | min | max | |
|---|---|---|---|---|
| mAP@0.5 | 0.8660 | 0.0159 | 0.8533 | 0.8868 |
| mAP@[.50:.95] | 0.5159 | 0.0199 | 0.5003 | 0.5436 |
| IoU of matches | 0.8041 | 0.0026 | 0.8002 | 0.8068 |
Published here is the median by mAP@0.5, not the best. The spread is wide enough that a single reported number would be whichever seed got chosen: the best run scores 0.8868, three points above this one, on the same code and the same data.
And one row of that table is six times steadier than another. How well the boxes fit barely
varies between seeds; how many are found does. Which rank the model gives a borderline cell
moves average precision and leaves the overlap alone. On a dataset this size, read
mean_matched_iou first.
What it does not do
Platelets are the unreliable class, and that is where almost all of the variance lives. Across the five seeds:
| class | AP spread | IoU spread |
|---|---|---|
| RBC | ± 0.0100 | ± 0.0048 |
| WBC | ± 0.0052 | ± 0.0049 |
| Platelets | ± 0.0409 | ± 0.0208 |
Platelets are small — a median side of 41 pixels against 103 for a red cell — and their boxes fit worst (IoU 0.715) and vary four to eight times more than the other two classes. Precision on them at confidence 0.5 is 0.543, meaning roughly half the platelets it reports are not there. If platelets are what you are counting, these weights are a starting point and not an answer.
It has seen one small collection. 205 training images of blood smears at one magnification, stained one way. BCCD is a demonstration dataset, not a clinical one, and generalisation was measured on its own held-out split and nowhere else.
It is not a diagnostic tool. Nothing about this was validated for clinical use, and a cell count from it is not a laboratory result.
About 0.1% of the training boxes could not be represented at all — 3 of 2,804. Two cells
of similar shape whose centres land in the same grid cell share a target slot, so the second
is dropped: never shown to the model, never counted as missed. At 416 on this data it is
negligible, and pyplatypus reports the figure before training rather than after.
It is an alpha. The weights will keep working — a pinned commit does not change — but the
code that loads them still moves, and detection is not in a released pyplatypus yet.
Reproducing it
# BCCD from https://github.com/Shenggan/BCCD_Dataset
python examples/detect_blood_cells.py --data path/to/BCCD --epochs 150 --seed 1 \
--save weights/bccd-yolo3.safetensors
150 epochs, about 37 minutes on a GTX 1070. The script writes the YAML of the run beside its
results, so the recipe in bccd-yolo3.json can be read back and run again.