platypus weights

Trained weights for pyplatypus and the platypus R package - segmentation and object detection. Every weights file has a .json beside it recording what it was trained for and on.

Weights are loaded by name, and a name is pinned to a commit of this repository, so it always means the same numbers:

from pyplatypus import Engine, from_dict

engine = Engine(from_dict({
    "data": {"train_path": "images/", "validation_path": "images/",
             "colormap": [[0, 0, 0], [255, 255, 255]]},
    "models": [{"name": "nuclei", "input_shape": [256, 256], "channels": 3, "n_class": 2,
                "blocks": 4, "filters": 16, "weights": "dsbowl-unet", "fit": False}],
}))
engine.fit()                      # loads the weights, trains nothing
masks = engine.predict("nuclei", split="validation")
library(platypus)

spec <- platypus_spec(
  data = segmentation_data("images/", "images/", colormap = binary_colormap),
  models = list(u_net("nuclei", input_shape = c(256, 256), blocks = 4, filters = 16,
                      weights = "dsbowl-unet", fit = FALSE))
)
masks <- predict(platypus_fit(spec), split = "validation")

Fetching by name needs pip install "pyplatypus[hub]". A local path to a downloaded file works without it, which is the route for machines with no internet access.


dsbowl-unet

A plain U-Net that separates cell nuclei from background in light microscopy images.

Architecture U-Net, 4 blocks, 16 base filters
Input 256 × 256, 3 channels
Output 2 classes: background, nucleus
Parameters 1,942,594
File 7.8 MB, safetensors
Loss CCE-Dice

What it was trained on

BBBC038v1, the 2018 Data Science Bowl image set, stage1_train, from the Broad Bioimage Benchmark Collection. 670 images of nuclei in varied cell types, magnifications and imaging modalities, split 536 / 134 by image with seed 1.

The data is CC0: "the various contributors of the imagesets have waived all copyright and related or neighboring rights to BBBC038v1". Taken from the Broad rather than from the Kaggle mirror on purpose — the images are the same, but the Kaggle copy is governed by competition rules accepted at download, and weights are a derivative of the data behind them.

Citation, which CC0 does not require and which is owed anyway:

Caicedo et al., Nucleus segmentation across imaging experiments: the 2018 Data Science Bowl, Nature Methods, 2019. Image set BBBC038v1, Broad Bioimage Benchmark Collection.

How it scores

On the 134 held-out images, per image rather than pooled:

metric mean sd median min max
Dice 0.9205 0.0532 0.9314 0.7240 0.9861
IoU 0.8570 0.0865 0.8716 0.5675 0.9725

The minimum is there deliberately. A mean of 0.92 with a worst case of 0.72 is a different model from a mean of 0.92 with a worst case of 0.90, and only one of those numbers tells you which you have. The five worst images are listed in dsbowl-unet.json.

Why a U-Net, and not one of the others

pyplatypus builds four architectures from one specification, so the question was settled by running all four on this data - same split, same augmentation, same 60 epochs, one command:

model parameters s/epoch epochs Dice mean sd median min
U-Net 1,942,594 15.5 60 0.9217 0.054 0.9353 0.7227
U-Net++ 2,263,730 18.8 44 0.9212 0.055 0.9323 0.7459
LinkNet 1,746,754 14.0 60 0.9203 0.053 0.9293 0.7205
Res-U-Net 2,029,682 14.2 55 0.9187 0.058 0.9293 0.6706

The spread across all four is 0.0030. The spread across images within any one of them is 0.054. The architectures are eighteen times closer to each other than the images are, which means they are not distinguishable on this data.

One number makes it plainer. These published weights score 0.9205; the U-Net in the table, same configuration and a different random seed, scored 0.9217. That 0.0012 of seed noise is 40% of the entire spread between the four architectures — switching from LinkNet to U-Net++ buys about what re-running the same model buys.

So only one set of weights is published, and the choice of U-Net is arbitrary but measured. If you are deciding where to spend effort on a problem like this one, the table says it is not on the architecture: more data, better augmentation or a longer run will all move the number further than swapping the decoder.

Reproduce it with python examples/compare_dsbowl_architectures.py.

What it does not do

It does not separate touching nuclei. This is semantic segmentation: every nucleus pixel gets the same class, so two nuclei in contact come back as one region. The 2018 Data Science Bowl was scored on instances, and the numbers above are not comparable to that leaderboard — they answer "which pixels are nuclei", not "how many nuclei are there". If you need counts, or per-nucleus measurements, you need an instance step on top (watershed on a distance transform is the usual starting point) and these weights are a reasonable input to it, not a replacement for it.

It has seen one collection. BBBC038v1 is varied by design, which is the point of the challenge, but generalisation was measured on a held-out split of the same collection and nothing else. On a microscope or a stain unlike anything in it, measure before trusting.

It is an alpha. pyplatypus is pre-1.0 and its API still moves. The weights will keep working - a pinned commit does not change - but the code that loads them may.

Reproducing it

pip install "pyplatypus[hub]"
python examples/publish_dsbowl_weights.py --out weights/

The script downloads BBBC038v1 from the Broad, splits it with the same seed, trains, measures per image, and writes the weights with this provenance in their sidecar. Trained on a GTX 1070: 60 epochs, about 15 seconds each.


bccd-yolo3

A YOLOv3 that finds red blood cells, white blood cells and platelets in blood-smear photographs.

Architecture YOLOv3, Darknet-53 backbone, 3 anchors per grid
Input 416 × 416, 3 channels, letterboxed
Output 3 classes: RBC, WBC, Platelets
Parameters 61,534,504
File 246 MB, safetensors
Anchors fitted to this data, recorded in the sidecar

The anchors are not optional. A detector's weights decode boxes relative to the anchors they were trained with, so the same numbers read with other anchors produce boxes scaled by a fixed factor - plausible boxes, plausible scores, wrong places, and nothing in the output to say so. pyplatypus takes them from the sidecar when it loads, and refuses a specification that names its own.

What it was trained on

BCCD: 364 blood-smear photographs with 4,888 boxes, in the dataset's own train/validation/test split of 205 / 87 / 72. MIT licence, copyright 2017 shenggan, with the original images and annotations from cosmicad and akshaylamba, re-organised into Pascal VOC format.

Taken from that repository rather than the Kaggle mirror. The images are the same, but the Kaggle copy is governed by competition rules accepted at download, and weights are a derivative of the data behind them — the same reasoning as dsbowl-unet above.

The split is the dataset's own, which is the only reason these numbers can be compared with anybody else's on BCCD.

How it scores

On the 72 held-out test images, at IoU 0.5 and averaged over 0.50–0.95, with precision and recall read at confidence 0.5:

class AP@0.5 IoU of matches truth precision recall
RBC 0.7934 0.810 805 0.715 0.793
WBC 0.9587 0.849 71 0.922 1.000
Platelets 0.8191 0.715 69 0.543 0.913
all 0.8571 0.8058 945

mAP@[.50:.95] is 0.5003. The gap between that and 0.8571 is localisation: the objects are found and the boxes fit them loosely, which mean_matched_iou says in one number instead of through the difference between two.

These are the median seed, and that matters here

Five runs of this exact recipe, differing only in the seed:

mean sd min max
mAP@0.5 0.8660 0.0159 0.8533 0.8868
mAP@[.50:.95] 0.5159 0.0199 0.5003 0.5436
IoU of matches 0.8041 0.0026 0.8002 0.8068

Published here is the median by mAP@0.5, not the best. The spread is wide enough that a single reported number would be whichever seed got chosen: the best run scores 0.8868, three points above this one, on the same code and the same data.

And one row of that table is six times steadier than another. How well the boxes fit barely varies between seeds; how many are found does. Which rank the model gives a borderline cell moves average precision and leaves the overlap alone. On a dataset this size, read mean_matched_iou first.

What it does not do

Platelets are the unreliable class, and that is where almost all of the variance lives. Across the five seeds:

class AP spread IoU spread
RBC ± 0.0100 ± 0.0048
WBC ± 0.0052 ± 0.0049
Platelets ± 0.0409 ± 0.0208

Platelets are small — a median side of 41 pixels against 103 for a red cell — and their boxes fit worst (IoU 0.715) and vary four to eight times more than the other two classes. Precision on them at confidence 0.5 is 0.543, meaning roughly half the platelets it reports are not there. If platelets are what you are counting, these weights are a starting point and not an answer.

It has seen one small collection. 205 training images of blood smears at one magnification, stained one way. BCCD is a demonstration dataset, not a clinical one, and generalisation was measured on its own held-out split and nowhere else.

It is not a diagnostic tool. Nothing about this was validated for clinical use, and a cell count from it is not a laboratory result.

About 0.1% of the training boxes could not be represented at all — 3 of 2,804. Two cells of similar shape whose centres land in the same grid cell share a target slot, so the second is dropped: never shown to the model, never counted as missed. At 416 on this data it is negligible, and pyplatypus reports the figure before training rather than after.

It is an alpha. The weights will keep working — a pinned commit does not change — but the code that loads them still moves, and detection is not in a released pyplatypus yet.

Reproducing it

# BCCD from https://github.com/Shenggan/BCCD_Dataset
python examples/detect_blood_cells.py --data path/to/BCCD --epochs 150 --seed 1 \
    --save weights/bccd-yolo3.safetensors

150 epochs, about 37 minutes on a GTX 1070. The script writes the YAML of the run beside its results, so the recipe in bccd-yolo3.json can be read back and run again.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support