leobk commited on
Commit
cf24457
·
verified ·
1 Parent(s): 7bb00d9

Add v2 microridge model: weights, plans, and provenance record

Browse files

2D nnU-Net (Dataset502_MicroridgeField, nnUNetTrainer_100epochs, fold 0)
segmenting background / cell_region / cell_membrane / microridge.

Frozen test over 3 unseen fields: cell_region Dice 0.944, microridge Dice
0.877, cell_membrane boundary F1 0.659.

The checkpoint is unmodified so registry.json's checkpoint_sha256 verifies.

README.md CHANGED
@@ -1,3 +1,163 @@
1
  ---
2
  license: cc-by-nc-sa-4.0
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: cc-by-nc-sa-4.0
3
+ library_name: nnunetv2
4
+ pipeline_tag: image-segmentation
5
+ tags:
6
+ - biology
7
+ - microscopy
8
+ - cell-segmentation
9
+ - microridge
10
+ - nnunet
11
+ - zebrafish
12
  ---
13
+
14
+ # MicroridgeVectorAI — v2
15
+
16
+ A 2D nnU-Net that segments **actin microridges**, **cell regions** and **cell
17
+ membranes** in projected single-channel microscopy of epithelial tissue.
18
+
19
+ Companion application: <https://github.com/LBK888/MicroridgeVectorAI> (CellVector,
20
+ AGPL-3.0). The model runs standalone with nnU-Net v2 alone — CellVector is not
21
+ required.
22
+
23
+ | | |
24
+ |---|---|
25
+ | Model id | `8f1dbd06-6ec8-4c22-8e26-412cfee26ea9` |
26
+ | Task | 2D semantic segmentation, 4 classes |
27
+ | Architecture | nnU-Net v2 PlainConvUNet, 8 stages, patch 512x512 |
28
+ | Trainer / folds | `nnUNetTrainer_100epochs`, fold 0 |
29
+ | Input | single-channel 2D image, any size |
30
+ | Snapshot hash | `dbe134f83c52b8ecae6cb62182310205496497ec297406ed8a2c912b60ba8cc9` |
31
+
32
+ Labels: `0` background, `1` cell_region, `2` cell_membrane, `3` microridge.
33
+
34
+ ## Scores
35
+
36
+ Frozen test — 36 tiles from **3 fields the model never saw**. Splits are grouped
37
+ by source field, so no tile of a training field appears in the test set.
38
+
39
+ | Metric | Value |
40
+ |---|---|
41
+ | cell_region Dice | 0.944 |
42
+ | cell_membrane Dice | 0.471 |
43
+ | cell_membrane boundary F1 (1 px tolerance) | 0.659 |
44
+ | microridge Dice | 0.877 |
45
+ | microridge precision / recall | 0.911 / 0.855 |
46
+ | microridge skeleton length error | 0.100 |
47
+
48
+ nnU-Net's own fold-0 validation (89 tiles): cell_region 0.964, cell_membrane
49
+ 0.495, microridge 0.912.
50
+
51
+ **Reading the membrane number.** A Dice of 0.47 on a 3-pixel line covering ~1%
52
+ of the frame is not the same failure as 0.47 on a region class: thin-structure
53
+ Dice collapses when a prediction is offset by a pixel even where it follows the
54
+ right path. The boundary F1 of 0.659, which allows one pixel of tolerance, is
55
+ the more informative figure, and the gap between them says the membrane is
56
+ mostly in the right place but not pixel-exact.
57
+
58
+ ## Limitations
59
+
60
+ - **Cell contours derived from the label map merge.** Reconstructing cells as
61
+ the connected components of `label in (1, 3)` yields a single blob, because
62
+ the predicted membrane is thin and not perfectly closed. Use the predicted
63
+ membrane (class 2) for cell geometry, or split the region with a watershed
64
+ seeded inside cells. Do not expect instance-separated cells out of the box.
65
+ - **The ground truth was not human-reviewed.** Labels were imported from
66
+ published raster masks and corrected only for import artifacts, not by an
67
+ expert. Treat this model as a proposal generator to be corrected, which is how
68
+ the companion application uses it.
69
+ - **Trained on 13 fields.** Train and validation loss diverge (-0.755 vs
70
+ -0.629), which is what a small number of independent acquisitions looks like.
71
+ More fields will help more than more epochs.
72
+ - **One fold, not an ensemble.** Only fold 0 was trained.
73
+ - Validated on zebrafish periderm-style epithelial microridge imagery. Behaviour
74
+ on other tissue, magnification or modality is unknown.
75
+
76
+ ## Training data
77
+
78
+ Wide-field frames cut into 477 tiles of at most 512x512 from 19 fields, keeping
79
+ only regions whose raster truth is trustworthy. Uneven illumination leaves part
80
+ of such a frame too dark for the upstream segmentation to resolve anything, and
81
+ that failure is silent — the skeleton mask is empty while the cell mask still
82
+ looks complete. Blocks were kept only where skeleton density cleared both an
83
+ absolute floor and a share of the frame's own 90th percentile, **and** at least
84
+ 95% of the block was attributed to a cell. 68.3% of the field pixels survived.
85
+
86
+ Labels were rasterized from vector geometry with a 3 px membrane and a 5 px
87
+ microridge stroke.
88
+
89
+ ## Files
90
+
91
+ ```text
92
+ registry.json provenance record, metrics, checksums
93
+ nnUNet_results/Dataset502_MicroridgeField/
94
+ └─ nnUNetTrainer_100epochs__nnUNetPlans__2d/
95
+ ├─ dataset.json channel names and label map
96
+ ├─ plans.json preprocessing and architecture
97
+ └─ fold_0/checkpoint_final.pth weights
98
+ ```
99
+
100
+ Those three files under the trainer folder are the complete inference set. The
101
+ directory names encode the configuration — nnU-Net parses
102
+ `Dataset<ID>_<name>/<trainer>__<plans>__<configuration>` — so do not rename them.
103
+
104
+ The checkpoint is shipped unmodified so the `checkpoint_sha256` in
105
+ `registry.json` verifies. About half of it is optimizer state; stripping to
106
+ `network_weights`, `init_args`, `trainer_name` and
107
+ `inference_allowed_mirroring_axes` halves the size but invalidates that
108
+ checksum.
109
+
110
+ ## Usage
111
+
112
+ ```bash
113
+ pip install nnunetv2 huggingface_hub
114
+ hf download leobk/MicroridgeVectorAI --local-dir microridge-model
115
+ ```
116
+
117
+ ```python
118
+ import torch, numpy as np, tifffile
119
+ from nnunetv2.inference.predict_from_raw_data import nnUNetPredictor
120
+
121
+ MODEL = ("microridge-model/nnUNet_results/Dataset502_MicroridgeField"
122
+ "/nnUNetTrainer_100epochs__nnUNetPlans__2d")
123
+
124
+ predictor = nnUNetPredictor(device=torch.device("cuda"))
125
+ predictor.initialize_from_trained_model_folder(
126
+ MODEL, use_folds=(0,), checkpoint_name="checkpoint_final.pth"
127
+ )
128
+
129
+ image = tifffile.imread("frame.tif").astype("float32")
130
+ segmentation = predictor.predict_single_npy_array(
131
+ image[None, None], {"spacing": (999.0, 1.0, 1.0)}, None, None, False
132
+ )
133
+ ```
134
+
135
+ No nnU-Net environment variables are needed for this path. Roughly 13 s for a
136
+ 512x512 tile on an RTX 4080 SUPER.
137
+
138
+ ### Reimplementing the pipeline
139
+
140
+ The network takes `(1, 1, H, W)` and returns 4 logit channels, and exports to
141
+ TorchScript. If you drive it yourself, reproduce all of:
142
+
143
+ - **Normalization** — z-score using *each image's own* mean and standard
144
+ deviation (`use_mask_for_norm=False`). No dataset statistics;
145
+ `foreground_intensity_properties_per_channel` in `plans.json` is for CT
146
+ normalization and unused here.
147
+ - **Sliding window** — 512x512 patches, step 0.5, Gaussian-weighted overlap.
148
+ - **Test-time augmentation** — mirroring over axes `(0, 1)`.
149
+ - **Output** — argmax over the 4 channels.
150
+
151
+ Skipping the normalization or the Gaussian window degrades results noticeably
152
+ and without any error.
153
+
154
+ ## Licensing note
155
+
156
+ These weights are released under **CC BY-NC-SA 4.0**: attribution required,
157
+ **non-commercial use only**, derivatives under the same terms. Note that this
158
+ differs from the companion application's code licence (AGPL-3.0) — the code and
159
+ the weights are covered separately.
160
+
161
+ The weights were trained on third-party imagery; if that source data carries its
162
+ own terms, they may constrain redistribution of this model independently of this
163
+ label.
nnUNet_results/Dataset502_MicroridgeField/nnUNetTrainer_100epochs__nnUNetPlans__2d/dataset.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "channel_names": {
3
+ "0": "actin_microridge"
4
+ },
5
+ "file_ending": ".tif",
6
+ "labels": {
7
+ "background": 0,
8
+ "cell_membrane": 2,
9
+ "cell_region": 1,
10
+ "microridge": 3
11
+ },
12
+ "numTraining": 441,
13
+ "overwrite_image_reader_writer": "NaturalImage2DIO"
14
+ }
nnUNet_results/Dataset502_MicroridgeField/nnUNetTrainer_100epochs__nnUNetPlans__2d/fold_0/checkpoint_final.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0e9e440280e99deffcd772cfde6e59aaa6e0f3578f59662a12246226692dbdde
3
+ size 370825919
nnUNet_results/Dataset502_MicroridgeField/nnUNetTrainer_100epochs__nnUNetPlans__2d/plans.json ADDED
@@ -0,0 +1,207 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "dataset_name": "Dataset502_MicroridgeField",
3
+ "plans_name": "nnUNetPlans",
4
+ "original_median_spacing_after_transp": [
5
+ 999.0,
6
+ 1.0,
7
+ 1.0
8
+ ],
9
+ "original_median_shape_after_transp": [
10
+ 1,
11
+ 512,
12
+ 491
13
+ ],
14
+ "image_reader_writer": "NaturalImage2DIO",
15
+ "transpose_forward": [
16
+ 0,
17
+ 1,
18
+ 2
19
+ ],
20
+ "transpose_backward": [
21
+ 0,
22
+ 1,
23
+ 2
24
+ ],
25
+ "configurations": {
26
+ "2d": {
27
+ "data_identifier": "nnUNetPlans_2d",
28
+ "preprocessor_name": "DefaultPreprocessor",
29
+ "batch_size": 12,
30
+ "patch_size": [
31
+ 512,
32
+ 512
33
+ ],
34
+ "median_image_size_in_voxels": [
35
+ 512.0,
36
+ 491.0
37
+ ],
38
+ "spacing": [
39
+ 1.0,
40
+ 1.0
41
+ ],
42
+ "normalization_schemes": [
43
+ "ZScoreNormalization"
44
+ ],
45
+ "use_mask_for_norm": [
46
+ false
47
+ ],
48
+ "resampling_fn_data": "resample_data_or_seg_to_shape",
49
+ "resampling_fn_seg": "resample_data_or_seg_to_shape",
50
+ "resampling_fn_data_kwargs": {
51
+ "is_seg": false,
52
+ "order": 3,
53
+ "order_z": 0,
54
+ "force_separate_z": null
55
+ },
56
+ "resampling_fn_seg_kwargs": {
57
+ "is_seg": true,
58
+ "order": 1,
59
+ "order_z": 0,
60
+ "force_separate_z": null
61
+ },
62
+ "resampling_fn_probabilities": "resample_data_or_seg_to_shape",
63
+ "resampling_fn_probabilities_kwargs": {
64
+ "is_seg": false,
65
+ "order": 1,
66
+ "order_z": 0,
67
+ "force_separate_z": null
68
+ },
69
+ "architecture": {
70
+ "network_class_name": "dynamic_network_architectures.architectures.unet.PlainConvUNet",
71
+ "arch_kwargs": {
72
+ "n_stages": 8,
73
+ "features_per_stage": [
74
+ 32,
75
+ 64,
76
+ 128,
77
+ 256,
78
+ 512,
79
+ 512,
80
+ 512,
81
+ 512
82
+ ],
83
+ "conv_op": "torch.nn.modules.conv.Conv2d",
84
+ "kernel_sizes": [
85
+ [
86
+ 3,
87
+ 3
88
+ ],
89
+ [
90
+ 3,
91
+ 3
92
+ ],
93
+ [
94
+ 3,
95
+ 3
96
+ ],
97
+ [
98
+ 3,
99
+ 3
100
+ ],
101
+ [
102
+ 3,
103
+ 3
104
+ ],
105
+ [
106
+ 3,
107
+ 3
108
+ ],
109
+ [
110
+ 3,
111
+ 3
112
+ ],
113
+ [
114
+ 3,
115
+ 3
116
+ ]
117
+ ],
118
+ "strides": [
119
+ [
120
+ 1,
121
+ 1
122
+ ],
123
+ [
124
+ 2,
125
+ 2
126
+ ],
127
+ [
128
+ 2,
129
+ 2
130
+ ],
131
+ [
132
+ 2,
133
+ 2
134
+ ],
135
+ [
136
+ 2,
137
+ 2
138
+ ],
139
+ [
140
+ 2,
141
+ 2
142
+ ],
143
+ [
144
+ 2,
145
+ 2
146
+ ],
147
+ [
148
+ 2,
149
+ 2
150
+ ]
151
+ ],
152
+ "n_conv_per_stage": [
153
+ 2,
154
+ 2,
155
+ 2,
156
+ 2,
157
+ 2,
158
+ 2,
159
+ 2,
160
+ 2
161
+ ],
162
+ "n_conv_per_stage_decoder": [
163
+ 2,
164
+ 2,
165
+ 2,
166
+ 2,
167
+ 2,
168
+ 2,
169
+ 2
170
+ ],
171
+ "conv_bias": true,
172
+ "norm_op": "torch.nn.modules.instancenorm.InstanceNorm2d",
173
+ "norm_op_kwargs": {
174
+ "eps": 1e-05,
175
+ "affine": true
176
+ },
177
+ "dropout_op": null,
178
+ "dropout_op_kwargs": null,
179
+ "nonlin": "torch.nn.LeakyReLU",
180
+ "nonlin_kwargs": {
181
+ "inplace": true
182
+ }
183
+ },
184
+ "_kw_requires_import": [
185
+ "conv_op",
186
+ "norm_op",
187
+ "dropout_op",
188
+ "nonlin"
189
+ ]
190
+ },
191
+ "batch_dice": true
192
+ }
193
+ },
194
+ "experiment_planner_used": "ExperimentPlanner",
195
+ "label_manager": "LabelManager",
196
+ "foreground_intensity_properties_per_channel": {
197
+ "0": {
198
+ "max": 18781.0,
199
+ "mean": 540.381103515625,
200
+ "median": 382.0,
201
+ "min": 104.0,
202
+ "percentile_00_5": 116.0,
203
+ "percentile_99_5": 2714.0,
204
+ "std": 473.2962646484375
205
+ }
206
+ }
207
+ }
registry.json ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "models": [
3
+ {
4
+ "architecture": "plainconv",
5
+ "artifact_sha256": {
6
+ "dataset.json": "aa6242b569dc3fa70ee707db38970fd7e97fbfb45f48aa7f1a68a1c7a5a899cf",
7
+ "dataset_fingerprint.json": "9234acf421a087a276d8271f214cbddfdd97c1b71f4974e18a71fe44f680c7b9",
8
+ "fold_0/checkpoint_best.pth": "189e5d97d34dd8750bf279e47a78b8038ebb31d95b72ddcf0425f55c2785eb75",
9
+ "fold_0/checkpoint_final.pth": "0e9e440280e99deffcd772cfde6e59aaa6e0f3578f59662a12246226692dbdde",
10
+ "plans.json": "7e403e7cb95ad1efbf184de013d651595561f248bd885d461d5b0f8fa2e73874"
11
+ },
12
+ "augmentation_profile_hash": "6f738a972b6f7eba84ee6a2cd5cbd01f83dbed2afa9633439b89e76bb3a5a4dd",
13
+ "backend_id": "nnunetv2-local",
14
+ "backend_version": "2.8.1",
15
+ "benchmark_id": null,
16
+ "checkpoint_path": "fold_0/checkpoint_final.pth",
17
+ "checkpoint_sha256": "0e9e440280e99deffcd772cfde6e59aaa6e0f3578f59662a12246226692dbdde",
18
+ "configuration": {
19
+ "configuration": "2d",
20
+ "dataset_id": 502,
21
+ "dataset_name": "MicroridgeField",
22
+ "epochs": 100,
23
+ "executed_locally": true,
24
+ "fold": 0,
25
+ "membrane_width_px": 3,
26
+ "microridge_width_px": 5,
27
+ "plan_command": "nnUNetv2_plan_and_preprocess -d 502 -c 2d --verify_dataset_integrity",
28
+ "plans": "nnUNetPlans",
29
+ "source_data": "full view field tiles, trustworthy regions only",
30
+ "split_grouping": "specimen:field_frame_uri",
31
+ "torch_version": "2.12.1+cu130",
32
+ "train_command": "nnUNetv2_train 502 2d 0 -tr nnUNetTrainer_100epochs --npz",
33
+ "trainer": "nnUNetTrainer_100epochs",
34
+ "training_job_id": "0305ef07-38d7-41d3-bc44-a1e51a2c0d4f"
35
+ },
36
+ "created_at": "2026-08-31T10:00:21.285927Z",
37
+ "dataset_artifact_hash": "90178985de2ef8d24b5a7338b7e9fec83dfe6164e0558ef30d449d8e13439547",
38
+ "folds_completed": [
39
+ 0
40
+ ],
41
+ "incomplete_provenance_fields": [],
42
+ "input_contract": "image/2d",
43
+ "label_contract": "cellvector.annotation/1.0.0",
44
+ "metrics": {
45
+ "frozen_test_cases": 36,
46
+ "frozen_test_cell_membrane_boundary_f1": 0.6589374464002449,
47
+ "frozen_test_cell_membrane_dice": 0.4708712201149024,
48
+ "frozen_test_cell_region_dice": 0.9438357584291484,
49
+ "frozen_test_cell_region_iou": 0.8940270075353971,
50
+ "frozen_test_microridge_dice": 0.8767592825579081,
51
+ "frozen_test_microridge_precision": 0.9106816936813859,
52
+ "frozen_test_microridge_recall": 0.8550323207264082,
53
+ "frozen_test_microridge_skeleton_length_error": 0.0997634178921583,
54
+ "nnunet_validation_dice_cell_membrane": 0.4954518160625487,
55
+ "nnunet_validation_dice_cell_region": 0.9644053946394331,
56
+ "nnunet_validation_dice_microridge": 0.9114529420774974
57
+ },
58
+ "model_id": "8f1dbd06-6ec8-4c22-8e26-412cfee26ea9",
59
+ "parent_model_id": null,
60
+ "provenance_status": "verified",
61
+ "selected_at": null,
62
+ "selection_actor": null,
63
+ "selection_reason": null,
64
+ "snapshot_hash": "dbe134f83c52b8ecae6cb62182310205496497ec297406ed8a2c912b60ba8cc9",
65
+ "software_smoke_test": false,
66
+ "status": "trained",
67
+ "training_job_id": "0305ef07-38d7-41d3-bc44-a1e51a2c0d4f",
68
+ "updated_at": "2026-08-31T10:00:21.285927Z"
69
+ }
70
+ ],
71
+ "schema_version": "1.0.0"
72
+ }