ResEncL-MissingPiece-VoCo
Copyright German Cancer Research Center (DKFZ) and contributors. Please make sure that your usage of these models is in compliance with their license.
ResEncL-MissingPiece-VoCo is a ResEnc-L backbone for 3D medical object detection, pre-trained with volume contrastive learning (VoCo) on CT-RATE + ABCD Study.
It is one of the checkpoints released with The Missing Piece: A Case for Pre-training in 3D Medical Object Detection (MICCAI 2025).
Model family
All checkpoints of the collection:
| Model | Pre-training | Architecture | Repository |
|---|---|---|---|
| ResEncL-MissingPiece-MAE | MAE (self-supervised) | ResEnc-L | MIC-DKFZ/ResEncL-MissingPiece-MAE |
| ResEncL-MissingPiece-MG | Models Genesis (self-supervised) | ResEnc-L | MIC-DKFZ/ResEncL-MissingPiece-MG |
| ResEncL-MissingPiece-S3D | Spark 3D (self-supervised) | ResEnc-L | MIC-DKFZ/ResEncL-MissingPiece-S3D |
| ResEncL-MissingPiece-VoCo | VoCo (self-supervised) | ResEnc-L | this repository |
| ResEncL-MissingPiece-MultiTalent | MultiTalent (supervised) | ResEnc-L | MIC-DKFZ/ResEncL-MissingPiece-MultiTalent |
| RetinaUNet-MissingPiece-MultiTalent | MultiTalent (supervised) | Retina U-Net (nnDetection ConvBackbone + FPN) |
MIC-DKFZ/RetinaUNet-MissingPiece-MultiTalent |
The nnFoundation models nnFoundationCNN (ResEnc-L) and nnFoundationViT (Primus) can be finetuned for detection with nnDetection in the same way -- see the finetuning docs.
Using this checkpoint
Fine-tune it for detection with nnDetection -- Finetuning pretrained backbones:
hf download MIC-DKFZ/ResEncL-MissingPiece-VoCo checkpoint_final.pth --local-dir ./checkpoints/ResEncL-MissingPiece-VoCo
nndet_train Task<XXX>_YourDataset residual_encoder_retinaunet_focal_v002 0 \
-o module=RetinaUNetFocalV002_ResEnc_TL exp.tag=_VoCo \
+transfer_learning_ckpt=./checkpoints/ResEncL-MissingPiece-VoCo/checkpoint_final.pth \
--transfer_learning --load_adapt_plan
--load_adapt_plan rebuilds the backbone to match this checkpoint before loading the weights. For a Deformable DETR head instead, use residual_encoder_def_detr_v002 with -o module=BoxDeformableDETRV002_ResEnc_TL.
Repository contents
| File | Purpose |
|---|---|
checkpoint_final.pth |
the pre-trained weights |
adaptation_plan.json |
architecture + preprocessing plan; nnDetection reads it (from the checkpoint) to rebuild the backbone |
config.json |
placeholder so the Hub records download counts |
Expected input
3D volumes, preprocessed with nnDetection's standard pipeline (nndet_prep), using the
dataset's default planned spacing.
Downstream patch size: 128 × 128 × 128.
For reference, pre-training used Z-score normalised volumes resampled to 1 × 1 × 1 mm, with a patch size of 192 × 192 × 64.
Checkpoint format
checkpoint_final.pth is a torch.save dictionary that loads safely with weights_only=True:
| Key | Contents |
|---|---|
network_weights |
the pre-trained state_dict (only the encoder and input stem are transferred downstream) |
trainer_name |
the pre-training trainer |
nnssl_adaptation_plan |
same content as adaptation_plan.json |
citations |
the references to cite when using these weights (printed to the training log on load) |
Note: this checkpoint stores a number of
state_dictentries as aliases of the same underlying tensor (theall_modules.*keys mirror the named conv/norm modules). This is expected for ResEnc -- do not deduplicate these keys.
Citation
If you use these weights, please cite The Missing Piece (MICCAI 2025):
The Missing Piece BibTeX
@inproceedings{eckstein2025missing,
title = {The Missing Piece: A Case for Pre-training in 3D Medical Object Detection},
author = {Eckstein, Katharina and Ulrich, Constantin and Baumgartner, Michael and K{\"a}chele, Jessica and Bounias, Dimitrios and Wald, Tassilo and Floca, Ralf and Maier-Hein, Klaus H.},
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2025},
series = {Lecture Notes in Computer Science},
volume = {15963},
pages = {615--626},
year = {2025},
publisher = {Springer Nature Switzerland},
doi = {10.1007/978-3-032-04965-0_58}
}
Please also cite the architecture, pre-training method and pre-training data behind this checkpoint:
Further references
Architecture -- ResEncL
- Isensee, F., Wald, T., Ulrich, C., Baumgartner, M., Roy, S., Maier-Hein, K., & Jaeger, P. F. (2024). nnU-Net revisited: A call for rigorous validation in 3D medical image segmentation. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 (Lecture Notes in Computer Science, pp. 488–498). Springer. https://doi.org/10.1007/978-3-031-72114-4_47
- Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203–211. https://doi.org/10.1038/s41592-020-01008-z
Pretraining Method -- VoCo
- Wu, L., Zhuang, J., & Chen, H. (2024). VoCo: A simple-yet-effective volume contrastive learning framework for 3D medical image analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 22873–22882).
Pre-Training Dataset -- CT-RATE
- Hamamci, I. E., Er, S., Wang, C., Almas, F., Simsek, A. G., Esirgun, S. N., ... & Menze, B. (2026). Generalist foundation models from a multimodal dataset for 3D computed tomography. Nature Biomedical Engineering, 10, 1610–1628. https://doi.org/10.1038/s41551-025-01599-y
Pre-Training Dataset -- ABCD Study
- Casey, B. J., Cannonier, T., Conley, M. I., Cohen, A. O., Barch, D. M., Heitzeg, M. M., ... & Dale, A. M. (2018). The Adolescent Brain Cognitive Development (ABCD) study: Imaging acquisition across 21 sites. Developmental Cognitive Neuroscience, 32, 43–54. https://doi.org/10.1016/j.dcn.2018.03.001
- Saragosa-Harris, N. M., Chaku, N., MacSweeney, N., Guazzelli Williamson, V., Scheuplein, M., Feola, B., ... & Michalska, K. J. (2022). A practical guide for researchers and reviewers using the ABCD Study and other large longitudinal datasets. Developmental Cognitive Neuroscience, 55, 101115. https://doi.org/10.1016/j.dcn.2022.101115
Framework -- nnssl
- Wald, T., Ulrich, C., Lukyanenko, S., Goncharov, A., Paderno, A., Maerkisch, L., ... & Maier-Hein, K. (2024). Revisiting MAE pre-training for 3D medical image segmentation. CVPR.
- Downloads last month
- 2