ResEncL-MissingPiece-VoCo

Copyright German Cancer Research Center (DKFZ) and contributors. Please make sure that your usage of these models is in compliance with their license.

MICCAI 2025

ResEncL-MissingPiece-VoCo is a ResEnc-L backbone for 3D medical object detection, pre-trained with volume contrastive learning (VoCo) on CT-RATE + ABCD Study. It is one of the checkpoints released with The Missing Piece: A Case for Pre-training in 3D Medical Object Detection (MICCAI 2025).

Model family

All checkpoints of the collection:

Model Pre-training Architecture Repository
ResEncL-MissingPiece-MAE MAE (self-supervised) ResEnc-L MIC-DKFZ/ResEncL-MissingPiece-MAE
ResEncL-MissingPiece-MG Models Genesis (self-supervised) ResEnc-L MIC-DKFZ/ResEncL-MissingPiece-MG
ResEncL-MissingPiece-S3D Spark 3D (self-supervised) ResEnc-L MIC-DKFZ/ResEncL-MissingPiece-S3D
ResEncL-MissingPiece-VoCo VoCo (self-supervised) ResEnc-L this repository
ResEncL-MissingPiece-MultiTalent MultiTalent (supervised) ResEnc-L MIC-DKFZ/ResEncL-MissingPiece-MultiTalent
RetinaUNet-MissingPiece-MultiTalent MultiTalent (supervised) Retina U-Net (nnDetection ConvBackbone + FPN) MIC-DKFZ/RetinaUNet-MissingPiece-MultiTalent

The nnFoundation models nnFoundationCNN (ResEnc-L) and nnFoundationViT (Primus) can be finetuned for detection with nnDetection in the same way -- see the finetuning docs.

Using this checkpoint

Fine-tune it for detection with nnDetection -- Finetuning pretrained backbones:

hf download MIC-DKFZ/ResEncL-MissingPiece-VoCo checkpoint_final.pth --local-dir ./checkpoints/ResEncL-MissingPiece-VoCo

nndet_train Task<XXX>_YourDataset residual_encoder_retinaunet_focal_v002 0 \
    -o module=RetinaUNetFocalV002_ResEnc_TL exp.tag=_VoCo \
    +transfer_learning_ckpt=./checkpoints/ResEncL-MissingPiece-VoCo/checkpoint_final.pth \
    --transfer_learning --load_adapt_plan

--load_adapt_plan rebuilds the backbone to match this checkpoint before loading the weights. For a Deformable DETR head instead, use residual_encoder_def_detr_v002 with -o module=BoxDeformableDETRV002_ResEnc_TL.

Repository contents

File Purpose
checkpoint_final.pth the pre-trained weights
adaptation_plan.json architecture + preprocessing plan; nnDetection reads it (from the checkpoint) to rebuild the backbone
config.json placeholder so the Hub records download counts

Expected input

3D volumes, preprocessed with nnDetection's standard pipeline (nndet_prep), using the dataset's default planned spacing. Downstream patch size: 128 × 128 × 128.

For reference, pre-training used Z-score normalised volumes resampled to 1 × 1 × 1 mm, with a patch size of 192 × 192 × 64.

Checkpoint format

checkpoint_final.pth is a torch.save dictionary that loads safely with weights_only=True:

Key Contents
network_weights the pre-trained state_dict (only the encoder and input stem are transferred downstream)
trainer_name the pre-training trainer
nnssl_adaptation_plan same content as adaptation_plan.json
citations the references to cite when using these weights (printed to the training log on load)

Note: this checkpoint stores a number of state_dict entries as aliases of the same underlying tensor (the all_modules.* keys mirror the named conv/norm modules). This is expected for ResEnc -- do not deduplicate these keys.

Citation

If you use these weights, please cite The Missing Piece (MICCAI 2025):

The Missing Piece BibTeX
@inproceedings{eckstein2025missing,
  title     = {The Missing Piece: A Case for Pre-training in 3D Medical Object Detection},
  author    = {Eckstein, Katharina and Ulrich, Constantin and Baumgartner, Michael and K{\"a}chele, Jessica and Bounias, Dimitrios and Wald, Tassilo and Floca, Ralf and Maier-Hein, Klaus H.},
  booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2025},
  series    = {Lecture Notes in Computer Science},
  volume    = {15963},
  pages     = {615--626},
  year      = {2025},
  publisher = {Springer Nature Switzerland},
  doi       = {10.1007/978-3-032-04965-0_58}
}

Please also cite the architecture, pre-training method and pre-training data behind this checkpoint:

Further references

Architecture -- ResEncL

  • Isensee, F., Wald, T., Ulrich, C., Baumgartner, M., Roy, S., Maier-Hein, K., & Jaeger, P. F. (2024). nnU-Net revisited: A call for rigorous validation in 3D medical image segmentation. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 (Lecture Notes in Computer Science, pp. 488–498). Springer. https://doi.org/10.1007/978-3-031-72114-4_47
  • Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203–211. https://doi.org/10.1038/s41592-020-01008-z

Pretraining Method -- VoCo

  • Wu, L., Zhuang, J., & Chen, H. (2024). VoCo: A simple-yet-effective volume contrastive learning framework for 3D medical image analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 22873–22882).

Pre-Training Dataset -- CT-RATE

  • Hamamci, I. E., Er, S., Wang, C., Almas, F., Simsek, A. G., Esirgun, S. N., ... & Menze, B. (2026). Generalist foundation models from a multimodal dataset for 3D computed tomography. Nature Biomedical Engineering, 10, 1610–1628. https://doi.org/10.1038/s41551-025-01599-y

Pre-Training Dataset -- ABCD Study

  • Casey, B. J., Cannonier, T., Conley, M. I., Cohen, A. O., Barch, D. M., Heitzeg, M. M., ... & Dale, A. M. (2018). The Adolescent Brain Cognitive Development (ABCD) study: Imaging acquisition across 21 sites. Developmental Cognitive Neuroscience, 32, 43–54. https://doi.org/10.1016/j.dcn.2018.03.001
  • Saragosa-Harris, N. M., Chaku, N., MacSweeney, N., Guazzelli Williamson, V., Scheuplein, M., Feola, B., ... & Michalska, K. J. (2022). A practical guide for researchers and reviewers using the ABCD Study and other large longitudinal datasets. Developmental Cognitive Neuroscience, 55, 101115. https://doi.org/10.1016/j.dcn.2022.101115

Framework -- nnssl

  • Wald, T., Ulrich, C., Lukyanenko, S., Goncharov, A., Paderno, A., Maerkisch, L., ... & Maier-Hein, K. (2024). Revisiting MAE pre-training for 3D medical image segmentation. CVPR.
Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train MIC-DKFZ/ResEncL-MissingPiece-VoCo

Collection including MIC-DKFZ/ResEncL-MissingPiece-VoCo

Papers for MIC-DKFZ/ResEncL-MissingPiece-VoCo