--- license: apache-2.0 library_name: pytorch tags: - multimodal - semantic-segmentation - vision - cvpr-2023 - cmnext - segformer - swin-transformer pipeline_tag: image-segmentation --- # CMNeXt: Delivering Arbitrary-Modal Semantic Segmentation [CVPR 2023] Official Model Checkpoints Repository for **CMNeXt**. - **Project Page:** https://jamycheung.github.io/DELIVER.html - **GitHub Repository:** https://github.com/InSAI-Lab/DELIVER - **Paper:** [arXiv:2303.01480](https://arxiv.org/abs/2303.01480) - **Dataset Repository:** [InSAI-Lab/DELIVER](https://huggingface.co/datasets/InSAI-Lab/DELIVER) --- ## Model Description **CMNeXt** is an arbitrary cross-modal semantic segmentation model operating in the Hub2Fuse paradigm with asymmetric branches: - Multi-Head Self-Attention (MHSA) blocks in the RGB branch - Parallel Pooling Mixer (PPX) blocks in the accompanying modality branch - Self-Query Hub selects informative supplementary features - Feature Rectification Module (FRM) and Feature Fusion Module (FFM) fuse features dynamically across 1 to 81 modalities ## Checkpoints Organization This repository contains trained weights and pretrained backbones matching the official DELIVER codebase structure: ```text ├── DELIVER/ │ ├── cmnext_b2_deliver_rgb.pth │ ├── cmnext_b2_deliver_rgbd.pth │ ├── cmnext_b2_deliver_rgbde.pth │ ├── cmnext_b2_deliver_rgbdel.pth │ ├── cmnext_b2_deliver_rgbdl.pth │ ├── cmnext_b2_deliver_rgbe.pth │ └── cmnext_b2_deliver_rgbl.pth ├── KITTI360/ │ ├── cmnext_b2_kitti360_rgb.pth │ ├── cmnext_b2_kitti360_rgbd.pth │ ├── cmnext_b2_kitti360_rgbde.pth │ ├── cmnext_b2_kitti360_rgbdel.pth │ ├── cmnext_b2_kitti360_rgbdl.pth │ ├── cmnext_b2_kitti360_rgbe.pth │ └── cmnext_b2_kitti360_rgbl.pth ├── MCubeS/ │ ├── cmnext_b2_mcubes_rgb.pth │ ├── cmnext_b2_mcubes_rgba.pth │ ├── cmnext_b2_mcubes_rgbad.pth │ └── cmnext_b2_mcubes_rgbadn.pth ├── MFNet/ │ └── cmnext_b4_mfnet_rgbt.pth ├── NYU_Depth_V2/ │ └── cmnext_b4_nyu_rgbd.pth ├── UrbanLF/ │ ├── cmnext_b4_urbanlf_real_rgblf1.pth │ ├── cmnext_b4_urbanlf_real_rgblf33.pth │ ├── cmnext_b4_urbanlf_real_rgblf8.pth │ ├── cmnext_b4_urbanlf_real_rgblf80.pth │ ├── cmnext_b4_urbanlf_syn_rgblf1.pth │ ├── cmnext_b4_urbanlf_syn_rgblf33.pth │ ├── cmnext_b4_urbanlf_syn_rgblf8.pth │ └── cmnext_b4_urbanlf_syn_rgblf80.pth └── pretrained/ ├── segformers/ │ ├── mit_b0.pth ~ mit_b5.pth └── swintransformer/ ├── swin_base_patch4_window12_384_22k.pth ├── swin_large_patch4_window12_384_22k.pth └── swin_small_patch4_window7_224.pth ``` ## Benchmark Results ### DELIVER Benchmark | Model-Modal | #Params(M) | GFLOPs | mIoU (%) | Checkpoint | | :--- | :--- | :--- | :--- | :--- | | CMNeXt-RGB | 25.79 | 38.93 | 57.20 | `DELIVER/cmnext_b2_deliver_rgb.pth` | | CMNeXt-RGB-E | 58.69 | 62.94 | 57.48 | `DELIVER/cmnext_b2_deliver_rgbe.pth` | | CMNeXt-RGB-L | 58.69 | 62.94 | 58.04 | `DELIVER/cmnext_b2_deliver_rgbl.pth` | | CMNeXt-RGB-D | 58.69 | 62.94 | 63.58 | `DELIVER/cmnext_b2_deliver_rgbd.pth` | | CMNeXt-RGB-D-E | 58.72 | 64.19 | 64.44 | `DELIVER/cmnext_b2_deliver_rgbde.pth` | | CMNeXt-RGB-D-L | 58.72 | 64.19 | 65.50 | `DELIVER/cmnext_b2_deliver_rgbdl.pth` | | CMNeXt-RGB-D-E-L | 58.73 | 65.42 | **66.30** | `DELIVER/cmnext_b2_deliver_rgbdel.pth` | ### Other Benchmarks - **KITTI-360:** CMNeXt-RGB-D-E-L achieves 67.84% mIoU - **NYU Depth V2:** CMNeXt-RGB-D (MiT-B4) achieves 56.90% mIoU - **MFNet (RGB-T):** CMNeXt-RGB-T (MiT-B4) achieves 59.90% mIoU - **UrbanLF:** Up to 83.22% mIoU (Real) / 81.02% mIoU (Synthetic) - **MCubeS:** CMNeXt-RGB-A-D-N achieves 51.54% mIoU ## Download & Usage Using `hf`: ```bash # Download all checkpoints into output directory hf download InSAI-Lab/CMNeXt --local-dir output/ ``` Evaluate with DELIVER repository: ```bash cd DELIVER CUDA_VISIBLE_DEVICES=0 python tools/val_mm.py --cfg configs/deliver_rgbdel.yaml ``` ## Citation ```bibtex @inproceedings{zhang2023delivering, title={Delivering Arbitrary-Modal Semantic Segmentation}, author={Zhang, Jiaming and Liu, Ruiping and Shi, Hao and Yang, Kailun and Rei{\ss}, Simon and Peng, Kunyu and Fu, Haodong and Wang, Kaiwei and Stiefelhagen, Rainer}, booktitle={CVPR}, year={2023} } ```