--- license: agpl-3.0 library_name: pytorch inference: false datasets: - CRASAR/CRASAR-U-DROIDs base_model: - timm/convnextv2_nano.fcmae_ft_in22k_in1k tags: - remote-sensing - building-damage-assessment - disaster-response - suas - ordinal-classification - convnextv2 --- # Mask Centered Damage Net (MCDN) [![Python](https://img.shields.io/badge/Python-3.13-3776AB?logo=python&logoColor=white)](https://www.python.org/) [![PyTorch](https://img.shields.io/badge/PyTorch-2.11-EE4C2C?logo=pytorch&logoColor=orange)](https://pytorch.org/) [![Codebase](https://img.shields.io/badge/Codebase-MobileSensorLab%2FMCDN-181717?logo=github)](https://github.com/MobileSensorLab/MCDN) [![Dataset](https://img.shields.io/badge/%F0%9F%A4%97%20Dataset-CRASAR%2FCRASAR--U--DROIDs-FFD21E)](https://huggingface.co/datasets/CRASAR/CRASAR-U-DROIDs) [![License: AGPL v3](https://img.shields.io/badge/License-GNU%20AGPL%20v3-663366.svg?logo=gnu&logoColor=white)](https://www.gnu.org/licenses/agpl-3.0) Mask Centered Damage Net (MCDN) is a ~33.5M-parameter unitemporal structural damage classifier designed for use on sUAS orthomosaics, built on a [ConvNeXt v2 Nano](https://huggingface.co/timm/convnextv2_nano.fcmae_ft_in22k_in1k) backbone. It grades individual structures on the four-level Joint Damage Scale (*No Damage*, *Minor*, *Major*, *Destroyed*) using only a post-disaster sUAS orthomosaic and a cache of building footprints, with no pre-disaster imagery. Most automated damage assessment is bitemporal: a Siamese network compares pre- and post-event satellite imagery and grades the change. Applied to sUAS post-disaster imagery, that approach assumes a pre-event raster can be delivered to the point of analysis, which a disaster zone with degraded communications generally cannot do, and would struggle to compare a sub-5 cm/px sUAS orthomosaic to a satellite (30–80 cm/px) or crewed (15–30 cm/px) prior. MCDN replaces the pre-event raster with the structure footprint, a vector prior small enough to cache on a field laptop in advance and already available for most of the built world. The footprint localizes, centering each input chip on the structure being evaluated, and directs attention: the rasterized mask is ingested as a fourth input channel and separately weights the model's spatial pooling, so features under the roof dominate the pooled representation and surrounding debris contributes less. A FiLM gate conditions the head on disaster typology, and a squared Earth Mover's Distance loss preserves the ordinal structure of the grades. This repository holds the trained weights. Code, evaluation scripts, per-seed metrics, and the full README live in the [MCDN repository](https://github.com/MobileSensorLab/MCDN); the training data is [CRASAR/CRASAR-U-DROIDs](https://huggingface.co/datasets/CRASAR/CRASAR-U-DROIDs). ## Headline performance MCDN was evaluated on four holdouts: the CRASAR-U-DROIDs default train/test split, whose test events are all absent from training, and three Leave One Event Out (LOEO) holdouts that each withhold a single event. Per-seed columns are mean ± SD over the ten seeds published here; ensemble columns average those ten networks under 8-view D4 test-time augmentation. | Holdout | Per-seed QWK | Per-seed Macro-F1 | Ensemble QWK | Ensemble Macro-F1 | |---|:---:|:---:|:---:|:---:| | DROIDs default split | 0.865 ± 0.002 | 0.768 ± 0.007 | **0.869** | **0.774** | | LOEO Hurricane Michael | 0.861 ± 0.005 | 0.779 ± 0.005 | **0.861** | **0.779** | | LOEO Mayfield Tornado | 0.849 ± 0.007 | 0.751 ± 0.014 | **0.862** | **0.768** | | LOEO Hurricane Ida | 0.752 ± 0.010 | 0.686 ± 0.008 | **0.762** | **0.693** | ![10-seed ensemble confusion matrices on the four holdouts, row-normalized](assets/confusion_triptych.svg) Three of the four holdouts fall within 0.01 of one another. Hurricane Ida is the exception, and its deficit traces to a single class boundary: most of Ida's severe errors are structures with intact roofs surrounded by storm surge, labeled *Minor* or *Major* for damage to the interior and lower structure that a nadir roof view does not show, so the model under-grades them as *No Damage*. Across all holdouts, errors concentrate on adjacent grades; on the default split, 727 of 820 errors are one grade off while only 93 (2.4% of structures) are two or more. A single network without test-time augmentation scores within 0.02 QWK of the ten-seed ensemble on every holdout; ensembling contributes seed-to-seed stability more than accuracy. On a desktop RTX 5090 a single network grades about 830 structures per second (108 with TTA) and the full 80-pass ensemble just under 10, with peak GPU memory at or below 3.3 GB. ## What is here 370 `best_model.pt` files across 10 ablation arms and 37 arm/holdout cells (42.9 GB), ten fixed seeds per cell (0, 11, 22, 33, 44, 55, 66, 77, 88, 99). Each seed directory also carries the `config_resolved.yaml` the run trained under, with its run-bookkeeping paths (`data.dir`, `checkpoint_root`) normalized to the code repository's layout; every other field is verbatim. ``` //seed_/best_model.pt //seed_/config_resolved.yaml MANIFEST.json # every file with size and SHA-256 MANIFEST.sha256 # sha256sum -c compatible ``` Holdout directories name the event(s) withheld from training: `Hurricane_Ida`, `Hurricane_Michael`, `Mayfield_Tornado`, and the dataset's default test split `Hurricane_Idalia+Hurricane_Michael+Mayfield_Tornado+Mussett_Bayou_Fire`. ### Arms C = footprint input channel, P = mask-weighted pooling, T = typology FiLM. Directory names follow the training presets, which name the component *removed*; the label column follows the paper. | Directory | Paper label | Components | Holdouts | Checkpoints | |---|---|---|---|---| | `all_features` | Full configuration | C+P+T | all 4 | 40 | | `pooling_typology` | Pooling + typology | P+T | all 4 | 40 | | `typology` | No typology | C+P | all 4 | 40 | | `pooling_only` | Pooling only | P | all 4 | 40 | | `mask_channel_only` | Channel only | C | all 4 | 40 | | `mask` | Typology only | T | all 4 | 40 | | `rgb_only` | RGB only | none | all 4 | 40 | | `ce_loss` | Cross-entropy loss | C+P+T | all 4 | 40 | | `no_smoothing` | No label smoothing | C+P+T | all 4 | 40 | | `resolution` | Crewed-aircraft imagery (15–30 cm/px), full configuration | C+P+T | `Hurricane_Michael` | 10 | ## Fetching From a clone of the code repository, which places files where the evaluation scripts expect them, verifies every checkpoint against `MANIFEST.json`, and skips files already present: ```bash python -m scripts.fetch_checkpoints --list python -m scripts.fetch_checkpoints --arm all_features --fold Hurricane_Michael python -m scripts.fetch_checkpoints --all ``` Or directly: ```python from huggingface_hub import snapshot_download snapshot_download("mobilesensorlab/mcdn", allow_patterns=["all_features/Hurricane_Michael/*"], local_dir="outputs/ablation") ``` Reproducing the published metrics needs the dataset's sUAS test pool (~66 GB) alongside the weights; the code repository's `scripts.fetch_dataset --sensor uas --split test` retrieves exactly that. ## Loading Checkpoints are PyTorch state dicts for `src.model.mcdn.MCDN`; build the model from the seed's `config_resolved.yaml` and call `load_state_dict`. See `src/postproc/ensemble.py` in the code repository for the ten-seed, eight-view test-time-augmentation ensemble used for every reported number, and `scripts/profile_mcdn_inference.py` for single-network inference. ## Citation The accompanying manuscript, *Mask Centered Damage Net: Building Damage Classification from Unitemporal sUAS Imagery and Footprint Priors* (A. Kaplan and E. Best, Mobile Sensor Lab, University at Albany), is under review at IEEE JSTARS. Until it is published, please cite this repository and the [code repository](https://github.com/MobileSensorLab/MCDN); a DOI and BibTeX entry will be added on acceptance. ## License MCDN and its checkpoints licensed under the GNU Affero General Public License v3.0.