Mask Centered Damage Net (MCDN)

Python PyTorch Codebase Dataset

License: AGPL v3

Mask Centered Damage Net (MCDN) is a ~33.5M-parameter unitemporal structural damage classifier designed for use on sUAS orthomosaics, built on a ConvNeXt v2 Nano backbone. It grades individual structures on the four-level Joint Damage Scale (No Damage, Minor, Major, Destroyed) using only a post-disaster sUAS orthomosaic and a cache of building footprints, with no pre-disaster imagery.

Most automated damage assessment is bitemporal: a Siamese network compares pre- and post-event satellite imagery and grades the change. Applied to sUAS post-disaster imagery, that approach assumes a pre-event raster can be delivered to the point of analysis, which a disaster zone with degraded communications generally cannot do, and would struggle to compare a sub-5 cm/px sUAS orthomosaic to a satellite (30–80 cm/px) or crewed (15–30 cm/px) prior. MCDN replaces the pre-event raster with the structure footprint, a vector prior small enough to cache on a field laptop in advance and already available for most of the built world. The footprint localizes, centering each input chip on the structure being evaluated, and directs attention: the rasterized mask is ingested as a fourth input channel and separately weights the model's spatial pooling, so features under the roof dominate the pooled representation and surrounding debris contributes less. A FiLM gate conditions the head on disaster typology, and a squared Earth Mover's Distance loss preserves the ordinal structure of the grades.

This repository holds the trained weights. Code, evaluation scripts, per-seed metrics, and the full README live in the MCDN repository; the training data is CRASAR/CRASAR-U-DROIDs.

Headline performance

MCDN was evaluated on four holdouts: the CRASAR-U-DROIDs default train/test split, whose test events are all absent from training, and three Leave One Event Out (LOEO) holdouts that each withhold a single event. Per-seed columns are mean ± SD over the ten seeds published here; ensemble columns average those ten networks under 8-view D4 test-time augmentation.

Holdout Per-seed QWK Per-seed Macro-F1 Ensemble QWK Ensemble Macro-F1
DROIDs default split 0.865 ± 0.002 0.768 ± 0.007 0.869 0.774
LOEO Hurricane Michael 0.861 ± 0.005 0.779 ± 0.005 0.861 0.779
LOEO Mayfield Tornado 0.849 ± 0.007 0.751 ± 0.014 0.862 0.768
LOEO Hurricane Ida 0.752 ± 0.010 0.686 ± 0.008 0.762 0.693

10-seed ensemble confusion matrices on the four holdouts, row-normalized

Three of the four holdouts fall within 0.01 of one another. Hurricane Ida is the exception, and its deficit traces to a single class boundary: most of Ida's severe errors are structures with intact roofs surrounded by storm surge, labeled Minor or Major for damage to the interior and lower structure that a nadir roof view does not show, so the model under-grades them as No Damage. Across all holdouts, errors concentrate on adjacent grades; on the default split, 727 of 820 errors are one grade off while only 93 (2.4% of structures) are two or more.

A single network without test-time augmentation scores within 0.02 QWK of the ten-seed ensemble on every holdout; ensembling contributes seed-to-seed stability more than accuracy. On a desktop RTX 5090 a single network grades about 830 structures per second (108 with TTA) and the full 80-pass ensemble just under 10, with peak GPU memory at or below 3.3 GB.

What is here

370 best_model.pt files across 10 ablation arms and 37 arm/holdout cells (42.9 GB), ten fixed seeds per cell (0, 11, 22, 33, 44, 55, 66, 77, 88, 99). Each seed directory also carries the config_resolved.yaml the run trained under, with its run-bookkeeping paths (data.dir, checkpoint_root) normalized to the code repository's layout; every other field is verbatim.

<arm>/<holdout>/seed_<NN>/best_model.pt
<arm>/<holdout>/seed_<NN>/config_resolved.yaml
MANIFEST.json      # every file with size and SHA-256
MANIFEST.sha256    # sha256sum -c compatible

Holdout directories name the event(s) withheld from training: Hurricane_Ida, Hurricane_Michael, Mayfield_Tornado, and the dataset's default test split Hurricane_Idalia+Hurricane_Michael+Mayfield_Tornado+Mussett_Bayou_Fire.

Arms

C = footprint input channel, P = mask-weighted pooling, T = typology FiLM. Directory names follow the training presets, which name the component removed; the label column follows the paper.

Directory Paper label Components Holdouts Checkpoints
all_features Full configuration C+P+T all 4 40
pooling_typology Pooling + typology P+T all 4 40
typology No typology C+P all 4 40
pooling_only Pooling only P all 4 40
mask_channel_only Channel only C all 4 40
mask Typology only T all 4 40
rgb_only RGB only none all 4 40
ce_loss Cross-entropy loss C+P+T all 4 40
no_smoothing No label smoothing C+P+T all 4 40
resolution Crewed-aircraft imagery (15–30 cm/px), full configuration C+P+T Hurricane_Michael 10

Fetching

From a clone of the code repository, which places files where the evaluation scripts expect them, verifies every checkpoint against MANIFEST.json, and skips files already present:

python -m scripts.fetch_checkpoints --list
python -m scripts.fetch_checkpoints --arm all_features --fold Hurricane_Michael
python -m scripts.fetch_checkpoints --all

Or directly:

from huggingface_hub import snapshot_download

snapshot_download("mobilesensorlab/mcdn", allow_patterns=["all_features/Hurricane_Michael/*"], local_dir="outputs/ablation")

Reproducing the published metrics needs the dataset's sUAS test pool (~66 GB) alongside the weights; the code repository's scripts.fetch_dataset --sensor uas --split test retrieves exactly that.

Loading

Checkpoints are PyTorch state dicts for src.model.mcdn.MCDN; build the model from the seed's config_resolved.yaml and call load_state_dict. See src/postproc/ensemble.py in the code repository for the ten-seed, eight-view test-time-augmentation ensemble used for every reported number, and scripts/profile_mcdn_inference.py for single-network inference.

Citation

The accompanying manuscript, Mask Centered Damage Net: Building Damage Classification from Unitemporal sUAS Imagery and Footprint Priors (A. Kaplan and E. Best, Mobile Sensor Lab, University at Albany), is under review at IEEE JSTARS. Until it is published, please cite this repository and the code repository; a DOI and BibTeX entry will be added on acceptance.

License

MCDN and its checkpoints licensed under the GNU Affero General Public License v3.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mobilesensorlab/MCDN

Finetuned
(2)
this model

Dataset used to train mobilesensorlab/MCDN