File size: 8,232 Bytes
a13cbcb aec40c8 e8f3769 a13cbcb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 | ---
license: agpl-3.0
library_name: pytorch
inference: false
datasets:
- CRASAR/CRASAR-U-DROIDs
base_model:
- timm/convnextv2_nano.fcmae_ft_in22k_in1k
tags:
- remote-sensing
- building-damage-assessment
- disaster-response
- suas
- ordinal-classification
- convnextv2
---
# Mask Centered Damage Net (MCDN)
[](https://www.python.org/)
[](https://pytorch.org/)
[](https://github.com/MobileSensorLab/MCDN)
[](https://huggingface.co/datasets/CRASAR/CRASAR-U-DROIDs)
[](https://www.gnu.org/licenses/agpl-3.0)
Mask Centered Damage Net (MCDN) is a ~33.5M-parameter unitemporal structural damage classifier designed for use on sUAS
orthomosaics, built on a [ConvNeXt v2 Nano](https://huggingface.co/timm/convnextv2_nano.fcmae_ft_in22k_in1k) backbone.
It grades individual structures on the four-level Joint Damage Scale (*No Damage*, *Minor*, *Major*, *Destroyed*) using only
a post-disaster sUAS orthomosaic and a cache of building footprints, with no pre-disaster imagery.
Most automated damage assessment is bitemporal: a Siamese network compares pre- and post-event satellite imagery and grades
the change. Applied to sUAS post-disaster imagery, that approach assumes a pre-event raster can be delivered to the point of
analysis, which a disaster zone with degraded communications generally cannot do, and would struggle to compare a sub-5 cm/px
sUAS orthomosaic to a satellite (30–80 cm/px) or crewed (15–30 cm/px) prior. MCDN replaces the pre-event raster with the
structure footprint, a vector prior small enough to cache on a field laptop in advance and already available for most of the
built world. The footprint localizes, centering each input chip on the structure being evaluated, and directs attention: the
rasterized mask is ingested as a fourth input channel and separately weights the model's spatial pooling, so features under
the roof dominate the pooled representation and surrounding debris contributes less. A FiLM gate conditions the head on
disaster typology, and a squared Earth Mover's Distance loss preserves the ordinal structure of the grades.
This repository holds the trained weights. Code, evaluation scripts, per-seed metrics, and the full README live in the
[MCDN repository](https://github.com/MobileSensorLab/MCDN); the training data is
[CRASAR/CRASAR-U-DROIDs](https://huggingface.co/datasets/CRASAR/CRASAR-U-DROIDs).
## Headline performance
MCDN was evaluated on four holdouts: the CRASAR-U-DROIDs default train/test split, whose test events are all absent from
training, and three Leave One Event Out (LOEO) holdouts that each withhold a single event. Per-seed columns are mean ± SD over
the ten seeds published here; ensemble columns average those ten networks under 8-view D4 test-time augmentation.
| Holdout | Per-seed QWK | Per-seed Macro-F1 | Ensemble QWK | Ensemble Macro-F1 |
|---|:---:|:---:|:---:|:---:|
| DROIDs default split | 0.865 ± 0.002 | 0.768 ± 0.007 | **0.869** | **0.774** |
| LOEO Hurricane Michael | 0.861 ± 0.005 | 0.779 ± 0.005 | **0.861** | **0.779** |
| LOEO Mayfield Tornado | 0.849 ± 0.007 | 0.751 ± 0.014 | **0.862** | **0.768** |
| LOEO Hurricane Ida | 0.752 ± 0.010 | 0.686 ± 0.008 | **0.762** | **0.693** |

Three of the four holdouts fall within 0.01 of one another. Hurricane Ida is the exception, and its deficit traces to a single
class boundary: most of Ida's severe errors are structures with intact roofs surrounded by storm surge, labeled *Minor* or
*Major* for damage to the interior and lower structure that a nadir roof view does not show, so the model under-grades them as
*No Damage*. Across all holdouts, errors concentrate on adjacent grades; on the default split, 727 of 820 errors are one grade
off while only 93 (2.4% of structures) are two or more.
A single network without test-time augmentation scores within 0.02 QWK of the ten-seed ensemble on every holdout; ensembling
contributes seed-to-seed stability more than accuracy. On a desktop RTX 5090 a single network grades about 830 structures per
second (108 with TTA) and the full 80-pass ensemble just under 10, with peak GPU memory at or below 3.3 GB.
## What is here
370 `best_model.pt` files across 10 ablation arms and 37 arm/holdout cells
(42.9 GB), ten fixed seeds per cell (0, 11, 22, 33, 44, 55, 66, 77, 88, 99). Each seed directory also
carries the `config_resolved.yaml` the run trained under, with its run-bookkeeping paths (`data.dir`, `checkpoint_root`)
normalized to the code repository's layout; every other field is verbatim.
```
<arm>/<holdout>/seed_<NN>/best_model.pt
<arm>/<holdout>/seed_<NN>/config_resolved.yaml
MANIFEST.json # every file with size and SHA-256
MANIFEST.sha256 # sha256sum -c compatible
```
Holdout directories name the event(s) withheld from training: `Hurricane_Ida`, `Hurricane_Michael`, `Mayfield_Tornado`, and
the dataset's default test split `Hurricane_Idalia+Hurricane_Michael+Mayfield_Tornado+Mussett_Bayou_Fire`.
### Arms
C = footprint input channel, P = mask-weighted pooling, T = typology FiLM. Directory names follow the training presets, which
name the component *removed*; the label column follows the paper.
| Directory | Paper label | Components | Holdouts | Checkpoints |
|---|---|---|---|---|
| `all_features` | Full configuration | C+P+T | all 4 | 40 |
| `pooling_typology` | Pooling + typology | P+T | all 4 | 40 |
| `typology` | No typology | C+P | all 4 | 40 |
| `pooling_only` | Pooling only | P | all 4 | 40 |
| `mask_channel_only` | Channel only | C | all 4 | 40 |
| `mask` | Typology only | T | all 4 | 40 |
| `rgb_only` | RGB only | none | all 4 | 40 |
| `ce_loss` | Cross-entropy loss | C+P+T | all 4 | 40 |
| `no_smoothing` | No label smoothing | C+P+T | all 4 | 40 |
| `resolution` | Crewed-aircraft imagery (15–30 cm/px), full configuration | C+P+T | `Hurricane_Michael` | 10 |
## Fetching
From a clone of the code repository, which places files where the evaluation scripts expect them, verifies every checkpoint
against `MANIFEST.json`, and skips files already present:
```bash
python -m scripts.fetch_checkpoints --list
python -m scripts.fetch_checkpoints --arm all_features --fold Hurricane_Michael
python -m scripts.fetch_checkpoints --all
```
Or directly:
```python
from huggingface_hub import snapshot_download
snapshot_download("mobilesensorlab/mcdn", allow_patterns=["all_features/Hurricane_Michael/*"], local_dir="outputs/ablation")
```
Reproducing the published metrics needs the dataset's sUAS test pool (~66 GB) alongside the weights; the code repository's
`scripts.fetch_dataset --sensor uas --split test` retrieves exactly that.
## Loading
Checkpoints are PyTorch state dicts for `src.model.mcdn.MCDN`; build the model from the seed's `config_resolved.yaml` and call
`load_state_dict`. See `src/postproc/ensemble.py` in the code repository for the ten-seed, eight-view test-time-augmentation
ensemble used for every reported number, and `scripts/profile_mcdn_inference.py` for single-network inference.
## Citation
The accompanying manuscript, *Mask Centered Damage Net: Building Damage Classification from Unitemporal sUAS Imagery and
Footprint Priors* (A. Kaplan and E. Best, Mobile Sensor Lab, University at Albany), is under review at IEEE JSTARS. Until it
is published, please cite this repository and the [code repository](https://github.com/MobileSensorLab/MCDN); a DOI and BibTeX entry will be added on
acceptance.
## License
MCDN and its checkpoints licensed under the GNU Affero General Public License v3.0.
|