Feature Extraction
Transformers
Safetensors
English
remote-sensing
earth-observation
self-supervised-learning
satellite
multispectral
convnext
mae
mmearth
mp-mae
Instructions to use Haight/MMEarth-transformers with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Haight/MMEarth-transformers with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="Haight/MMEarth-transformers")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Haight/MMEarth-transformers", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 6,419 Bytes
5a3f0b7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 | ---
license: mit
language:
- en
tags:
- remote-sensing
- earth-observation
- self-supervised-learning
- satellite
- multispectral
- feature-extraction
- convnext
- mae
- mmearth
- mp-mae
- transformers
library_name: transformers
pipeline_tag: feature-extraction
---
# MMEarth Transformers Models
Hugging Face–compatible checkpoints converted from the official [MMEarth](https://arxiv.org/abs/2405.02771) MP-MAE pretrained weights. Each subfolder is a standalone model repo layout (`config.json`, `model.safetensors`, preprocessor, and remote code) for geospatial feature extraction.
## Model Description
These models are ConvNeXt V2 encoders pretrained with Multi Pretext Masked Autoencoding (MP-MAE) on the [MMEarth](https://github.com/vishalned/MMEarth-data) multi-modal geospatial dataset. Checkpoints cover different pretext task configurations (all modalities, S2-only, RGB/BGR, image-level, pixel-level) and model sizes (atto, tiny).
All folders ship self-contained remote code (`modeling_mmearth.py`, processor, pipeline) and load with `trust_remote_code=True`.
**Developed by:** [MMEarth Authors](https://github.com/vishalned/MMEarth-train)
**Converted for Hugging Face by:** BiliSakura
**License (weights):** MIT
**Original paper:** [MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning](https://arxiv.org/abs/2405.02771) (ECCV 2024)
## Available checkpoints (10 models)
| Folder | Input | Size | Dataset | Loss | Image | Patch | Ch |
|--------|-------|------|---------|------|-------|-------|----|
| `mmearth-convnextv2-atto-all-mod-1m-64-uncertainty-56x8` | all_mod | atto | 1M_64 | uncertainty | 56 | 8 | 12 |
| `mmearth-convnextv2-atto-all-mod-1m-64-unweighted-56x8` | all_mod | atto | 1M_64 | unweighted | 56 | 8 | 12 |
| `mmearth-convnextv2-atto-all-mod-1m-128-uncertainty-112x16` | all_mod | atto | 1M_128 | uncertainty | 112 | 16 | 12 |
| `mmearth-convnextv2-atto-all-mod-100k-128-uncertainty-112x16` | all_mod | atto | 100k_128 | uncertainty | 112 | 16 | 12 |
| `mmearth-convnextv2-tiny-all-mod-1m-64-uncertainty-56x8` | all_mod | tiny | 1M_64 | uncertainty | 56 | 8 | 12 |
| `mmearth-convnextv2-atto-s2-1m-64-uncertainty-56x8` | S2 | atto | 1M_64 | uncertainty | 56 | 8 | 12 |
| `mmearth-convnextv2-atto-rgb-1m-64-uncertainty-56x8` | rgb (BGR) | atto | 1M_64 | uncertainty | 56 | 8 | 3 |
| `mmearth-convnextv2-atto-rgb-1m-128-uncertainty-112x16` | rgb (BGR) | atto | 1M_128 | uncertainty | 112 | 16 | 3 |
| `mmearth-convnextv2-atto-img-mod-1m-64-uncertainty-56x8` | img_mod | atto | 1M_64 | uncertainty | 56 | 8 | 12 |
| `mmearth-convnextv2-atto-pix-mod-1m-64-uncertainty-56x8` | pix_mod | atto | 1M_64 | uncertainty | 56 | 8 | 12 |
Legacy `.pth` filename mapping is in [`conversion_manifest.json`](conversion_manifest.json).
## Usage
Processors default to **`do_resize: false`**. Inputs keep native height and width. Apply per-band MMEarth normalization when you have dataset statistics (`image_mean` / `image_std`).
```python
from transformers import pipeline
import numpy as np
MODEL = "/path/to/MMEarth-transformers/mmearth-convnextv2-atto-rgb-1m-64-uncertainty-56x8"
pipe = pipeline(
task="mmearth-feature-extraction",
model=MODEL,
trust_remote_code=True,
)
# RGB/BGR: 3 bands at native size (56×56 for this checkpoint)
image = np.random.rand(56, 56, 3).astype(np.float32) * 1000
features = pipe(image, pool=True, return_tensors=True)
print(features.shape) # torch.Size([1, 320])
```
12-band Sentinel-2 (all_mod / S2 checkpoints):
```python
MODEL = "/path/to/MMEarth-transformers/mmearth-convnextv2-atto-all-mod-1m-64-uncertainty-56x8"
pipe = pipeline(task="mmearth-feature-extraction", model=MODEL, trust_remote_code=True)
image = np.random.rand(56, 56, 12).astype(np.float32) * 1000
features = pipe(image, pool=True, return_tensors=True)
print(features.shape) # torch.Size([1, 320])
```
Dense spatial token map:
```python
tokens = pipe(image, pool=False, return_tensors=True)
print(tokens.shape) # [1, num_patches, hidden_size]
```
To resize to the pretraining reference size:
```python
features = pipe(image, pool=True, return_tensors=True, image_processor_kwargs={"do_resize": True})
```
Load components directly:
```python
from transformers import AutoModel, AutoImageProcessor
model = AutoModel.from_pretrained(MODEL, trust_remote_code=True)
processor = AutoImageProcessor.from_pretrained(MODEL, trust_remote_code=True)
```
## Custom pipeline
Each checkpoint registers a custom pipeline in `config.json`:
```json
"custom_pipelines": {
"mmearth-feature-extraction": {
"impl": "pipeline_mmearth.MMEarthImageFeatureExtractionPipeline",
"pt": ["AutoModel"]
}
}
```
This follows the [HuggingFace custom pipeline pattern](https://huggingface.co/docs/transformers/add_new_pipeline): remote code ships with the model folder, and `trust_remote_code=True` loads `MMEarthImageFeatureExtractionPipeline`, which extends the standard `ImageFeatureExtractionPipeline` with numpy array and file path support.
The built-in `image-feature-extraction` task also works:
```python
pipe = pipeline(task="image-feature-extraction", model=MODEL, trust_remote_code=True)
```
## Normalization
MMEarth pretraining normalizes each band with dataset-specific mean/std from `data_*_band_stats.json`. The converted preprocessor defaults to `do_normalize: false` because band statistics are not embedded in the legacy checkpoints. Provide your own `image_mean` / `image_std` when preprocessing:
```python
features = pipe(
image,
pool=True,
return_tensors=True,
image_processor_kwargs={
"do_normalize": True,
"image_mean": [...], # one value per channel
"image_std": [...],
},
)
```
RGB checkpoints were trained with **BGR** channel order (bands B4, B3, B2). The processor swaps RGB→BGR when `channel_order="bgr"`.
## Dependencies
- `transformers`, `torch`, `timm`, `safetensors`
- `opencv-python` (multispectral resize with more than 4 channels when `do_resize=True`)
## Citation
```bibtex
@inproceedings{nedungadi2024mmearth,
title={MMEarth: Exploring multi-modal pretext tasks for geospatial representation learning},
author={Nedungadi, Vishal and Kariryaa, Ankit and Oehmcke, Stefan and Belongie, Serge and Igel, Christian and Lang, Nico},
booktitle={European Conference on Computer Vision},
pages={164--182},
year={2024},
organization={Springer}
}
```
|