File size: 6,521 Bytes
f1d3656 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 | ---
license: other
language:
- en
- zh
tags:
- OneScience
- Earth Science
- Weather Forecasting
- ERA5
- Representation Learning
frameworks: PyTorch
datasets:
- OneScience/ERA5
---
<p align="center">
<strong><span style="font-size: 30px;">AtmoRep</span></strong>
</p>
# Model Overview
AtmoRep is a stochastic atmospheric dynamics model based on large-scale representation learning that captures the distribution of atmospheric states through masked-token training and ensemble outputs.
Paper: *AtmoRep: A stochastic model of atmosphere dynamics using large scale representation learning*
https://arxiv.org/abs/2308.13280
# Model Description
This directory retains the official vorticity single-field model weights, configuration, and normalization, and provides a tiny AtmoRep-style model for local training and inference verification.
# Use Cases
| Scenario | Description |
| :---: | :--- |
| Official Resource Validation | Load official `.mod` weights and verify configuration. |
| Local Rapid Verification | Run masked-token training and inference with the tiny model. |
| ERA5 Atmospheric Representation Learning | Subsequently interface with official GRIB or Zarr data. |
# Usage
## 1. OneCode
[Click to experience intelligent one-click AI4S programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation & Usage
**Hardware Requirements**
- The tiny model runs on CPU.
- GPU is recommended for the official model and real-data inference.
### Download the Model Package
```bash
hf download --model OneScience-Group/AtmoRep --local-dir ./AtmoRep
cd AtmoRep
```
### Set Up the Runtime Environment
**DCU Environment**
```bash
conda create -n onescience311 python=3.11 -y
conda activate onescience311
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
**GPU Environment**
```bash
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
If dependencies are missing from the official paths, install them with:
```bash
pip install zarr wandb cfgrib xarray dask netCDF4 torchinfo
```
### Weights & Data
The current directory preserves:
```text
resources/id4nvwbetz/AtmoRep_id4nvwbetz.mod
resources/id4nvwbetz/model_id4nvwbetz.json
resources/data/normalization/vorticity/
```
Before re-downloading official resources please note: the current vendor snapshot is not a Git checkout, and the existing download script cannot re-clone into a non-empty `vendor/atmorep-official`. The resources bundled with this package do not require re-downloading; the script is only suitable for an empty target directory.
```bash
bash scripts/download_official_resources.sh .
```
### Tiny Training
```bash
python scripts/train.py
```
This command performs multi-epoch masked-token training using independent train/validation fake Datasets, including DataLoader, AdamW, validation, learning rate scheduling, early stopping, best/latest checkpointing, and training history. Default parameters reside in `conf/config.yaml`.
Resuming training:
```bash
python scripts/train.py --resume weight/training/latest.pth --epochs 20
```
Training artifacts are `weight/training/latest.pth`, `best.pth`, and `history.json`; the inference-compatible weights `weight/tiny_atmorep.pth` are also updated. Each Dataset sample consists of `fields [T,V,H,W]` and a non-empty `mask [N]`; train/validation use different seeds. This is a complete training pipeline for the tiny model, not a reproduction of the paper's 3.5-billion-parameter official model training.
### Tiny Inference
```bash
python scripts/inference.py
```
Inference results are saved as:
```text
result/prediction.pt
result/target.pt
```
These include the ensemble, ensemble mean, ensemble std, and mask.
### Result Inspection
```bash
python scripts/result.py
```
This command generates `result/metrics.json` and `result/comparison.png`. The reported metrics are ensemble/mean/spread RMSE in normalized token space, not the paper's physical-unit RMSE, ACC, CRPS, or spread-skill.
### Paper vs. Current Implementation I/O
| Item | Paper / Official Model | Tiny Smoke Model |
| --- | --- | --- |
| Input | ERA5 local 4D neighborhoods, 5 mode levels, multiple physical fields | `[B,4,1,8,8]` single-field random tensor |
| Token | Variable-correlated 4D tokens with absolute space-time and level conditioning | `1×4×4` patches, 16 tokens, relative coordinates with single-level conditioning |
| Output | Multi-head ensembles supporting reconstruction, nowcasting, and interpolation | 4-member masked-token ensemble |
| Training | Large-scale masked-token distribution learning | Multi-epoch Dataset training with independent validation and checkpoint resumption |
| Weights | `resources/id4nvwbetz` official vorticity weights | `weight/tiny_atmorep.pth`; the two are mutually incompatible |
The complete tiny execution flow is `train.py -> inference.py -> result.py`. The result analysis reads `weight/training/history.json` when present; inference output includes `ensemble`, `ensemble_mean`, `ensemble_std`, `mask`, and `target`. Random data is generated by the Dataset by index; there is currently no independent fake dataset that can be used with the official Zarr sampler. The model package is distributed without local training weights or `result/` artifacts — these are created at the paths described above after running the commands.
### Official Real-Data Inference
A directly runnable official real-data inference command is not yet available. It further requires ERA5 vorticity GRIB/Zarr data, the ecCodes environment, and parameterization of the file paths in the official `evaluate.py`.
### Real Data
The official vorticity model requires ERA5 vorticity, model levels 96/105/114/123/137, hourly temporal axes, and 0.25° global GRIB/Zarr data.
# OneScience Official Information
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
# Citation & License
- Official code is MIT License.
- Official model weights are declared CC BY 4.0.
- ERA5 is subject to Copernicus/ECMWF data terms.
|