File size: 6,649 Bytes
275b5a1
 
 
b74f2b3
 
 
 
275b5a1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
923fdff
 
 
275b5a1
 
923fdff
275b5a1
 
923fdff
 
 
 
275b5a1
 
 
923fdff
275b5a1
 
 
 
 
 
 
 
23ee5ab
923fdff
 
 
 
23ee5ab
 
 
 
275b5a1
 
 
 
 
 
 
 
 
 
923fdff
 
 
 
 
 
275b5a1
 
b74f2b3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
# SWaG Empty-8 Training Bundle

`Train/` is a portable download, training, and sampling bundle.
The download, training, and sampling workflow does not use absolute local
paths, `~/.seisbench`, or the SeisBench Python API. The optional distribution
evaluator below intentionally uses SeisBench's EQTransformer and its model
cache. The default workflow downloads already processed data from the Hugging
Face Dataset `DancingNow/swag-processed-data`; it does not reprocess anything.
Checkpoints, metrics, and generated samples stay below `Train/`.

## Model interface

The model is the no-high-resolution-skip SWaG backbone with eight reserved
continuous condition slots:

```python
prediction = model(noisy_waveforms, diffusion_timesteps, conditions)
```

`conditions` has shape `[batch, 8]`. During base training all eight values are
zero. The condition projection is zero-initialized, preserving a clean base
initialization while reserving the interface for later transfer learning.

## Training data interface

The processed STEAD files use this waveform interface and component order:

```text
data:   float32 [N, 6000, 3]  # ENZ, 100 Hz, 60 seconds
labels: float32 [N, 2]        # [P sample, S sample]
```

Each 6000-sample window is standardized to zero mean and unit standard
deviation. No bandpass filter is applied. Only rows marked
`earthquake_local` with valid P/S arrivals are written to the processed STEAD
files; noise is never used for training.

The Iquique files are downloaded from the `iquique/` directory of the Hugging
Face dataset `DancingNow/swag-processed-data`. They were processed previously
from the original SeisBench Iquique release and already use the same
`[N,6000,3]` ENZ waveform interface and P/S labels. They are downloaded as-is
and are not processed again by this bundle's default workflow.

## Install

Run from the directory containing `Train/`:

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r Train/requirements.txt
```

The system must also provide the `hf` command and a Hugging Face login. Use a
new virtual environment as shown; reusing an environment with incompatible
preinstalled PyTorch packages can cause import errors unrelated to this bundle.

For a private dataset, log in before downloading:

```bash
hf auth login
```

For CUDA training, install the PyTorch build matching the machine's CUDA driver
if the default pip build is not suitable.

## Download Processed Data

Download the already processed STEAD and Iquique files:

```bash
bash Train/run_download_prepare.sh
```

Outputs are written to:

```text
Train/data/processed/from_raw/stead/train/stead_100hz_60s_train.h5
Train/data/processed/from_raw/stead/test/stead_100hz_60s_test.h5
Train/data/processed/from_raw/iquique/
```

The command downloads from `DancingNow/swag-processed-data` into the exact
directory expected by the training configuration. It performs no cropping,
standardization, label filtering, or other conversion.

The repository is currently Private, so the target computer must be logged in
to the `DancingNow` account (or otherwise have read access). The download is
resumable through the Hugging Face cache; rerunning the command is safe.

To use another revision or destination:

```bash
HF_REVISION=main \
HF_LOCAL_DIR=/path/to/project/Train/data/processed/from_raw \
bash Train/run_download_prepare.sh
```

## Training

The default runnable configuration uses 8 GPUs, a global batch size of 256
(32 samples per GPU), no gradient accumulation, 15 epochs, and saves
checkpoints every 5 epochs:

```bash
bash Train/run_train_local.sh --num-gpus 8
```

The GPU count is a launcher parameter. For example, use `--num-gpus 1`,
`--num-gpus 4`, or `--num-gpus 8`. It defaults to 8 and may also be set with
`NUM_GPUS`. The global batch size must be divisible by the GPU count.

For a new machine, the complete download-to-training sequence can be started with:

```bash
bash Train/run_all_local.sh --num-gpus 8
```

The training entrypoint itself is independent of a scheduler:

```bash
python Train/training/train.py --config Train/configs/stead_empty8_local.yaml
```

The transfer-learning configuration in `Train/Transfer_and_Test/configs/config.yaml`
also uses a global training batch size of 256. Generation uses a global batch
size of 256 in both local and transfer-learning launchers. Training and
generation distribute their work across the requested GPUs and only the main
process writes the final checkpoint or merged HDF5 output.

If another machine cannot fit 256 waveforms in GPU memory, reduce the configured
batch size before running. If distributed training is required, launch this same
Python entrypoint with that environment's own distributed launcher or job
scheduler. No scheduler-specific submission script is included.

## Generate 100 samples

After epoch 15 finishes:

```bash
bash Train/run_generate_100.sh
```

This also defaults to 8 GPUs. To select a different count:

```bash
bash Train/run_generate_100.sh --num-gpus 4
```

This performs ancestral DDPM sampling with 25 steps and seed 2026. The output
is `Train/results/stead_empty8_local/generated_100/generated_ddpm25_seed2026.h5`.

## Unconditional distribution evaluation

The bundle also includes an unconditional real-vs-generated distribution
evaluator. It ignores labels and conditions, extracts hidden features from the
SeisBench EQTransformer, and computes FID, PRDC Density, Coverage, precision,
and recall. The default comparison uses 10,000 randomly selected waveforms per
side, `nearest_k=5`, and seed 2026:

```bash
python Train/Evaluation/evaluate_distribution.py \
  --real-h5 Train/data/processed/from_raw/stead/test/stead_100hz_60s_test.h5 \
  --fake-h5 Train/results/stead_empty8_local/generated_100/generated_ddpm25_seed2026.h5 \
  --output-dir Train/results/stead_empty8_local/distribution_eval \
  --sample-size 10000
```

The output is `distribution_metrics_eqt.csv`, together with the cached feature
arrays. This is an unconditional distribution comparison; no P/S labels are
read. Use `--num-gpus 8` to run the EQT feature extractor with DataParallel on
8 visible GPUs. The default uses the bundled
`Train/Evaluation/eqtransformer_original_state_dict.pt`, which is an exact
state-dict copy of `seisbench.models.EQTransformer.from_pretrained("original")`
and is the same feature model used by the archived SWaG distribution
evaluation.

The file `Train/Transfer_and_Test/eqt/best_val_state_dict.pt` is an Iquique
transfer-learning EQT picker for phase-pick evaluation. It is not the model
used for the unconditional FID/Density/Coverage comparison.