--- license: cc-by-4.0 datasets: - risenyard/egms-qa-dataset tags: - safetensors - remote-sensing - insar - ground-motion - time-series - egms --- # EGMS-QA Encoder Pretrained encoder for EGMS displacement time series within 7 km tiles. It produces 256-dimensional point representations, which are spatially pooled into 65 tile tokens. This repository contains the released weights, model settings, training recipe, and evaluation results. ## Use the model See the [architecture on GitHub](https://github.com/risenyard/egms-qa/blob/main/src/egms_encoder/README.md#architecture), then follow [installation and a one-tile run](https://github.com/risenyard/egms-qa/blob/main/src/egms_encoder/README.md#installation) to try the encoder. The same guide covers [local inputs](https://github.com/risenyard/egms-qa/blob/main/src/egms_encoder/README.md#use-local-inputs) and [training, resuming, and token extraction](https://github.com/risenyard/egms-qa/blob/main/src/egms_encoder/README.md#reproduce-training). Prepared tiles and precomputed tokens are available in the [Dataset repository](https://huggingface.co/datasets/risenyard/egms-qa-dataset#encoder-data). ## Files | file | purpose | |---|---| | [encoder.safetensors](encoder.safetensors) | encoder weights | | [config.json](config.json) | model architecture and input dimensions | | [normalization.json](normalization.json) | input mean, standard deviation, and residual scale | | [training_args.json](training_args.json) | training recipe and checkpoint-selection record | | [eval_results.json](eval_results.json) | masked-reconstruction metrics | Inference requires the weights, model config, and normalization. The Dataset repository provides the measurements and split manifest. ## Input requirements | input or output | contract | |---|---| | tile displacement | vertical displacement in mm, `[N,294]` | | coordinates | EPSG:3035 easting and northing in meters, `[N,2]` | | model preprocessing | checkpoint normalization and centered coordinates | | point representations | `[N,256]` | | pooled tokens | `[65,256]` and a 65-element validity mask | The Dataset stores `[0,294)`, corresponding to `[8,302)` on the 304-step source-preparation axis. Its data config retains the source offset and six-day cadence for physical-time calculations. Use this checkpoint's normalization for inference. When training a new encoder on another corpus, fit normalization on that corpus's training split. ## Evaluation Evaluation covers 1,000 held-out tiles with 2,047,451 point histories. A central 88-step interval, approximately 30% of the 294-step input, is masked at the same positions for every point in a tile. | metric | value | |---|---:| | normalized MSE | 0.0702 | | MSE | 2.279 mm² | | RMSE | 1.510 mm | | MAE | 1.007 mm | | pooled EGMS residual | 1.433 mm | | per-point RMSE P10 / P50 / P90 | 0.54 / 1.03 / 2.43 mm | These values describe reconstruction of held-out observations under the specified masking protocol. ## Scope and license The encoder consumes prepared EGMS-QA tiles. Official-product downloading, format conversion, and preparation of another reference period require a separate workflow. Its outputs describe observed deformation and do not establish causes, predict future motion, or certify structural safety. The encoder is released under CC-BY-4.0.