Release 1.2: completed 270-storm benchmark and revised paper
Browse files- .gitattributes +1 -0
- README.md +26 -19
- RELEASE_NOTES_TRACKFORMER_1_2.md +16 -4
- docs/daily_storm_benchmark.md +3 -1
- evaluation/daily_storm_final.json +0 -0
- evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png +3 -0
- paper/trackformer.pdf +2 -2
- paper/trackformer.tex +291 -281
- release_tools/build_paper_figures.py +6 -3
- release_tools/plot_daily_storm_final.py +35 -0
.gitattributes
CHANGED
|
@@ -26,3 +26,4 @@ docs/trackformer_1_2_yagi.mp4 filter=lfs diff=lfs merge=lfs -text
|
|
| 26 |
docs/trackformer_1_2_fung_wong.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 27 |
evaluation/trackformer_1_2_fung_wong_video_poster.png filter=lfs diff=lfs merge=lfs -text
|
| 28 |
evaluation/release_data/fung_wong_video.npz filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 26 |
docs/trackformer_1_2_fung_wong.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 27 |
evaluation/trackformer_1_2_fung_wong_video_poster.png filter=lfs diff=lfs merge=lfs -text
|
| 28 |
evaluation/release_data/fung_wong_video.npz filter=lfs diff=lfs merge=lfs -text
|
| 29 |
+
evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -11,7 +11,7 @@ tags:
|
|
| 11 |
|
| 12 |
**[Download 1.2: weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Read the illustrated paper](paper/trackformer.pdf)** · **[Watch Fung-wong](https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4)** · **[Hugging Face model](https://huggingface.co/euler314/typhoon-predict)**
|
| 13 |
|
| 14 |
-
Release status: **research
|
| 15 |
|
| 16 |
Trackformer 1.2 is a **research-only Western Pacific tropical-cyclone forecast model**. It evolves a sea-level-pressure (MSLP) field and a moving storm-centred pressure core every six hours through +120 h. A track and central-pressure estimate are extracted from the evolving core, not independently drawn on top of a pressure image. The prior [Trackformer 1.1 release](https://github.com/yu314-coder/typhoon-predict/releases/tag/trackformer-1.1) remains available and unchanged.
|
| 17 |
|
|
@@ -54,7 +54,30 @@ The implementation is in [`models/trackformer_1_2_field/`](models/trackformer_1_
|
|
| 54 |
|
| 55 |
Grid spacing describes the representation, not independently demonstrated effective resolution. A mean of member centres is not necessarily the minimum of the displayed mean pressure field.
|
| 56 |
|
| 57 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
|
| 59 |
This comparison uses **270 forecast cases from 90 storms**, with the same issue rows, observed routes and 20 six-hour leads through +120 h. The number 270 counts forecast cases. The 1.2 prediction for each case is a **mean of 50 causal input-perturbation members**; 1.1 uses its archived causal route pipeline.
|
| 60 |
|
|
@@ -82,23 +105,7 @@ These are **six selected best-performing examples from six distinct storms**, no
|
|
| 82 |
|
| 83 |
The 1.2 mean has better aggregate direction, shape and position scores on this previously inspected development cohort. That does not establish an overall win on every storm or pressure metric. A same-270 1.1 central-pressure prediction array was not verified in these route artifacts, so it is not assigned a pressure-error bar. The separate 1.2 report gives 15.73 hPa central-pressure MAE and 2.54 hPa area-weighted regional MSLP MAE. TIP remains a separate diagnostic: its pressure error was worse for 1.2 (30.7 vs 20.1 hPa). See [metric definitions and limitations](docs/trackformer_1_2_evaluation.md).
|
| 84 |
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
A **1,473-case / 270-storm** benchmark is in progress; it is separate from the completed 270-case comparison above. Each storm contributes at most one forecast per UTC day, through +120 h. Daily errors are averaged within each storm, then storm scores receive equal weight. Selection was frozen before inference without filtering on forecast quality. See the [protocol](docs/daily_storm_benchmark.md) and [frozen cohort](evaluation/release_data/daily_storm_cohort.json).
|
| 88 |
-
|
| 89 |
-
**Dated snapshot — 2026-09-29T06:09:06.508182+00:00 (UTC), not a live counter:** 808/1,473 daily forecasts saved; **159/270 storms fully completed**. The 1.2 forecasts use 50 causal input-perturbation members. [Machine-readable snapshot](evaluation/daily_storm_progress_20260929.json).
|
| 90 |
-
|
| 91 |
-
| Preliminary equal-storm metric | 1.1 | 1.2 · mean of 50 |
|
| 92 |
-
| --- | ---: | ---: |
|
| 93 |
-
| Mean track error, +6 to +120 h ↓ | 843.9 km | 487.1 km |
|
| 94 |
-
| Track error at +120 h ↓ | 1733.3 km | 1068.8 km |
|
| 95 |
-
| Six-hour direction error ↓ | 51.81° | 35.33° |
|
| 96 |
-
| Centred route-shape similarity ↑ | 0.7497 | 0.8842 |
|
| 97 |
-
| Geographic path similarity ↑ | 0.5310 | 0.6611 |
|
| 98 |
-
|
| 99 |
-
**Incomplete development evidence:** only the 159 fully completed storms enter this table. Execution order is not random, and the remaining storms can change the result. These are not the final 270-storm scores or an untouched-test claim. Do not compare this subset directly with the differently sampled 270-case table above.
|
| 100 |
-
|
| 101 |
-
1.2-only pressure diagnostics: central-pressure MAE **12.62 hPa** over 40 storms with valid pressure labels; basin MSLP MAE **2.70 hPa** over 159 storms. There is no matched 1.1 pressure result in this run, and basin-wide error does not establish core-pressure-map accuracy.
|
| 102 |
|
| 103 |
## Fung-wong pressure forecast — MP4
|
| 104 |
|
|
|
|
| 11 |
|
| 12 |
**[Download 1.2: weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Read the illustrated paper](paper/trackformer.pdf)** · **[Watch Fung-wong](https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4)** · **[Hugging Face model](https://huggingface.co/euler314/typhoon-predict)**
|
| 13 |
|
| 14 |
+
Release status: **1.2 research release**, not operational certification. The package contains the released 1.2 weights, inference code, input contract, completed 270-storm evaluation and illustrated paper. The original 1.1 release remains separate.
|
| 15 |
|
| 16 |
Trackformer 1.2 is a **research-only Western Pacific tropical-cyclone forecast model**. It evolves a sea-level-pressure (MSLP) field and a moving storm-centred pressure core every six hours through +120 h. A track and central-pressure estimate are extracted from the evolving core, not independently drawn on top of a pressure image. The prior [Trackformer 1.1 release](https://github.com/yu314-coder/typhoon-predict/releases/tag/trackformer-1.1) remains available and unchanged.
|
| 17 |
|
|
|
|
| 54 |
|
| 55 |
Grid spacing describes the representation, not independently demonstrated effective resolution. A mean of member centres is not necessarily the minimum of the displayed mean pressure field.
|
| 56 |
|
| 57 |
+
## Completed benchmark: 270 storms · 1,473 daily forecasts
|
| 58 |
+
|
| 59 |
+
One typhoon on one UTC day is one case. Each issue forecasts **+6 to +120 h**. We average lead errors within each issue, daily scores within each storm, then the **270 storm scores equally**. This prevents long-lived storms dominating the benchmark. All 1,473 cases finished on **29 September 2026, 08:22 UTC**, with every saved forecast SHA-256 verified.
|
| 60 |
+
|
| 61 |
+
| Equal-storm metric | 1.1 | 1.2 · mean of 50 | Preferred |
|
| 62 |
+
| --- | ---: | ---: | --- |
|
| 63 |
+
| Mean track error, +6 to +120 h | 798.4 km | **471.2 km** | Lower |
|
| 64 |
+
| Track error at +120 h | 1,646.4 km | **1,031.6 km** | Lower |
|
| 65 |
+
| Six-hour direction error | 51.58° | **34.96°** | Lower |
|
| 66 |
+
| Centred route-shape similarity | 0.7544 | **0.8837** | Higher |
|
| 67 |
+
| Geographic path similarity | 0.5345 | **0.6560** | Higher |
|
| 68 |
+
| Fréchet distance | 1,658.1 km | **1,045.8 km** | Lower |
|
| 69 |
+
|
| 70 |
+

|
| 71 |
+
|
| 72 |
+
Mean track error is **41.0% lower** for 1.2 on this cohort. Direction, shape and geographic alignment remain separate measures, not a combined score. A shape score near one does not guarantee overlapping routes at matching times. The 1.2 forecasts are means of **50 distinct causal input perturbations**, not 50 independently trained networks. Forecast pipelines differ, so this is not an architecture ablation.
|
| 73 |
+
|
| 74 |
+
**Pressure coverage:** 1.2 central-pressure MAE is **12.62 hPa over 40 storms with valid labels**; basin-area-weighted MSLP MAE is **2.72 hPa over 270 storms**. There is no matched 1.1 pressure output in this run. Basin-wide error does not establish core-field accuracy.
|
| 75 |
+
|
| 76 |
+
**Limits:** the frozen cohort includes 40 recent storms and 230 historical storms from 1980–1999. It excludes this checkpoint's fitting/validation years, but has not been certified untouched across prior experiments. Historical hindcasts use retrospective analyses and a model trained on later years. Complete five-day labels are required, excluding short remaining lifetimes. No operational or no-overfitting claim follows.
|
| 77 |
+
|
| 78 |
+
[Final report and all storm scores](evaluation/daily_storm_final.json) · [Protocol](docs/daily_storm_benchmark.md) · [Frozen issues](evaluation/release_data/daily_storm_cohort.json) · [Reproduce chart](release_tools/plot_daily_storm_final.py)
|
| 79 |
+
|
| 80 |
+
## Earlier 270-case comparison: 90 storms
|
| 81 |
|
| 82 |
This comparison uses **270 forecast cases from 90 storms**, with the same issue rows, observed routes and 20 six-hour leads through +120 h. The number 270 counts forecast cases. The 1.2 prediction for each case is a **mean of 50 causal input-perturbation members**; 1.1 uses its archived causal route pipeline.
|
| 83 |
|
|
|
|
| 105 |
|
| 106 |
The 1.2 mean has better aggregate direction, shape and position scores on this previously inspected development cohort. That does not establish an overall win on every storm or pressure metric. A same-270 1.1 central-pressure prediction array was not verified in these route artifacts, so it is not assigned a pressure-error bar. The separate 1.2 report gives 15.73 hPa central-pressure MAE and 2.54 hPa area-weighted regional MSLP MAE. TIP remains a separate diagnostic: its pressure error was worse for 1.2 (30.7 vs 20.1 hPa). See [metric definitions and limitations](docs/trackformer_1_2_evaluation.md).
|
| 107 |
|
| 108 |
+
The earlier dated partial snapshot is retained for provenance only. The completed 270-storm result at the top of this document supersedes it; the earlier 270-case comparison remains a different cohort.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 109 |
|
| 110 |
## Fung-wong pressure forecast — MP4
|
| 111 |
|
RELEASE_NOTES_TRACKFORMER_1_2.md
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
# Trackformer 1.2 — research
|
| 2 |
|
| 3 |
**[Download weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Download the illustrated PDF](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer.pdf)** · **[Hugging Face model card](https://huggingface.co/euler314/typhoon-predict)**
|
| 4 |
|
|
@@ -8,7 +8,7 @@ The attached archive includes inference-only weights, exact source modules, an i
|
|
| 8 |
|
| 9 |
On the 270-case development cohort, 1.1 versus 1.2 mean-of-50 scores are: direction error **56.22° vs 44.82°**, route-shape similarity **0.7273 vs 0.8249**, path similarity **0.4983 vs 0.5727**, and mean track error **902.3 vs 714.4 km**. Direction uses 5,382 common valid steps. The old 1.1 report had stored latitude/longitude under a local-kilometre key; the new comparison repairs that coordinate interpretation and records the correction. Original artifacts remain available. These results are development comparisons, and no matched 1.1 pressure-field bar is claimed.
|
| 10 |
|
| 11 |
-
The route gallery shows six explicitly selected best-performing examples from distinct storms, with the selection rule and all case scores supplied. It is not representative evidence; the aggregate bars retain all 270 cases.
|
| 12 |
|
| 13 |
The README shows the user-selected **Fung-wong MP4** in place of Yagi; the interactive PRAPIROON showcase remains removed. The forecast starts 7 November 2025 at 00 UTC and uses actual saved 50-member mean pressure fields and routes through +120 h. Central-pressure MAE is 7.36 hPa against JMA best track from IBTrACS. The broad route is similar but not perfectly overlapping: mean geographic error is 130.9 km and +120 h error is 142.3 km. This is a selected development example, not typical skill. From +66 h the forecast centre leaves the regional patch; the video explicitly labels that only the saved coarse basin field is shown there. No detailed core is invented outside coverage.
|
| 14 |
|
|
@@ -18,6 +18,18 @@ The Surigae illustration is a single forecast issued 2026-09-27 12 UTC, with 4 h
|
|
| 18 |
|
| 19 |
This is not an operational warning service. Do not use for safety-critical decisions. A genuinely untouched storm-level holdout is still required before generalization claims.
|
| 20 |
|
| 21 |
-
##
|
| 22 |
|
| 23 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Trackformer 1.2 — research release
|
| 2 |
|
| 3 |
**[Download weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Download the illustrated PDF](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer.pdf)** · **[Hugging Face model card](https://huggingface.co/euler314/typhoon-predict)**
|
| 4 |
|
|
|
|
| 8 |
|
| 9 |
On the 270-case development cohort, 1.1 versus 1.2 mean-of-50 scores are: direction error **56.22° vs 44.82°**, route-shape similarity **0.7273 vs 0.8249**, path similarity **0.4983 vs 0.5727**, and mean track error **902.3 vs 714.4 km**. Direction uses 5,382 common valid steps. The old 1.1 report had stored latitude/longitude under a local-kilometre key; the new comparison repairs that coordinate interpretation and records the correction. Original artifacts remain available. These results are development comparisons, and no matched 1.1 pressure-field bar is claimed.
|
| 10 |
|
| 11 |
+
The route gallery shows six explicitly selected best-performing examples from distinct storms, with the selection rule and all case scores supplied. It is not representative evidence; the earlier aggregate bars retain all 270 cases. The new primary benchmark below uses **270 distinct storms / 1,473 daily issues**, with daily scores averaged per storm and storms weighted equally. The paper now presents this completed larger cohort, separately from the earlier 270-case results.
|
| 12 |
|
| 13 |
The README shows the user-selected **Fung-wong MP4** in place of Yagi; the interactive PRAPIROON showcase remains removed. The forecast starts 7 November 2025 at 00 UTC and uses actual saved 50-member mean pressure fields and routes through +120 h. Central-pressure MAE is 7.36 hPa against JMA best track from IBTrACS. The broad route is similar but not perfectly overlapping: mean geographic error is 130.9 km and +120 h error is 142.3 km. This is a selected development example, not typical skill. From +66 h the forecast centre leaves the regional patch; the video explicitly labels that only the saved coarse basin field is shown there. No detailed core is invented outside coverage.
|
| 14 |
|
|
|
|
| 18 |
|
| 19 |
This is not an operational warning service. Do not use for safety-critical decisions. A genuinely untouched storm-level holdout is still required before generalization claims.
|
| 20 |
|
| 21 |
+
## Completed daily benchmark and revised paper
|
| 22 |
|
| 23 |
+
All **1,473 daily forecasts across 270 distinct storms** completed at 2026-09-29T08:22:39Z. Every saved forecast SHA-256 was verified. Each storm has equal weight after its daily issues are averaged; all leads +6 through +120 h are included.
|
| 24 |
+
|
| 25 |
+
| Equal-storm measure | 1.1 | 1.2 mean of 50 |
|
| 26 |
+
| --- | ---: | ---: |
|
| 27 |
+
| Mean track error | 798.4 km | 471.2 km |
|
| 28 |
+
| +120 h track error | 1,646.4 km | 1,031.6 km |
|
| 29 |
+
| Direction error | 51.58° | 34.96° |
|
| 30 |
+
| Centred shape similarity | 0.7544 | 0.8837 |
|
| 31 |
+
| Geographic path similarity | 0.5345 | 0.6560 |
|
| 32 |
+
|
| 33 |
+
This is a **41.0% reduction in mean track error** on a broader development cohort, not a certified untouched test. The 1.2-only central-pressure MAE is 12.62 hPa over 40 pressure-labelled storms; basin MSLP MAE is 2.72 hPa over 270 storms. There is no matched 1.1 pressure comparison. Retrospective analyses, different input pipelines, prior model selection and complete-five-day eligibility limit interpretation.
|
| 34 |
+
|
| 35 |
+
The main README chart and redesigned eight-page paper now use these final equal-storm results. The earlier 270-case/90-storm comparison remains explicitly separate. `evaluation/daily_storm_final.json` includes all storm scores; `release_tools/plot_daily_storm_final.py` reproduces the new bars. The paper includes the architecture, equations, lead-error chart, selected route/pressure curves and model isobars. **Weights and inference implementation are unchanged.**
|
docs/daily_storm_benchmark.md
CHANGED
|
@@ -1,6 +1,8 @@
|
|
| 1 |
# Daily forecasts, equal-storm benchmark
|
| 2 |
|
| 3 |
-
Status:
|
|
|
|
|
|
|
| 4 |
|
| 5 |
## What counts as a case and a score
|
| 6 |
|
|
|
|
| 1 |
# Daily forecasts, equal-storm benchmark
|
| 2 |
|
| 3 |
+
Status: **complete** at 2026-09-29T08:22:39Z. All **1,473 daily cases / 270 storms** finished; the SHA-256 of every saved forecast was verified. [Final equal-storm results and all 270 storm scores](../evaluation/daily_storm_final.json).
|
| 4 |
+
|
| 5 |
+
Mean track error is 798.4 km for 1.1 and 471.2 km for 1.2 (50-member mean); +120 h error is 1,646.4 versus 1,031.6 km. Direction error is 51.58° versus 34.96°, centred shape similarity 0.7544 versus 0.8837, and geographic path similarity 0.5345 versus 0.6560. Central-pressure MAE for 1.2 is 12.62 hPa on 40 pressure-labelled storms; basin MSLP MAE is 2.72 hPa on 270 storms. No matched 1.1 pressure result or confidence interval is claimed.
|
| 6 |
|
| 7 |
## What counts as a case and a score
|
| 8 |
|
evaluation/daily_storm_final.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png
ADDED
|
Git LFS Details
|
paper/trackformer.pdf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d5797fc86961d7163bba57903d0bed3da47f29f585aa567f24b75a6403c4a8c2
|
| 3 |
+
size 416164
|
paper/trackformer.tex
CHANGED
|
@@ -1,65 +1,113 @@
|
|
| 1 |
-
\documentclass[10pt]{article}
|
| 2 |
-
\usepackage[
|
| 3 |
-
\usepackage{
|
| 4 |
-
\usepackage{
|
|
|
|
| 5 |
\usetikzlibrary{arrows.meta,positioning,calc}
|
| 6 |
-
\definecolor{modelteal}{HTML}{
|
| 7 |
\definecolor{baselinegray}{HTML}{65758B}
|
| 8 |
\definecolor{ink}{HTML}{182C42}
|
| 9 |
\definecolor{routepink}{HTML}{B31565}
|
| 10 |
-
\
|
|
|
|
| 11 |
\pagestyle{fancy}
|
| 12 |
\fancyhf{}
|
| 13 |
-
\fancyhead[L]{\
|
| 14 |
-
\fancyhead[R]{\
|
| 15 |
-
\fancyfoot[
|
|
|
|
|
|
|
| 16 |
\setlength{\headheight}{14pt}
|
| 17 |
-
\
|
| 18 |
-
\hypersetup{pdftitle={Trackformer 1.2: Coupled Pressure-Field Evolution for Tropical-Cyclone Forecasting},pdfauthor={Yu Yao-Hsing},pdfsubject={Architecture and development evaluation}}
|
| 19 |
-
\setlength{\parskip}{4pt}
|
| 20 |
\setlength{\parindent}{0pt}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
\newcommand{\sample}{\mathcal{S}}
|
| 22 |
-
\
|
| 23 |
-
\
|
| 24 |
-
\date{Research technical report --- 29 September 2026}
|
| 25 |
\begin{document}
|
| 26 |
-
\
|
| 27 |
-
\
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 51 |
|
| 52 |
At step $k$, the recurrent state is
|
| 53 |
\[
|
| 54 |
s_k=(g_k,h^g_k,c_k,h^c_k,a_k,o_k,z_k,m_k,v_k;G,R,I,M).
|
| 55 |
\]
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
\
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
\centering
|
| 64 |
\begin{tikzpicture}[font=\small,>=Latex,
|
| 65 |
box/.style={draw=ink!50,rounded corners=3pt,align=center,text width=4.6cm,minimum height=1.15cm,inner sep=6pt,fill=ink!3},
|
|
@@ -87,47 +135,36 @@ basin + regional pressure fields, coverage masks and auxiliary wind};
|
|
| 87 |
\node[below=0.25cm of readout,align=center,font=\scriptsize,text=ink] {Shared transition repeated 20 times: +6, +12, $\ldots$, +120 h.\\
|
| 88 |
50 perturbed input histories $\rightarrow$ 50 rollouts $\rightarrow$ separate route, pressure and common-grid field means.};
|
| 89 |
\end{tikzpicture}
|
| 90 |
-
\caption{
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
are excluded from inference. The ensemble contains 50 input realizations of one
|
| 94 |
-
network, not 50 trained networks.}
|
| 95 |
\label{fig:architecture}
|
| 96 |
\end{figure}
|
| 97 |
|
| 98 |
-
\
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
\
|
| 103 |
-
|
| 104 |
-
\
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
The internal core is a $65\times65$ grid at 20-km sampling over $\pm640$ km.
|
| 123 |
-
Its field is a learned reconstruction, not native 20-km observations.
|
| 124 |
-
For each of twenty leads ($6,12,\ldots,120$ h), outputs include basin fields,
|
| 125 |
-
regional pressure, core pressure/coordinates, storm centre, central pressure,
|
| 126 |
-
domain masks and auxiliary maximum wind. This implementation does not output
|
| 127 |
-
quadrant wind radii or a resolved surface-wind map.
|
| 128 |
-
|
| 129 |
-
\section{Neural modules}
|
| 130 |
-
The released non-tiny architecture has 21,452,595 trainable parameters.
|
| 131 |
\subsection{Multiresolution evolution networks}
|
| 132 |
The basin network receives 44 channels: eight fields, 32 memory features and
|
| 133 |
four static channels. Encoder widths are $72,144,288,432$; residual-block depths
|
|
@@ -169,7 +206,7 @@ retrieves context; a zero-initialized $64\to32$ projection adds it to basin
|
|
| 169 |
memory before the transition. No separate positional embedding is added:
|
| 170 |
geographic coordinates are already among the token inputs.
|
| 171 |
|
| 172 |
-
\
|
| 173 |
Write $\sample(F,\xi)$ for bilinear field sampling. Sampling uses FP32 geometry,
|
| 174 |
aligned corners and border padding for the basin; transported core anomalies
|
| 175 |
use zero padding. All nine pressure histories are sampled onto the core grid
|
|
@@ -188,77 +225,73 @@ $T=\operatorname{clip}((640-\sqrt{x^2+y^2})/160,0,1)$.
|
|
| 188 |
The anomaly is $a_0=c_0-\sample(g_{0,\mathrm{MSLP}},\xi_0)$.
|
| 189 |
This compact prior is applied only at initialization. Future scalar pressure
|
| 190 |
is not repeatedly inserted into the field.
|
| 191 |
-
|
| 192 |
-
\section{One six-hour transition}
|
| 193 |
-
|
| 194 |
-
|
| 195 |
-
|
| 196 |
-
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
|
| 200 |
-
|
| 201 |
-
|
| 202 |
-
\
|
| 203 |
-
|
| 204 |
-
\
|
| 205 |
-
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
| 210 |
-
|
| 211 |
|
| 212 |
\subsection{Moving-frame proposal}
|
| 213 |
-
A $44\to64\to3$ GELU MLP reads
|
| 214 |
previous motion divided by 100, and two issue-intensity masks. For output $b$,
|
| 215 |
-
\
|
| 216 |
m_{k+1}=0.7(21.6)U_k(z_k)+0.3m_k+100\tanh(b_{0:2}).
|
| 217 |
-
\
|
| 218 |
-
This
|
| 219 |
-
$z_k$
|
| 220 |
-
centre. The third output updates auxiliary wind:
|
| 221 |
$v_{k+1}=\max(0,v_k+5b_2)$.
|
| 222 |
|
| 223 |
-
\subsection{
|
| 224 |
-
The core network reads
|
| 225 |
-
|
| 226 |
$U^{\rm local}=\sample(U_k)+5\tanh(d^c_{1:3})$.
|
| 227 |
-
Subtracting $m_{k+1}/21.6$ removes bulk frame
|
| 228 |
-
also
|
| 229 |
-
|
| 230 |
-
\begin{align*}
|
| 231 |
a_{k+1}&=\sample(a_k,D_c)+0.1d^c_0,\\
|
| 232 |
c_{k+1}&=\sample(g_{k+1,\mathrm{MSLP}},\xi_{k+1})+Ta_{k+1}.
|
| 233 |
-
\end{align
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
|
| 237 |
-
\subsection{
|
| 238 |
-
Convert core pressure to hPa
|
| 239 |
-
|
| 240 |
-
\
|
| 241 |
-
w_i=\operatorname{softmax}_i\left(-P_i/2-r_i^2/(2\,180^2)\right).
|
| 242 |
-
\
|
| 243 |
-
|
| 244 |
-
is $
|
| 245 |
-
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
|
| 249 |
-
|
| 250 |
-
\[
|
| 251 |
\widehat p_R=\sample(g_{k+1,\mathrm{MSLP}},R)
|
| 252 |
-
|
| 253 |
-
\
|
| 254 |
-
then de-normalized to hPa. This basin-plus-core
|
| 255 |
-
|
| 256 |
-
|
| 257 |
-
|
| 258 |
-
|
| 259 |
-
|
| 260 |
-
|
| 261 |
-
\section{Autoregression and learning objectives}
|
| 262 |
Inference initializes once and repeats the transition twenty times, using only
|
| 263 |
fixed issue context and its own predicted future state. The implemented
|
| 264 |
single-step loss is
|
|
@@ -281,7 +314,7 @@ together, without adding an independent central-pressure output head.
|
|
| 281 |
Training-only normalization, storm-level split boundaries and causal input
|
| 282 |
selection remain essential implementation constraints.
|
| 283 |
|
| 284 |
-
\
|
| 285 |
Evaluation-mode inference is deterministic for fixed inputs and weights.
|
| 286 |
The documented ensemble perturbs only normalized historical weather: independent
|
| 287 |
Gaussian fields are spatially smoothed with kernels five (basin) and nine
|
|
@@ -301,38 +334,37 @@ These are perturbed-input realizations of one model, not 50 trained networks or
|
|
| 301 |
learned latent samples. Spread is not automatically calibrated uncertainty.
|
| 302 |
Member counts, perturbation settings and seed policy must accompany the outputs.
|
| 303 |
|
| 304 |
-
\
|
| 305 |
-
The public
|
| 306 |
-
checkpoint identity
|
| 307 |
-
|
| 308 |
-
\
|
| 309 |
-
\
|
| 310 |
-
|
| 311 |
-
|
| 312 |
-
|
| 313 |
-
|
| 314 |
-
|
| 315 |
-
|
| 316 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 317 |
\section{Development evaluation}
|
| 318 |
\label{sec:evaluation}
|
| 319 |
-
\subsection{
|
| 320 |
-
The completed comparison
|
| 321 |
-
with
|
| 322 |
-
|
| 323 |
-
|
| 324 |
-
|
| 325 |
-
|
| 326 |
-
|
| 327 |
-
or basin-specific validation claim.
|
| 328 |
-
|
| 329 |
-
The original 1.1 archive stored absolute latitude/longitude under a local-km
|
| 330 |
-
array name. The reproduced comparison converts those coordinates into the same
|
| 331 |
-
issue-centred local projection used by 1.2 and checks a round trip to the
|
| 332 |
-
original coordinates. Its corrected 902.3-km mean replaces the superseded
|
| 333 |
-
1,036.3-km score; original forecasts are preserved.
|
| 334 |
-
|
| 335 |
-
\begin{figure}[htbp]
|
| 336 |
\centering
|
| 337 |
\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
|
| 338 |
\begin{scope}[shift={(0,0)}]
|
|
@@ -340,15 +372,15 @@ original coordinates. Its corrected 902.3-km mean replaces the superseded
|
|
| 340 |
\node[rotate=90] at (-.75,1.75) {km; lower is better};
|
| 341 |
\draw[black!10] (0,0.0000)--(5,0.0000);
|
| 342 |
\node[left] at (0,0.0000) {0};
|
| 343 |
-
\draw[black!10] (0,1.
|
| 344 |
-
\node[left] at (0,1.
|
| 345 |
-
\draw[black!10] (0,3.
|
| 346 |
-
\node[left] at (0,3.
|
| 347 |
-
\fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.
|
| 348 |
-
\node[above] at (1.4,2.
|
| 349 |
\node[below=3pt] at (1.4,0) {1.1};
|
| 350 |
-
\fill[modelteal] (3.22,0) rectangle (3.98,
|
| 351 |
-
\node[above] at (3.6,
|
| 352 |
\node[below=3pt] at (3.6,0) {1.2 mean of 50};
|
| 353 |
\draw[black!55] (0,3.5)--(0,0)--(5,0);
|
| 354 |
\end{scope}
|
|
@@ -363,11 +395,11 @@ original coordinates. Its corrected 902.3-km mean replaces the superseded
|
|
| 363 |
\node[left] at (0,2.0000) {40};
|
| 364 |
\draw[black!10] (0,3.0000)--(5,3.0000);
|
| 365 |
\node[left] at (0,3.0000) {60};
|
| 366 |
-
\fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.
|
| 367 |
-
\node[above] at (1.4,2.
|
| 368 |
\node[below=3pt] at (1.4,0) {1.1};
|
| 369 |
-
\fill[modelteal] (3.22,0) rectangle (3.98,
|
| 370 |
-
\node[above] at (3.6,
|
| 371 |
\node[below=3pt] at (3.6,0) {1.2 mean of 50};
|
| 372 |
\draw[black!55] (0,3.5)--(0,0)--(5,0);
|
| 373 |
\end{scope}
|
|
@@ -401,49 +433,45 @@ original coordinates. Its corrected 902.3-km mean replaces the superseded
|
|
| 401 |
\node[font=\small,rotate=90] at (-1.05,1.850) {Mean position error (km)};
|
| 402 |
\begin{scope}
|
| 403 |
\clip (0,0) rectangle (12.2,3.7);
|
| 404 |
-
\draw[baselinegray,dashed,thick] (0.0000,0.
|
| 405 |
-
\draw[modelteal,solid,thick] (0.0000,0.
|
| 406 |
\end{scope}
|
| 407 |
\draw[baselinegray,dashed,thick] (0.250,-1.23)--(0.800,-1.23);
|
| 408 |
\node[anchor=west] at (0.900,-1.23) {Trackformer 1.1};
|
| 409 |
\draw[modelteal,solid,thick] (6.350,-1.23)--(6.900,-1.23);
|
| 410 |
\node[anchor=west] at (7.000,-1.23) {Trackformer 1.2: mean of 50};
|
| 411 |
\end{tikzpicture}
|
| 412 |
-
\caption{
|
| 413 |
-
|
| 414 |
-
|
| 415 |
-
|
| 416 |
-
scores, including poor and out-of-domain cases. No uncertainty intervals are
|
| 417 |
-
available in these saved results.}
|
| 418 |
\label{fig:benchmark}
|
| 419 |
\end{figure}
|
| 420 |
-
|
| 421 |
-
\
|
| 422 |
-
|
| 423 |
-
|
| 424 |
-
\
|
| 425 |
-
|
| 426 |
-
|
| 427 |
-
|
| 428 |
-
|
| 429 |
-
|
| 430 |
-
|
| 431 |
-
|
| 432 |
-
|
| 433 |
-
|
| 434 |
-
|
| 435 |
-
|
| 436 |
-
|
| 437 |
-
\
|
| 438 |
-
\section{
|
| 439 |
-
The
|
| 440 |
-
|
| 441 |
-
|
| 442 |
-
|
| 443 |
-
|
| 444 |
-
|
| 445 |
-
|
| 446 |
-
\begin{figure}[htbp]
|
| 447 |
\centering
|
| 448 |
\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
|
| 449 |
\node[font=\small\bfseries,align=center] at (6.100,7.800) {Fung-wong: geographic routes, +0 to +120 h};
|
|
@@ -569,10 +597,15 @@ The companion MP4 shows the actual common-grid model pressure fields and
|
|
| 569 |
4-hPa isobars; none of the pressure fields is moved to match observations.}
|
| 570 |
\label{fig:fungwong}
|
| 571 |
\end{figure}
|
| 572 |
-
|
| 573 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 574 |
\centering
|
| 575 |
-
\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
|
| 576 |
\node[font=\small\bfseries,align=center] at (5.750,8.900) {Fung-wong: model-mean MSLP at +48 h};
|
| 577 |
\draw[black!10] (0.676,0)--(0.676,8.5);
|
| 578 |
\node[below=3pt] at (0.676,0) {124};
|
|
@@ -639,59 +672,36 @@ The companion MP4 shows the actual common-grid model pressure fields and
|
|
| 639 |
\draw[routepink,solid,thick] (6.000,-1.23)--(6.550,-1.23);
|
| 640 |
\node[anchor=west] at (6.650,-1.23) {Forecast route to +48 h};
|
| 641 |
\end{tikzpicture}
|
| 642 |
-
\caption{
|
| 643 |
-
(9 November 2025, 00 UTC), contoured every 4 hPa and labelled every 8 hPa.
|
| 644 |
-
The source grid is sampled at $0.25^\circ$; this is not a claim of native
|
| 645 |
-
resolution. Lines overlay the observed and forecast paths only through this
|
| 646 |
-
valid time. The figure uses the portion inside the fixed regional patch;
|
| 647 |
-
its western edge is $123^\circ$E. The map has not been shifted to match the
|
| 648 |
-
observed centre, and its minimum need not equal member-mean central pressure.}
|
| 649 |
\label{fig:pressuremap}
|
| 650 |
\end{figure}
|
| 651 |
-
|
| 652 |
-
\
|
| 653 |
-
|
| 654 |
-
|
| 655 |
-
|
| 656 |
-
the
|
| 657 |
-
|
| 658 |
-
|
| 659 |
-
|
| 660 |
-
|
| 661 |
-
|
| 662 |
-
|
| 663 |
-
|
| 664 |
-
|
| 665 |
-
|
| 666 |
-
|
| 667 |
-
|
| 668 |
-
|
| 669 |
-
|
| 670 |
-
|
| 671 |
-
|
| 672 |
-
|
| 673 |
-
|
| 674 |
-
|
| 675 |
-
|
| 676 |
-
|
| 677 |
-
|
| 678 |
-
|
| 679 |
-
|
| 680 |
-
is an illustration of saved predictions, not a new inference experiment.
|
| 681 |
-
|
| 682 |
-
A separate expanded protocol freezes 1,473 daily issues from 270 distinct
|
| 683 |
-
storms: average leads within each issue, issues within each storm, then storms
|
| 684 |
-
equally. It is not the completed 270-case comparison reported here; its
|
| 685 |
-
partial results are not used in these figures. Stronger generalization claims
|
| 686 |
-
require a genuinely unused whole-storm holdout, fitting-only preprocessing,
|
| 687 |
-
matched data availability and lead masks, and uncertainty estimates based on
|
| 688 |
-
resampling whole storms rather than overlapping issues.
|
| 689 |
-
|
| 690 |
-
\section{Conclusion}
|
| 691 |
-
Trackformer 1.2 couples recurrent environmental evolution with a translating
|
| 692 |
-
pressure anomaly and pressure-associated readout. Development results improve
|
| 693 |
-
aggregate route measures relative to the archived 1.1 pipeline, while leaving
|
| 694 |
-
large absolute errors, short-lead regressions and intensity failures. The
|
| 695 |
-
architecture is a research model, not a substitute for official warnings;
|
| 696 |
-
selected visually aligned routes do not establish operational reliability.
|
| 697 |
\end{document}
|
|
|
|
| 1 |
+
\documentclass[10pt,letterpaper]{article}
|
| 2 |
+
\usepackage[left=0.82in,right=0.82in,top=0.78in,bottom=0.78in,headsep=0.22in]{geometry}
|
| 3 |
+
\usepackage[T1]{fontenc}
|
| 4 |
+
\usepackage{lmodern,amsmath,amssymb,booktabs,array,tabularx,microtype,xcolor}
|
| 5 |
+
\usepackage{tikz,caption,fancyhdr,float}
|
| 6 |
\usetikzlibrary{arrows.meta,positioning,calc}
|
| 7 |
+
\definecolor{modelteal}{HTML}{007D78}
|
| 8 |
\definecolor{baselinegray}{HTML}{65758B}
|
| 9 |
\definecolor{ink}{HTML}{182C42}
|
| 10 |
\definecolor{routepink}{HTML}{B31565}
|
| 11 |
+
\usepackage[colorlinks=true,linkcolor=modelteal,urlcolor=modelteal]{hyperref}
|
| 12 |
+
\hypersetup{pdftitle={Trackformer 1.2: Coupled Pressure-Field Forecasting},pdfauthor={Yu Yao-Hsing},pdfsubject={Architecture, forecast dynamics and development evaluation}}
|
| 13 |
\pagestyle{fancy}
|
| 14 |
\fancyhf{}
|
| 15 |
+
\fancyhead[L]{\sffamily\footnotesize\color{ink}TRACKFORMER / 1.2}
|
| 16 |
+
\fancyhead[R]{\sffamily\footnotesize\color{baselinegray}RESEARCH TECHNICAL REPORT}
|
| 17 |
+
\fancyfoot[L]{\sffamily\scriptsize\color{baselinegray}Architecture and development evaluation}
|
| 18 |
+
\fancyfoot[R]{\sffamily\small\thepage}
|
| 19 |
+
\renewcommand{\headrulewidth}{0.3pt}
|
| 20 |
\setlength{\headheight}{14pt}
|
| 21 |
+
\captionsetup{font={small},labelfont={bf,sf,color=ink},labelsep=period,justification=justified,singlelinecheck=false,skip=7pt}
|
|
|
|
|
|
|
| 22 |
\setlength{\parindent}{0pt}
|
| 23 |
+
\setlength{\parskip}{4pt plus 1pt}
|
| 24 |
+
\setlength{\intextsep}{10pt}
|
| 25 |
+
\setlength{\abovedisplayskip}{7pt}
|
| 26 |
+
\setlength{\belowdisplayskip}{7pt}
|
| 27 |
+
\setlength{\emergencystretch}{1.5em}
|
| 28 |
+
\renewcommand{\arraystretch}{1.18}
|
| 29 |
+
\widowpenalty=10000
|
| 30 |
+
\clubpenalty=10000
|
| 31 |
+
\raggedbottom
|
| 32 |
+
\makeatletter
|
| 33 |
+
\renewcommand\section{\@startsection{section}{1}{0pt}{12pt}{7pt}{\sffamily\bfseries\color{ink}\fontsize{14}{17}\selectfont}}
|
| 34 |
+
\renewcommand\subsection{\@startsection{subsection}{2}{0pt}{9pt}{4pt}{\sffamily\bfseries\color{ink}\fontsize{10.5}{13}\selectfont}}
|
| 35 |
+
\makeatother
|
| 36 |
\newcommand{\sample}{\mathcal{S}}
|
| 37 |
+
\newcommand{\reportpage}{\clearpage}
|
| 38 |
+
\newcommand{\keyline}[2]{\textbf{#1}\quad #2\par}
|
|
|
|
| 39 |
\begin{document}
|
| 40 |
+
\fontsize{10.3}{13}\selectfont
|
| 41 |
+
\thispagestyle{empty}
|
| 42 |
+
{\sffamily\small\color{modelteal}RESEARCH TECHNICAL REPORT\hfill 29 SEPTEMBER 2026\par}
|
| 43 |
+
\vspace{13pt}
|
| 44 |
+
{\sffamily\bfseries\color{ink}\fontsize{26}{29}\selectfont Trackformer 1.2\par}
|
| 45 |
+
\vspace{4pt}
|
| 46 |
+
{\sffamily\color{ink}\fontsize{17}{21}\selectfont Coupled pressure-field forecasting\par}
|
| 47 |
+
\vspace{6pt}
|
| 48 |
+
{\large Architecture, forecast dynamics and development evaluation\par}
|
| 49 |
+
\vspace{10pt}
|
| 50 |
+
{\sffamily\small Yu Yao-Hsing\hfill\href{https://github.com/yu314-coder/typhoon-predict}{Code, data and model release}\par}
|
| 51 |
+
\vspace{8pt}
|
| 52 |
+
{\color{modelteal}\hrule height 1pt}
|
| 53 |
+
\vspace{10pt}
|
| 54 |
+
{\sffamily\bfseries\color{ink}Abstract\par}
|
| 55 |
+
Trackformer 1.2 forecasts tropical-cyclone motion and central pressure by
|
| 56 |
+
jointly advancing a basin-scale atmospheric state and a translating pressure
|
| 57 |
+
core. Recurrent convolutional networks evolve the fields, while multiscale
|
| 58 |
+
attention supplies environmental context. A pressure-associated centre and
|
| 59 |
+
central-pressure sample are read from the evolving core at each six-hour
|
| 60 |
+
step. Twenty autoregressive transitions produce a five-day forecast.
|
| 61 |
+
This report specifies the released architecture, state updates, learning
|
| 62 |
+
objectives and 50-member input-perturbation policy. Across 1,473 daily forecasts
|
| 63 |
+
from 270 storms, equal-storm mean track error is 471.2 km for the 1.2 ensemble
|
| 64 |
+
and 798.4 km for the causal 1.1 route pipeline.
|
| 65 |
+
The pipelines differ; this is not an architecture ablation or an untouched-test
|
| 66 |
+
result. A selected Fung-wong case illustrates actual geographic alignment,
|
| 67 |
+
intensity evolution and model-generated pressure contours, including their
|
| 68 |
+
remaining errors and coverage limits.
|
| 69 |
+
|
| 70 |
+
\vspace{5pt}
|
| 71 |
+
\begin{tabularx}{\linewidth}{@{}XXX@{}}
|
| 72 |
+
\toprule
|
| 73 |
+
\sffamily\bfseries Forecast horizon & \sffamily\bfseries Representation & \sffamily\bfseries Evaluation policy\\
|
| 74 |
+
+6 to +120 h & Basin + moving pressure core & 50 input realizations\\
|
| 75 |
+
20 six-hour transitions & 21.45 million parameters & One shared set of weights\\
|
| 76 |
+
\bottomrule
|
| 77 |
+
\end{tabularx}
|
| 78 |
+
|
| 79 |
+
\section{Overview and forecast state}
|
| 80 |
+
The defining constraint is a coupled readout: the route and central pressure
|
| 81 |
+
must follow the pressure core generated by the model. A learned,
|
| 82 |
+
steering-conditioned proposal moves the core's reference frame; a local
|
| 83 |
+
pressure association then determines the reported centre. The method is
|
| 84 |
+
therefore neither a pressure-only tracker nor a complete numerical atmospheric
|
| 85 |
+
solver. Its displayed pressure is a forecast field, not an independently drawn
|
| 86 |
+
vortex placed along a predicted route.
|
| 87 |
|
| 88 |
At step $k$, the recurrent state is
|
| 89 |
\[
|
| 90 |
s_k=(g_k,h^g_k,c_k,h^c_k,a_k,o_k,z_k,m_k,v_k;G,R,I,M).
|
| 91 |
\]
|
| 92 |
+
\begin{tabularx}{\linewidth}{@{}lX@{}}
|
| 93 |
+
\toprule
|
| 94 |
+
\textbf{State} & \textbf{Meaning}\\
|
| 95 |
+
$g,h^g$ & Eight normalized basin fields and basin memory.\\
|
| 96 |
+
$c,h^c,a$ & Core pressure, core memory and pressure anomaly relative to the basin.\\
|
| 97 |
+
$o,z,m,v$ & Core-grid origin, associated centre, six-hour motion and auxiliary wind.\\
|
| 98 |
+
$G,R,I,M$ & Fixed geography, regional coordinates, issue intensity and validity masks.\\
|
| 99 |
+
\bottomrule
|
| 100 |
+
\end{tabularx}
|
| 101 |
+
|
| 102 |
+
\vspace{7pt}
|
| 103 |
+
\keyline{Scope.}{A research model for the Western Pacific, not an operational
|
| 104 |
+
warning system. Selected examples do not establish generalization or safety.}
|
| 105 |
+
\reportpage
|
| 106 |
+
\section{Architecture and input contract}
|
| 107 |
+
The basin and moving core share a recurrent forecast cycle, but retain separate
|
| 108 |
+
spatial representations and memories. Figure~\ref{fig:architecture} shows the
|
| 109 |
+
dependency structure; the operator definitions follow in Section~\ref{sec:operators}.
|
| 110 |
+
\begin{figure}[H]
|
| 111 |
\centering
|
| 112 |
\begin{tikzpicture}[font=\small,>=Latex,
|
| 113 |
box/.style={draw=ink!50,rounded corners=3pt,align=center,text width=4.6cm,minimum height=1.15cm,inner sep=6pt,fill=ink!3},
|
|
|
|
| 135 |
\node[below=0.25cm of readout,align=center,font=\scriptsize,text=ink] {Shared transition repeated 20 times: +6, +12, $\ldots$, +120 h.\\
|
| 136 |
50 perturbed input histories $\rightarrow$ 50 rollouts $\rightarrow$ separate route, pressure and common-grid field means.};
|
| 137 |
\end{tikzpicture}
|
| 138 |
+
\caption{Coupled forecast architecture. Basin evolution supplies environmental
|
| 139 |
+
fields to the moving core; the pressure-associated readout follows the motion
|
| 140 |
+
proposal. All leads and ensemble members use the same trained network.}
|
|
|
|
|
|
|
| 141 |
\label{fig:architecture}
|
| 142 |
\end{figure}
|
| 143 |
|
| 144 |
+
\subsection{Causal inputs and spatial representations}
|
| 145 |
+
\begin{tabularx}{\linewidth}{@{}p{3.1cm}p{3.6cm}X@{}}
|
| 146 |
+
\toprule
|
| 147 |
+
\textbf{Input / state} & \textbf{Shape / sampling} & \textbf{Role and boundary}\\
|
| 148 |
+
\midrule
|
| 149 |
+
Basin history & $B\times9\times8\times25\times33$ & Nine analyses at $-48,-42,\ldots,0$ h; $0$--$60^\circ$N, $100$--$180^\circ$E, $2.5^\circ$ spacing.\\
|
| 150 |
+
Regional history & $B\times9\times1\times121\times121$ & Issue-anchored MSLP patch, $\pm15^\circ$, sampled every $0.25^\circ$; availability mask selects native or coarse initialization.\\
|
| 151 |
+
Moving core & $65\times65$; 20-km spacing & Learned pressure reconstruction over $\pm640$ km; grid spacing does not establish effective resolution.\\
|
| 152 |
+
Issue state & Centre, motion, intensity, masks & Latitude/longitude, six-hour east/north displacement, wind in knots, pressure in hPa.\\
|
| 153 |
+
Static channels & Four channels per grid & Latitude/90, longitude/180$-1$, land fraction and elevation/8000 m.\\
|
| 154 |
+
\bottomrule
|
| 155 |
+
\end{tabularx}
|
| 156 |
+
|
| 157 |
+
Basin channels are $[\mathrm{MSLP},Z_{500},u_{850},v_{850},u_{500},v_{500},u_{200},v_{200}]$
|
| 158 |
+
in hPa, metres and metres per second. Normalization uses fixed training-only
|
| 159 |
+
statistics, $\widetilde x_j=(x_j-\mu_j)/\sigma_j$. Regional coordinates remain
|
| 160 |
+
available when native history is missing. Only information available at or
|
| 161 |
+
before issue time may enter inference; retrospective archive coverage does
|
| 162 |
+
not establish operational availability. Outputs include basin and regional
|
| 163 |
+
pressure, core coordinates, centre, central pressure, masks and auxiliary wind;
|
| 164 |
+
there are no quadrant wind radii or resolved surface-wind maps.
|
| 165 |
+
\reportpage
|
| 166 |
+
\section{Neural operators}\label{sec:operators}
|
| 167 |
+
The released model has 21,452,595 trainable parameters.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 168 |
\subsection{Multiresolution evolution networks}
|
| 169 |
The basin network receives 44 channels: eight fields, 32 memory features and
|
| 170 |
four static channels. Encoder widths are $72,144,288,432$; residual-block depths
|
|
|
|
| 206 |
memory before the transition. No separate positional embedding is added:
|
| 207 |
geographic coordinates are already among the token inputs.
|
| 208 |
|
| 209 |
+
\subsection{Observation-conditioned initialization}
|
| 210 |
Write $\sample(F,\xi)$ for bilinear field sampling. Sampling uses FP32 geometry,
|
| 211 |
aligned corners and border padding for the basin; transported core anomalies
|
| 212 |
use zero padding. All nine pressure histories are sampled onto the core grid
|
|
|
|
| 225 |
The anomaly is $a_0=c_0-\sample(g_{0,\mathrm{MSLP}},\xi_0)$.
|
| 226 |
This compact prior is applied only at initialization. Future scalar pressure
|
| 227 |
is not repeatedly inserted into the field.
|
| 228 |
+
\reportpage
|
| 229 |
+
\section{One six-hour forecast transition}
|
| 230 |
+
Each step advances the basin first, proposes core motion, transports the core
|
| 231 |
+
relative to that motion, and reads out position and pressure. All twenty leads
|
| 232 |
+
share the same operators. Sampling $\sample(F,\xi)$ is bilinear; coordinates
|
| 233 |
+
and units are explicit below.
|
| 234 |
+
|
| 235 |
+
\subsection{Environmental transport}
|
| 236 |
+
Learned softmax weights blend 850-, 500- and 200-hPa winds; their initial values
|
| 237 |
+
are $(0.4,0.4,0.2)$. For basin-network output $d$, the corrected flow and next
|
| 238 |
+
normalized basin field are
|
| 239 |
+
\begin{align}
|
| 240 |
+
U_k &= \sum_\ell\alpha_\ell(u_\ell,v_\ell)_k+15\tanh(d_{8:10}),\label{eq:flow}\\
|
| 241 |
+
g_{k+1} &= \sample(g_k,D(U_k))+0.2\tanh(d_{0:8}).\label{eq:basin}
|
| 242 |
+
\end{align}
|
| 243 |
+
$U_k$ is in m\,s$^{-1}$ and $D$ is the backtraced departure grid. Six hours
|
| 244 |
+
convert 1 m\,s$^{-1}$ to 21.6 km. Coordinate conversion uses 111.2 km per
|
| 245 |
+
latitude degree and $111.2\max(\cos\phi,0.2)$ km per longitude degree.
|
| 246 |
+
The attention-conditioned memory follows the same advection and then a gated
|
| 247 |
+
update. This transport is not an exactly conservative atmospheric solver.
|
| 248 |
|
| 249 |
\subsection{Moving-frame proposal}
|
| 250 |
+
A $44\to64\to3$ GELU MLP reads 32 sampled memory channels, eight basin fields,
|
| 251 |
previous motion divided by 100, and two issue-intensity masks. For output $b$,
|
| 252 |
+
\begin{equation}
|
| 253 |
m_{k+1}=0.7(21.6)U_k(z_k)+0.3m_k+100\tanh(b_{0:2}).
|
| 254 |
+
\end{equation}
|
| 255 |
+
This km-per-six-hour displacement defines the proposed origin relative to
|
| 256 |
+
$z_k$, not the final reported centre. Auxiliary wind becomes
|
|
|
|
| 257 |
$v_{k+1}=\max(0,v_k+5b_2)$.
|
| 258 |
|
| 259 |
+
\subsection{Core transport in the moving frame}
|
| 260 |
+
The core network reads previous core pressure, the old sampled environment,
|
| 261 |
+
memory and geography at the proposed origin. Its local velocity is
|
| 262 |
$U^{\rm local}=\sample(U_k)+5\tanh(d^c_{1:3})$.
|
| 263 |
+
Subtracting $m_{k+1}/21.6$ removes bulk frame motion. The relative departure
|
| 264 |
+
grid $D_c$ also accounts for the previous centre--origin offset:
|
| 265 |
+
\begin{align}
|
|
|
|
| 266 |
a_{k+1}&=\sample(a_k,D_c)+0.1d^c_0,\\
|
| 267 |
c_{k+1}&=\sample(g_{k+1,\mathrm{MSLP}},\xi_{k+1})+Ta_{k+1}.
|
| 268 |
+
\end{align}
|
| 269 |
+
The anomaly tendency $d^c_0$ is unbounded by $\tanh$. Core memory uses the same
|
| 270 |
+
relative sampling followed by its gated update.
|
| 271 |
+
|
| 272 |
+
\subsection{Pressure-associated readout}
|
| 273 |
+
Convert core pressure to hPa, $P=\sigma_p c_{k+1}+\mu_p$. For cells no more
|
| 274 |
+
than 300 km from the proposed origin, define
|
| 275 |
+
\begin{equation}
|
| 276 |
+
w_i=\operatorname{softmax}_i\!\left(-P_i/2-r_i^2/(2\,180^2)\right).
|
| 277 |
+
\end{equation}
|
| 278 |
+
The weighted mean of cell latitude/longitude gives $z_{k+1}$; central pressure
|
| 279 |
+
is $\sample(P,z_{k+1})$, not necessarily the grid minimum. Cells outside the
|
| 280 |
+
radius are excluded, reducing association with unrelated depressions.
|
| 281 |
+
|
| 282 |
+
\subsection{Geographic pressure reconstruction}
|
| 283 |
+
The issue-anchored regional field is
|
| 284 |
+
\begin{equation}
|
|
|
|
| 285 |
\widehat p_R=\sample(g_{k+1,\mathrm{MSLP}},R)
|
| 286 |
+
+\sample(Ta_{k+1},R;\mathrm{zero\ padding}),
|
| 287 |
+
\end{equation}
|
| 288 |
+
then de-normalized to hPa. This is the model's basin-plus-core field; future
|
| 289 |
+
observations or display-only vortices are never substituted. Coverage masks
|
| 290 |
+
must remain visible. Resampling and local centre association mean the route,
|
| 291 |
+
central pressure and displayed field minimum need not coincide exactly.
|
| 292 |
+
\reportpage
|
| 293 |
+
\section{Training, ensembles and reproducibility}
|
| 294 |
+
\subsection{Autoregression and learning objectives}
|
|
|
|
| 295 |
Inference initializes once and repeats the transition twenty times, using only
|
| 296 |
fixed issue context and its own predicted future state. The implemented
|
| 297 |
single-step loss is
|
|
|
|
| 314 |
Training-only normalization, storm-level split boundaries and causal input
|
| 315 |
selection remain essential implementation constraints.
|
| 316 |
|
| 317 |
+
\subsection{Single forecasts and 50-member means}
|
| 318 |
Evaluation-mode inference is deterministic for fixed inputs and weights.
|
| 319 |
The documented ensemble perturbs only normalized historical weather: independent
|
| 320 |
Gaussian fields are spatially smoothed with kernels five (basin) and nine
|
|
|
|
| 334 |
learned latent samples. Spread is not automatically calibrated uncertainty.
|
| 335 |
Member counts, perturbation settings and seed policy must accompany the outputs.
|
| 336 |
|
| 337 |
+
\subsection{Release identity and reproducibility}
|
| 338 |
+
The public model is \textbf{Trackformer 1.2}; the manifest retains internal
|
| 339 |
+
checkpoint identity 1.2.73 epoch 4 and exact source/weight hashes. The following
|
| 340 |
+
artifacts are distributed in the
|
| 341 |
+
\href{https://github.com/yu314-coder/typhoon-predict}{release repository}:
|
| 342 |
+
\begin{tabularx}{\linewidth}{@{}p{4.2cm}X@{}}
|
| 343 |
+
\toprule
|
| 344 |
+
\textbf{Artifact} & \textbf{Purpose}\\
|
| 345 |
+
\midrule
|
| 346 |
+
\texttt{model.py} & Environmental attention and additional temporal losses.\\
|
| 347 |
+
\texttt{baseline\_model.py} & Moving-core evolution and pressure-associated readout.\\
|
| 348 |
+
\texttt{v165\_base.py} & Sampling, multiscale evolution blocks and recurrent memory.\\
|
| 349 |
+
\texttt{manifest.json} & Released grids, normalization and checkpoint/source identity.\\
|
| 350 |
+
Saved evaluation arrays & 270-storm final metrics; Fung-wong route, fields and member provenance.\\
|
| 351 |
+
\bottomrule
|
| 352 |
+
\end{tabularx}
|
| 353 |
+
Model sources are in \path{models/trackformer_1_2_field/}; evaluation arrays
|
| 354 |
+
are in \path{evaluation/}. This document embeds figure coordinates and needs
|
| 355 |
+
no external images. The released plotting utility regenerates the figure
|
| 356 |
+
data from saved arrays; it does not perform new inference.
|
| 357 |
+
\reportpage
|
| 358 |
\section{Development evaluation}
|
| 359 |
\label{sec:evaluation}
|
| 360 |
+
\subsection{Matched cases, different forecast pipelines}
|
| 361 |
+
The completed comparison uses \textbf{1,473 daily issues from 270 storms},
|
| 362 |
+
with 20 six-hour leads per issue. Scores average leads within daily issues,
|
| 363 |
+
issues within storms, then storms equally. The cohort was frozen before inference.
|
| 364 |
+
Trackformer 1.2 averages 50 causal input perturbations; 1.1 uses its causal
|
| 365 |
+
route pipeline. Cases and targets match, but inputs differ. This is broader
|
| 366 |
+
development evidence, not an untouched test or architecture ablation.
|
| 367 |
+
\begin{figure}[H]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 368 |
\centering
|
| 369 |
\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
|
| 370 |
\begin{scope}[shift={(0,0)}]
|
|
|
|
| 372 |
\node[rotate=90] at (-.75,1.75) {km; lower is better};
|
| 373 |
\draw[black!10] (0,0.0000)--(5,0.0000);
|
| 374 |
\node[left] at (0,0.0000) {0};
|
| 375 |
+
\draw[black!10] (0,1.7500)--(5,1.7500);
|
| 376 |
+
\node[left] at (0,1.7500) {500};
|
| 377 |
+
\draw[black!10] (0,3.5000)--(5,3.5000);
|
| 378 |
+
\node[left] at (0,3.5000) {1000};
|
| 379 |
+
\fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.7944);
|
| 380 |
+
\node[above] at (1.4,2.7944) {798.4};
|
| 381 |
\node[below=3pt] at (1.4,0) {1.1};
|
| 382 |
+
\fill[modelteal] (3.22,0) rectangle (3.98,1.6492);
|
| 383 |
+
\node[above] at (3.6,1.6492) {471.2};
|
| 384 |
\node[below=3pt] at (3.6,0) {1.2 mean of 50};
|
| 385 |
\draw[black!55] (0,3.5)--(0,0)--(5,0);
|
| 386 |
\end{scope}
|
|
|
|
| 395 |
\node[left] at (0,2.0000) {40};
|
| 396 |
\draw[black!10] (0,3.0000)--(5,3.0000);
|
| 397 |
\node[left] at (0,3.0000) {60};
|
| 398 |
+
\fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.5790);
|
| 399 |
+
\node[above] at (1.4,2.5790) {51.58};
|
| 400 |
\node[below=3pt] at (1.4,0) {1.1};
|
| 401 |
+
\fill[modelteal] (3.22,0) rectangle (3.98,1.7480);
|
| 402 |
+
\node[above] at (3.6,1.7480) {34.96};
|
| 403 |
\node[below=3pt] at (3.6,0) {1.2 mean of 50};
|
| 404 |
\draw[black!55] (0,3.5)--(0,0)--(5,0);
|
| 405 |
\end{scope}
|
|
|
|
| 433 |
\node[font=\small,rotate=90] at (-1.05,1.850) {Mean position error (km)};
|
| 434 |
\begin{scope}
|
| 435 |
\clip (0,0) rectangle (12.2,3.7);
|
| 436 |
+
\draw[baselinegray,dashed,thick] (0.0000,0.1223)--(0.6421,0.2469)--(1.2842,0.3744)--(1.9263,0.5056)--(2.5684,0.6418)--(3.2105,0.7768)--(3.8526,0.9131)--(4.4947,1.0510)--(5.1368,1.1916)--(5.7789,1.3369)--(6.4211,1.4867)--(7.0632,1.6402)--(7.7053,1.8006)--(8.3474,1.9621)--(8.9895,2.1304)--(9.6316,2.3032)--(10.2737,2.4856)--(10.9158,2.6699)--(11.5579,2.8565)--(12.2000,3.0458);
|
| 437 |
+
\draw[modelteal,solid,thick] (0.0000,0.0893)--(0.6421,0.1686)--(1.2842,0.2389)--(1.9263,0.3043)--(2.5684,0.3683)--(3.2105,0.4346)--(3.8526,0.5041)--(4.4947,0.5762)--(5.1368,0.6539)--(5.7789,0.7389)--(6.4211,0.8312)--(7.0632,0.9296)--(7.7053,1.0329)--(8.3474,1.1399)--(8.9895,1.2537)--(9.6316,1.3707)--(10.2737,1.4963)--(10.9158,1.6281)--(11.5579,1.7666)--(12.2000,1.9085);
|
| 438 |
\end{scope}
|
| 439 |
\draw[baselinegray,dashed,thick] (0.250,-1.23)--(0.800,-1.23);
|
| 440 |
\node[anchor=west] at (0.900,-1.23) {Trackformer 1.1};
|
| 441 |
\draw[modelteal,solid,thick] (6.350,-1.23)--(6.900,-1.23);
|
| 442 |
\node[anchor=west] at (7.000,-1.23) {Trackformer 1.2: mean of 50};
|
| 443 |
\end{tikzpicture}
|
| 444 |
+
\caption{Completed equal-storm development comparison, with every daily issue
|
| 445 |
+
retained. Positions use a common issue-centred local-km projection. Heading
|
| 446 |
+
uses common steps where truth and both models move more than 1 km.
|
| 447 |
+
No uncertainty intervals are claimed.}
|
|
|
|
|
|
|
| 448 |
\label{fig:benchmark}
|
| 449 |
\end{figure}
|
| 450 |
+
\subsection{Interpretation: alignment, not shape alone}
|
| 451 |
+
Mean position error improves by 41.0\%, from 798.4 to 471.2 km; +120 h error
|
| 452 |
+
falls from 1,646.4 to 1,031.6 km. At +6 h, error is 48.3 versus 66.1 km.
|
| 453 |
+
Centred shape similarity is 0.8837 versus 0.7544 for 1.1; normalized
|
| 454 |
+
Fr\'echet path similarity is 0.6560 versus 0.5345. These are separate
|
| 455 |
+
measures, not accuracy percentages or evidence of a win on every storm.
|
| 456 |
+
|
| 457 |
+
Actual alignment requires predicted and observed positions to coincide at the
|
| 458 |
+
same valid times. Removing translation and scale can make displaced routes
|
| 459 |
+
look similar; the shape score cannot replace geographic error or an unshifted
|
| 460 |
+
overlay. Path similarity responds to displacement but is not a strict
|
| 461 |
+
same-time distance.
|
| 462 |
+
|
| 463 |
+
\keyline{Pressure coverage.}{The 1.2 central-pressure MAE is 12.62 hPa on
|
| 464 |
+
40 storms with valid labels; basin MSLP MAE is 2.72 hPa on 270 storms.
|
| 465 |
+
No matched 1.1 pressure comparison is available. All 1,473 saved forecast
|
| 466 |
+
hashes were verified; the earlier 270-case/90-storm study remains separate.}
|
| 467 |
+
\reportpage
|
| 468 |
+
\section{Fung-wong: route alignment and intensity}
|
| 469 |
+
The selected forecast begins at 00 UTC on 7 November 2025 and uses the saved
|
| 470 |
+
50-member policy. Figure~\ref{fig:fungwong} compares exact six-hour positions
|
| 471 |
+
without translating, rotating or scaling either route. Great-circle errors
|
| 472 |
+
are 130.9 km on average, 142.3 km at +120 h and 198.3 km at maximum; all 20
|
| 473 |
+
leads are within 200 km. These are separate from the benchmark's local-km scores.
|
| 474 |
+
\begin{figure}[H]
|
|
|
|
|
|
|
| 475 |
\centering
|
| 476 |
\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
|
| 477 |
\node[font=\small\bfseries,align=center] at (6.100,7.800) {Fung-wong: geographic routes, +0 to +120 h};
|
|
|
|
| 597 |
4-hPa isobars; none of the pressure fields is moved to match observations.}
|
| 598 |
\label{fig:fungwong}
|
| 599 |
\end{figure}
|
| 600 |
+
The broad curves align, but position, timing and intensity differences remain.
|
| 601 |
+
This user-selected development example is not representative performance.
|
| 602 |
+
The \href{https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4}{companion MP4}
|
| 603 |
+
animates the saved fields; no additional inference or route adjustment is used.
|
| 604 |
+
\reportpage
|
| 605 |
+
\section{Pressure-field interpretation and limitations}
|
| 606 |
+
\begin{figure}[H]
|
| 607 |
\centering
|
| 608 |
+
\begin{tikzpicture}[scale=0.9,x=1cm,y=1cm,font=\scriptsize]
|
| 609 |
\node[font=\small\bfseries,align=center] at (5.750,8.900) {Fung-wong: model-mean MSLP at +48 h};
|
| 610 |
\draw[black!10] (0.676,0)--(0.676,8.5);
|
| 611 |
\node[below=3pt] at (0.676,0) {124};
|
|
|
|
| 672 |
\draw[routepink,solid,thick] (6.000,-1.23)--(6.550,-1.23);
|
| 673 |
\node[anchor=west] at (6.650,-1.23) {Forecast route to +48 h};
|
| 674 |
\end{tikzpicture}
|
| 675 |
+
\caption{Saved model-mean regional MSLP at +48 h (9 November 2025, 00 UTC), with 4-hPa isobars labelled every 8 hPa. Routes extend only to this valid time. The western boundary is $123^\circ$E. The field is not shifted to the observed centre; its minimum need not equal member-mean central pressure.}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 676 |
\label{fig:pressuremap}
|
| 677 |
\end{figure}
|
| 678 |
+
\subsection{Intensity labels and map coverage}
|
| 679 |
+
Pressure truth is JMA best track from IBTrACS \texttt{TOKYO\_PRES}, not a JMA
|
| 680 |
+
forecast. Native-detail history was unavailable; $0.25^\circ$ output sampling
|
| 681 |
+
does not establish native resolution. From +66 h the centre leaves the fixed
|
| 682 |
+
regional patch. The video then shows only the saved coarse basin field there,
|
| 683 |
+
while the moving-core readout supplies central pressure. No detailed core is
|
| 684 |
+
invented beyond coverage.
|
| 685 |
+
|
| 686 |
+
There is no verified matched 1.1 pressure array for this benchmark. On the
|
| 687 |
+
repeatedly inspected TIP ten-start diagnostic, 1.2 pressure MAE is worse
|
| 688 |
+
(30.7 versus 20.1 hPa). Those overlapping starts are not independent tests;
|
| 689 |
+
route gains do not establish an all-metric improvement.
|
| 690 |
+
|
| 691 |
+
\subsection{What remains unverified}
|
| 692 |
+
The completed cohort includes 40 recent and 230 historical storms (1980--1999),
|
| 693 |
+
outside this checkpoint's fitting and validation years but not certified unused
|
| 694 |
+
across prior experiments. Complete five-day truth excludes short remaining
|
| 695 |
+
lifetimes. Historical hindcasts use a model trained on later years and
|
| 696 |
+
retrospective analyses. Generalization still requires an unused whole-storm
|
| 697 |
+
holdout, training-only preprocessing and whole-storm uncertainty resampling.
|
| 698 |
+
|
| 699 |
+
\subsection{Conclusion}
|
| 700 |
+
Trackformer 1.2 links recurrent environmental evolution to a translating
|
| 701 |
+
pressure anomaly and coupled geographic readout. Its development results
|
| 702 |
+
improve aggregate route measures over the archived 1.1 pipeline, but leave
|
| 703 |
+
large absolute errors and intensity failures; the older cohort also showed
|
| 704 |
+
a short-lead regression. The
|
| 705 |
+
released model is a reproducible research artifact, not a substitute for
|
| 706 |
+
official warnings.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 707 |
\end{document}
|
release_tools/build_paper_figures.py
CHANGED
|
@@ -50,11 +50,11 @@ def curves(x,y1,y2,xlim,ylim,xticks,yticks,width,height,title,xlabel,ylabel,labe
|
|
| 50 |
return '\n'.join(lines)
|
| 51 |
|
| 52 |
def generate():
|
| 53 |
-
m=json.loads((ROOT/'evaluation/
|
| 54 |
with np.load(ROOT/'evaluation/release_data/fung_wong_video.npz') as z:a={k:z[k] for k in z.files}
|
| 55 |
lead=np.arange(6,121,6)
|
| 56 |
bars=[r'\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]']
|
| 57 |
-
for offset,title,values,top,ticks,unit in [(0,'Mean position error',[
|
| 58 |
bars.append(rf'\begin{{scope}}[shift={{({offset},0)}}]')
|
| 59 |
bars.append(rf'\node[font=\small\bfseries] at (2.5,4.05) {{{title}}};')
|
| 60 |
bars.append(rf'\node[rotate=90] at (-.75,1.75) {{{unit}}};')
|
|
@@ -96,7 +96,10 @@ def main():
|
|
| 96 |
text=match.group()
|
| 97 |
for label,picture in figures.items():
|
| 98 |
if '\\label{'+label+'}' in text:
|
| 99 |
-
|
|
|
|
|
|
|
|
|
|
| 100 |
return text
|
| 101 |
new=re.sub(r'\\begin\{figure\}.*?\\end\{figure\}',replace,new,flags=re.S)
|
| 102 |
print('*** Begin Patch\n*** Update File: '+str(path))
|
|
|
|
| 50 |
return '\n'.join(lines)
|
| 51 |
|
| 52 |
def generate():
|
| 53 |
+
m=json.loads((ROOT/'evaluation/daily_storm_final.json').read_text())['aggregate_equal_storm']
|
| 54 |
with np.load(ROOT/'evaluation/release_data/fung_wong_video.npz') as z:a={k:z[k] for k in z.files}
|
| 55 |
lead=np.arange(6,121,6)
|
| 56 |
bars=[r'\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]']
|
| 57 |
+
for offset,title,values,top,ticks,unit in [(0,'Mean position error',[round(m[k]['mean_track_error_km'],1) for k in ['1.1','1.2']],1000,[0,500,1000],'km; lower is better'),(7.2,'Six-hour heading error',[round(m[k]['direction_error_deg'],2) for k in ['1.1','1.2']],70,[0,20,40,60],'degrees; lower is better')]:
|
| 58 |
bars.append(rf'\begin{{scope}}[shift={{({offset},0)}}]')
|
| 59 |
bars.append(rf'\node[font=\small\bfseries] at (2.5,4.05) {{{title}}};')
|
| 60 |
bars.append(rf'\node[rotate=90] at (-.75,1.75) {{{unit}}};')
|
|
|
|
| 96 |
text=match.group()
|
| 97 |
for label,picture in figures.items():
|
| 98 |
if '\\label{'+label+'}' in text:
|
| 99 |
+
if label == 'fig:pressuremap':
|
| 100 |
+
picture = picture.replace(r'\begin{tikzpicture}[x=1cm', r'\begin{tikzpicture}[scale=0.9,x=1cm')
|
| 101 |
+
opening = text[:text.index('\\begin{tikzpicture}')]
|
| 102 |
+
return opening+picture+'\n'+text[text.index('\\caption'):]
|
| 103 |
return text
|
| 104 |
new=re.sub(r'\\begin\{figure\}.*?\\end\{figure\}',replace,new,flags=re.S)
|
| 105 |
print('*** Begin Patch\n*** Update File: '+str(path))
|
release_tools/plot_daily_storm_final.py
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Plot the completed equal-storm development benchmark; no new inference."""
|
| 2 |
+
from pathlib import Path
|
| 3 |
+
import json
|
| 4 |
+
import matplotlib
|
| 5 |
+
matplotlib.use('Agg')
|
| 6 |
+
import matplotlib.pyplot as plt
|
| 7 |
+
|
| 8 |
+
ROOT = Path(__file__).resolve().parents[1]
|
| 9 |
+
report = json.loads((ROOT / 'evaluation/daily_storm_final.json').read_text())
|
| 10 |
+
assert report['status'] == 'complete' and report['completed_storms'] == 270
|
| 11 |
+
metrics = report['aggregate_equal_storm']
|
| 12 |
+
panels = [
|
| 13 |
+
('mean_track_error_km', 'Mean position error ↓', 'km', 1),
|
| 14 |
+
('track_error_120h_km', '+120 h position error ↓', 'km', 1),
|
| 15 |
+
('direction_error_deg', 'Six-hour heading error ↓', 'degrees', 2),
|
| 16 |
+
('shape_similarity', 'Centred route shape ↑', 'similarity (not accuracy)', 4),
|
| 17 |
+
('path_similarity', 'Geographic path similarity ↑', 'similarity (not accuracy)', 4),
|
| 18 |
+
('frechet_distance_km', 'Fréchet distance ↓', 'km', 1),
|
| 19 |
+
]
|
| 20 |
+
plt.rcParams.update({'font.family': 'DejaVu Sans', 'font.size': 10, 'axes.spines.top': False, 'axes.spines.right': False})
|
| 21 |
+
fig, axes = plt.subplots(2, 3, figsize=(13, 8))
|
| 22 |
+
for ax, (key, title, unit, digits) in zip(axes.flat, panels):
|
| 23 |
+
values = [metrics[m][key] for m in ['1.1', '1.2']]
|
| 24 |
+
bars = ax.bar(['1.1', '1.2 · mean of 50'], values, width=.55, color=['#65758b', '#007d78'])
|
| 25 |
+
ax.bar_label(bars, labels=[f'{v:,.{digits}f}' for v in values], padding=6, fontweight='bold')
|
| 26 |
+
ax.set_title(title, loc='left', pad=14, fontweight='bold')
|
| 27 |
+
ax.set_ylabel(unit)
|
| 28 |
+
ax.set_ylim(0, 1.06 if 'similarity' in key else max(values)*1.22)
|
| 29 |
+
ax.set_axisbelow(True)
|
| 30 |
+
ax.grid(axis='y', alpha=.16)
|
| 31 |
+
fig.suptitle('Trackformer 1.2 vs 1.1 | 270 storms · 1,473 daily forecasts', x=.07, ha='left', fontsize=17, fontweight='bold')
|
| 32 |
+
fig.text(.07, .922, 'Complete +6 to +120 h rollouts · each storm has equal weight · broader development evidence', color='#526174')
|
| 33 |
+
fig.text(.07, .022, 'Different input pipelines; not an architecture ablation or certified untouched test. No matched 1.1 pressure comparison.', fontsize=9, color='#526174')
|
| 34 |
+
fig.subplots_adjust(top=.84, bottom=.11, left=.07, right=.98, hspace=.42, wspace=.35)
|
| 35 |
+
fig.savefig(ROOT / 'evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png', dpi=180, facecolor='white')
|