euler314 commited on
Commit
a9c91b1
·
verified ·
1 Parent(s): 9cbb93f

Release 1.2: completed 270-storm benchmark and revised paper

Browse files
.gitattributes CHANGED
@@ -26,3 +26,4 @@ docs/trackformer_1_2_yagi.mp4 filter=lfs diff=lfs merge=lfs -text
26
  docs/trackformer_1_2_fung_wong.mp4 filter=lfs diff=lfs merge=lfs -text
27
  evaluation/trackformer_1_2_fung_wong_video_poster.png filter=lfs diff=lfs merge=lfs -text
28
  evaluation/release_data/fung_wong_video.npz filter=lfs diff=lfs merge=lfs -text
 
 
26
  docs/trackformer_1_2_fung_wong.mp4 filter=lfs diff=lfs merge=lfs -text
27
  evaluation/trackformer_1_2_fung_wong_video_poster.png filter=lfs diff=lfs merge=lfs -text
28
  evaluation/release_data/fung_wong_video.npz filter=lfs diff=lfs merge=lfs -text
29
+ evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -11,7 +11,7 @@ tags:
11
 
12
  **[Download 1.2: weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Read the illustrated paper](paper/trackformer.pdf)** · **[Watch Fung-wong](https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4)** · **[Hugging Face model](https://huggingface.co/euler314/typhoon-predict)**
13
 
14
- Release status: **research prerelease**, not operational certification. The package contains the released 1.2 weights, inference code, input contract, saved evaluation data and illustrated paper. The original 1.1 release remains separate.
15
 
16
  Trackformer 1.2 is a **research-only Western Pacific tropical-cyclone forecast model**. It evolves a sea-level-pressure (MSLP) field and a moving storm-centred pressure core every six hours through +120 h. A track and central-pressure estimate are extracted from the evolving core, not independently drawn on top of a pressure image. The prior [Trackformer 1.1 release](https://github.com/yu314-coder/typhoon-predict/releases/tag/trackformer-1.1) remains available and unchanged.
17
 
@@ -54,7 +54,30 @@ The implementation is in [`models/trackformer_1_2_field/`](models/trackformer_1_
54
 
55
  Grid spacing describes the representation, not independently demonstrated effective resolution. A mean of member centres is not necessarily the minimum of the displayed mean pressure field.
56
 
57
- ## 270-case comparison: direction, route shape and position
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
 
59
  This comparison uses **270 forecast cases from 90 storms**, with the same issue rows, observed routes and 20 six-hour leads through +120 h. The number 270 counts forecast cases. The 1.2 prediction for each case is a **mean of 50 causal input-perturbation members**; 1.1 uses its archived causal route pipeline.
60
 
@@ -82,23 +105,7 @@ These are **six selected best-performing examples from six distinct storms**, no
82
 
83
  The 1.2 mean has better aggregate direction, shape and position scores on this previously inspected development cohort. That does not establish an overall win on every storm or pressure metric. A same-270 1.1 central-pressure prediction array was not verified in these route artifacts, so it is not assigned a pressure-error bar. The separate 1.2 report gives 15.73 hPa central-pressure MAE and 2.54 hPa area-weighted regional MSLP MAE. TIP remains a separate diagnostic: its pressure error was worse for 1.2 (30.7 vs 20.1 hPa). See [metric definitions and limitations](docs/trackformer_1_2_evaluation.md).
84
 
85
- ### Expanded daily-issue benchmark: 270 distinct typhoons
86
-
87
- A **1,473-case / 270-storm** benchmark is in progress; it is separate from the completed 270-case comparison above. Each storm contributes at most one forecast per UTC day, through +120 h. Daily errors are averaged within each storm, then storm scores receive equal weight. Selection was frozen before inference without filtering on forecast quality. See the [protocol](docs/daily_storm_benchmark.md) and [frozen cohort](evaluation/release_data/daily_storm_cohort.json).
88
-
89
- **Dated snapshot — 2026-09-29T06:09:06.508182+00:00 (UTC), not a live counter:** 808/1,473 daily forecasts saved; **159/270 storms fully completed**. The 1.2 forecasts use 50 causal input-perturbation members. [Machine-readable snapshot](evaluation/daily_storm_progress_20260929.json).
90
-
91
- | Preliminary equal-storm metric | 1.1 | 1.2 · mean of 50 |
92
- | --- | ---: | ---: |
93
- | Mean track error, +6 to +120 h ↓ | 843.9 km | 487.1 km |
94
- | Track error at +120 h ↓ | 1733.3 km | 1068.8 km |
95
- | Six-hour direction error ↓ | 51.81° | 35.33° |
96
- | Centred route-shape similarity ↑ | 0.7497 | 0.8842 |
97
- | Geographic path similarity ↑ | 0.5310 | 0.6611 |
98
-
99
- **Incomplete development evidence:** only the 159 fully completed storms enter this table. Execution order is not random, and the remaining storms can change the result. These are not the final 270-storm scores or an untouched-test claim. Do not compare this subset directly with the differently sampled 270-case table above.
100
-
101
- 1.2-only pressure diagnostics: central-pressure MAE **12.62 hPa** over 40 storms with valid pressure labels; basin MSLP MAE **2.70 hPa** over 159 storms. There is no matched 1.1 pressure result in this run, and basin-wide error does not establish core-pressure-map accuracy.
102
 
103
  ## Fung-wong pressure forecast — MP4
104
 
 
11
 
12
  **[Download 1.2: weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Read the illustrated paper](paper/trackformer.pdf)** · **[Watch Fung-wong](https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4)** · **[Hugging Face model](https://huggingface.co/euler314/typhoon-predict)**
13
 
14
+ Release status: **1.2 research release**, not operational certification. The package contains the released 1.2 weights, inference code, input contract, completed 270-storm evaluation and illustrated paper. The original 1.1 release remains separate.
15
 
16
  Trackformer 1.2 is a **research-only Western Pacific tropical-cyclone forecast model**. It evolves a sea-level-pressure (MSLP) field and a moving storm-centred pressure core every six hours through +120 h. A track and central-pressure estimate are extracted from the evolving core, not independently drawn on top of a pressure image. The prior [Trackformer 1.1 release](https://github.com/yu314-coder/typhoon-predict/releases/tag/trackformer-1.1) remains available and unchanged.
17
 
 
54
 
55
  Grid spacing describes the representation, not independently demonstrated effective resolution. A mean of member centres is not necessarily the minimum of the displayed mean pressure field.
56
 
57
+ ## Completed benchmark: 270 storms · 1,473 daily forecasts
58
+
59
+ One typhoon on one UTC day is one case. Each issue forecasts **+6 to +120 h**. We average lead errors within each issue, daily scores within each storm, then the **270 storm scores equally**. This prevents long-lived storms dominating the benchmark. All 1,473 cases finished on **29 September 2026, 08:22 UTC**, with every saved forecast SHA-256 verified.
60
+
61
+ | Equal-storm metric | 1.1 | 1.2 · mean of 50 | Preferred |
62
+ | --- | ---: | ---: | --- |
63
+ | Mean track error, +6 to +120 h | 798.4 km | **471.2 km** | Lower |
64
+ | Track error at +120 h | 1,646.4 km | **1,031.6 km** | Lower |
65
+ | Six-hour direction error | 51.58° | **34.96°** | Lower |
66
+ | Centred route-shape similarity | 0.7544 | **0.8837** | Higher |
67
+ | Geographic path similarity | 0.5345 | **0.6560** | Higher |
68
+ | Fréchet distance | 1,658.1 km | **1,045.8 km** | Lower |
69
+
70
+ ![Complete equal-storm benchmark: direction, route alignment, shape and position](evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png)
71
+
72
+ Mean track error is **41.0% lower** for 1.2 on this cohort. Direction, shape and geographic alignment remain separate measures, not a combined score. A shape score near one does not guarantee overlapping routes at matching times. The 1.2 forecasts are means of **50 distinct causal input perturbations**, not 50 independently trained networks. Forecast pipelines differ, so this is not an architecture ablation.
73
+
74
+ **Pressure coverage:** 1.2 central-pressure MAE is **12.62 hPa over 40 storms with valid labels**; basin-area-weighted MSLP MAE is **2.72 hPa over 270 storms**. There is no matched 1.1 pressure output in this run. Basin-wide error does not establish core-field accuracy.
75
+
76
+ **Limits:** the frozen cohort includes 40 recent storms and 230 historical storms from 1980–1999. It excludes this checkpoint's fitting/validation years, but has not been certified untouched across prior experiments. Historical hindcasts use retrospective analyses and a model trained on later years. Complete five-day labels are required, excluding short remaining lifetimes. No operational or no-overfitting claim follows.
77
+
78
+ [Final report and all storm scores](evaluation/daily_storm_final.json) · [Protocol](docs/daily_storm_benchmark.md) · [Frozen issues](evaluation/release_data/daily_storm_cohort.json) · [Reproduce chart](release_tools/plot_daily_storm_final.py)
79
+
80
+ ## Earlier 270-case comparison: 90 storms
81
 
82
  This comparison uses **270 forecast cases from 90 storms**, with the same issue rows, observed routes and 20 six-hour leads through +120 h. The number 270 counts forecast cases. The 1.2 prediction for each case is a **mean of 50 causal input-perturbation members**; 1.1 uses its archived causal route pipeline.
83
 
 
105
 
106
  The 1.2 mean has better aggregate direction, shape and position scores on this previously inspected development cohort. That does not establish an overall win on every storm or pressure metric. A same-270 1.1 central-pressure prediction array was not verified in these route artifacts, so it is not assigned a pressure-error bar. The separate 1.2 report gives 15.73 hPa central-pressure MAE and 2.54 hPa area-weighted regional MSLP MAE. TIP remains a separate diagnostic: its pressure error was worse for 1.2 (30.7 vs 20.1 hPa). See [metric definitions and limitations](docs/trackformer_1_2_evaluation.md).
107
 
108
+ The earlier dated partial snapshot is retained for provenance only. The completed 270-storm result at the top of this document supersedes it; the earlier 270-case comparison remains a different cohort.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
109
 
110
  ## Fung-wong pressure forecast — MP4
111
 
RELEASE_NOTES_TRACKFORMER_1_2.md CHANGED
@@ -1,4 +1,4 @@
1
- # Trackformer 1.2 — research candidate
2
 
3
  **[Download weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Download the illustrated PDF](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer.pdf)** · **[Hugging Face model card](https://huggingface.co/euler314/typhoon-predict)**
4
 
@@ -8,7 +8,7 @@ The attached archive includes inference-only weights, exact source modules, an i
8
 
9
  On the 270-case development cohort, 1.1 versus 1.2 mean-of-50 scores are: direction error **56.22° vs 44.82°**, route-shape similarity **0.7273 vs 0.8249**, path similarity **0.4983 vs 0.5727**, and mean track error **902.3 vs 714.4 km**. Direction uses 5,382 common valid steps. The old 1.1 report had stored latitude/longitude under a local-kilometre key; the new comparison repairs that coordinate interpretation and records the correction. Original artifacts remain available. These results are development comparisons, and no matched 1.1 pressure-field bar is claimed.
10
 
11
- The route gallery shows six explicitly selected best-performing examples from distinct storms, with the selection rule and all case scores supplied. It is not representative evidence; the aggregate bars retain all 270 cases. A separate larger benchmark protocol freezes **270 distinct storms / 1,473 daily issues**, with daily scores averaged per storm and storms weighted equally. That expanded protocol is separate from the completed 270-case results in this paper.
12
 
13
  The README shows the user-selected **Fung-wong MP4** in place of Yagi; the interactive PRAPIROON showcase remains removed. The forecast starts 7 November 2025 at 00 UTC and uses actual saved 50-member mean pressure fields and routes through +120 h. Central-pressure MAE is 7.36 hPa against JMA best track from IBTrACS. The broad route is similar but not perfectly overlapping: mean geographic error is 130.9 km and +120 h error is 142.3 km. This is a selected development example, not typical skill. From +66 h the forecast centre leaves the regional patch; the video explicitly labels that only the saved coarse basin field is shown there. No detailed core is invented outside coverage.
14
 
@@ -18,6 +18,18 @@ The Surigae illustration is a single forecast issued 2026-09-27 12 UTC, with 4 h
18
 
19
  This is not an operational warning service. Do not use for safety-critical decisions. A genuinely untouched storm-level holdout is still required before generalization claims.
20
 
21
- ## Expanded benchmark progress (dated snapshot)
22
 
23
- At 2026-09-29T06:09:06.508182+00:00, 808/1,473 daily forecasts and 159/270 storms were complete. This is an incomplete development snapshot, not final validation. The README includes separate preliminary direction, shape and position metrics and pressure-label coverage. The machine-readable snapshot is `evaluation/daily_storm_progress_20260929.json`. Model weights and published paper are unchanged in this documentation update.
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Trackformer 1.2 — research release
2
 
3
  **[Download weights + code](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer_1_2_field_20260929.tar.gz)** · **[Download the illustrated PDF](https://github.com/yu314-coder/typhoon-predict/releases/download/trackformer-1.2/trackformer.pdf)** · **[Hugging Face model card](https://huggingface.co/euler314/typhoon-predict)**
4
 
 
8
 
9
  On the 270-case development cohort, 1.1 versus 1.2 mean-of-50 scores are: direction error **56.22° vs 44.82°**, route-shape similarity **0.7273 vs 0.8249**, path similarity **0.4983 vs 0.5727**, and mean track error **902.3 vs 714.4 km**. Direction uses 5,382 common valid steps. The old 1.1 report had stored latitude/longitude under a local-kilometre key; the new comparison repairs that coordinate interpretation and records the correction. Original artifacts remain available. These results are development comparisons, and no matched 1.1 pressure-field bar is claimed.
10
 
11
+ The route gallery shows six explicitly selected best-performing examples from distinct storms, with the selection rule and all case scores supplied. It is not representative evidence; the earlier aggregate bars retain all 270 cases. The new primary benchmark below uses **270 distinct storms / 1,473 daily issues**, with daily scores averaged per storm and storms weighted equally. The paper now presents this completed larger cohort, separately from the earlier 270-case results.
12
 
13
  The README shows the user-selected **Fung-wong MP4** in place of Yagi; the interactive PRAPIROON showcase remains removed. The forecast starts 7 November 2025 at 00 UTC and uses actual saved 50-member mean pressure fields and routes through +120 h. Central-pressure MAE is 7.36 hPa against JMA best track from IBTrACS. The broad route is similar but not perfectly overlapping: mean geographic error is 130.9 km and +120 h error is 142.3 km. This is a selected development example, not typical skill. From +66 h the forecast centre leaves the regional patch; the video explicitly labels that only the saved coarse basin field is shown there. No detailed core is invented outside coverage.
14
 
 
18
 
19
  This is not an operational warning service. Do not use for safety-critical decisions. A genuinely untouched storm-level holdout is still required before generalization claims.
20
 
21
+ ## Completed daily benchmark and revised paper
22
 
23
+ All **1,473 daily forecasts across 270 distinct storms** completed at 2026-09-29T08:22:39Z. Every saved forecast SHA-256 was verified. Each storm has equal weight after its daily issues are averaged; all leads +6 through +120 h are included.
24
+
25
+ | Equal-storm measure | 1.1 | 1.2 mean of 50 |
26
+ | --- | ---: | ---: |
27
+ | Mean track error | 798.4 km | 471.2 km |
28
+ | +120 h track error | 1,646.4 km | 1,031.6 km |
29
+ | Direction error | 51.58° | 34.96° |
30
+ | Centred shape similarity | 0.7544 | 0.8837 |
31
+ | Geographic path similarity | 0.5345 | 0.6560 |
32
+
33
+ This is a **41.0% reduction in mean track error** on a broader development cohort, not a certified untouched test. The 1.2-only central-pressure MAE is 12.62 hPa over 40 pressure-labelled storms; basin MSLP MAE is 2.72 hPa over 270 storms. There is no matched 1.1 pressure comparison. Retrospective analyses, different input pipelines, prior model selection and complete-five-day eligibility limit interpretation.
34
+
35
+ The main README chart and redesigned eight-page paper now use these final equal-storm results. The earlier 270-case/90-storm comparison remains explicitly separate. `evaluation/daily_storm_final.json` includes all storm scores; `release_tools/plot_daily_storm_final.py` reproduces the new bars. The paper includes the architecture, equations, lead-error chart, selected route/pressure curves and model isobars. **Weights and inference implementation are unchanged.**
docs/daily_storm_benchmark.md CHANGED
@@ -1,6 +1,8 @@
1
  # Daily forecasts, equal-storm benchmark
2
 
3
- Status: inference in progress. This document specifies the frozen evaluation, not its final results.
 
 
4
 
5
  ## What counts as a case and a score
6
 
 
1
  # Daily forecasts, equal-storm benchmark
2
 
3
+ Status: **complete** at 2026-09-29T08:22:39Z. All **1,473 daily cases / 270 storms** finished; the SHA-256 of every saved forecast was verified. [Final equal-storm results and all 270 storm scores](../evaluation/daily_storm_final.json).
4
+
5
+ Mean track error is 798.4 km for 1.1 and 471.2 km for 1.2 (50-member mean); +120 h error is 1,646.4 versus 1,031.6 km. Direction error is 51.58° versus 34.96°, centred shape similarity 0.7544 versus 0.8837, and geographic path similarity 0.5345 versus 0.6560. Central-pressure MAE for 1.2 is 12.62 hPa on 40 pressure-labelled storms; basin MSLP MAE is 2.72 hPa on 270 storms. No matched 1.1 pressure result or confidence interval is claimed.
6
 
7
  ## What counts as a case and a score
8
 
evaluation/daily_storm_final.json ADDED
The diff for this file is too large to render. See raw diff
 
evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png ADDED

Git LFS Details

  • SHA256: 053110494b50bb92dd1c40a806530d1c557befcc92fb229276489d5dec8c3094
  • Pointer size: 131 Bytes
  • Size of remote file: 169 kB
paper/trackformer.pdf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8a8f055d83de393ae6f494dddc39187cf361a96a3d6d5ab920ef77f130af9369
3
- size 330235
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d5797fc86961d7163bba57903d0bed3da47f29f585aa567f24b75a6403c4a8c2
3
+ size 416164
paper/trackformer.tex CHANGED
@@ -1,65 +1,113 @@
1
- \documentclass[10pt]{article}
2
- \usepackage[margin=0.9in]{geometry}
3
- \usepackage{amsmath,amssymb,booktabs,array,microtype,xcolor}
4
- \usepackage{tikz,caption,fancyhdr}
 
5
  \usetikzlibrary{arrows.meta,positioning,calc}
6
- \definecolor{modelteal}{HTML}{00857D}
7
  \definecolor{baselinegray}{HTML}{65758B}
8
  \definecolor{ink}{HTML}{182C42}
9
  \definecolor{routepink}{HTML}{B31565}
10
- \captionsetup{font=small,labelfont=bf}
 
11
  \pagestyle{fancy}
12
  \fancyhf{}
13
- \fancyhead[L]{\small\textsc{Trackformer 1.2}}
14
- \fancyhead[R]{\small Architecture and development evaluation}
15
- \fancyfoot[C]{\small\thepage}
 
 
16
  \setlength{\headheight}{14pt}
17
- \usepackage[colorlinks=true,linkcolor=teal,urlcolor=teal]{hyperref}
18
- \hypersetup{pdftitle={Trackformer 1.2: Coupled Pressure-Field Evolution for Tropical-Cyclone Forecasting},pdfauthor={Yu Yao-Hsing},pdfsubject={Architecture and development evaluation}}
19
- \setlength{\parskip}{4pt}
20
  \setlength{\parindent}{0pt}
 
 
 
 
 
 
 
 
 
 
 
 
 
21
  \newcommand{\sample}{\mathcal{S}}
22
- \title{\textbf{Trackformer 1.2: Coupled Pressure-Field\\Evolution for Tropical-Cyclone Forecasting}\\[5pt]\large Architecture, Forecast Dynamics and Development Evaluation}
23
- \author{Yu Yao-Hsing\\\small\url{https://github.com/yu314-coder/typhoon-predict}}
24
- \date{Research technical report --- 29 September 2026}
25
  \begin{document}
26
- \maketitle
27
- \begin{abstract}
28
- Trackformer 1.2 advances an atmospheric basin state and a translating pressure
29
- core to forecast tropical-cyclone motion and central pressure. Recurrent
30
- convolutional networks evolve the fields; multiscale attention supplies
31
- environmental context. The reported centre is associated with a depression in
32
- the predicted core, and central pressure is sampled from that same field.
33
- Twenty six-hour transitions produce a five-day forecast. We describe the
34
- released architecture, input contract, update equations, objectives and
35
- 50-member input-perturbation ensemble, illustrated by a module diagram.
36
- On a previously inspected development cohort of 270 forecast cases from
37
- 90 storms, mean track error is 714.4 km for the 1.2 ensemble versus 902.3 km
38
- for the archived 1.1 route pipeline. These are output comparisons, not a
39
- controlled architecture ablation or evidence of untouched generalization.
40
- A selected Fung-wong example illustrates geographic route and pressure
41
- comparisons while retaining visible position, timing and intensity differences. All figures retain the saved
42
- predictions and explicitly separate aggregate evidence from selected examples.
43
- \end{abstract}
44
-
45
- \section{Design and forecast state}
46
- The system is not a route predictor followed by an independently drawn pressure
47
- image. It evolves pressure and atmospheric states, proposes a translation of a
48
- storm-following frame, then associates a centre with the predicted pressure
49
- within that frame. Translation is learned and steering-conditioned: this is
50
- neither a pressure-only tracker nor a complete atmospheric numerical solver.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51
 
52
  At step $k$, the recurrent state is
53
  \[
54
  s_k=(g_k,h^g_k,c_k,h^c_k,a_k,o_k,z_k,m_k,v_k;G,R,I,M).
55
  \]
56
- Here $g$ is the normalized eight-channel basin field; $h^g$ its memory;
57
- $c$ the core pressure; $h^c$ core memory; $a$ its anomaly relative to the basin;
58
- $o$ the core-grid origin; $z$ the associated storm centre; $m$ displacement per
59
- six hours; and $v$ auxiliary maximum wind. Static geography $G$, regional
60
- coordinates $R$, issue intensity $I$ and masks $M$ stay fixed.
61
-
62
- \begin{figure}[htbp]
 
 
 
 
 
 
 
 
 
 
 
 
63
  \centering
64
  \begin{tikzpicture}[font=\small,>=Latex,
65
  box/.style={draw=ink!50,rounded corners=3pt,align=center,text width=4.6cm,minimum height=1.15cm,inner sep=6pt,fill=ink!3},
@@ -87,47 +135,36 @@ basin + regional pressure fields, coverage masks and auxiliary wind};
87
  \node[below=0.25cm of readout,align=center,font=\scriptsize,text=ink] {Shared transition repeated 20 times: +6, +12, $\ldots$, +120 h.\\
88
  50 perturbed input histories $\rightarrow$ 50 rollouts $\rightarrow$ separate route, pressure and common-grid field means.};
89
  \end{tikzpicture}
90
- \caption{Released model structure. The basin branch supplies environmental
91
- information to the moving core; motion is proposed before the pressure core determines the reported
92
- centre. Boxes show functional groups, not independent models. Future observations
93
- are excluded from inference. The ensemble contains 50 input realizations of one
94
- network, not 50 trained networks.}
95
  \label{fig:architecture}
96
  \end{figure}
97
 
98
- \section{Inputs, geometry and outputs}
99
- Nine past analyses at $-48,-42,\ldots,0$ h form a tensor of shape
100
- $B\times9\times8\times25\times33$. The basin spans $60$ to $0^\circ$N and
101
- $100$ to $180^\circ$E at $2.5^\circ$ spacing. Channels are
102
- \[
103
- [\mathrm{MSLP},Z_{500},u_{850},v_{850},u_{500},v_{500},u_{200},v_{200}],
104
- \]
105
- in hPa, metres and metres per second. Channel normalization is
106
- $\widetilde x_j=(x_j-\mu_j)/\sigma_j$, with the released training-only constants.
107
- Statistics are not refitted for each forecast.
108
-
109
- Optional regional pressure history has shape $B\times9\times1\times121\times121$
110
- on an issue-anchored patch sampled every $0.25^\circ$, with offsets of
111
- $\pm15^\circ$. A detail-availability flag chooses native regional history or
112
- coarse basin pressure during initialization. The regional coordinate grid is
113
- still supplied when native observations are unavailable.
114
-
115
- Four static channels encode latitude/$90$, longitude/$180-1$, land fraction
116
- and elevation/$8000$ m. The issue state supplies latitude/longitude, east/north
117
- motion in km per six hours, maximum wind in knots, central pressure in hPa and
118
- two intensity-availability masks. Only observations and analyses at or before
119
- issue time may enter inference; retrospective data availability is a separate
120
- operational question.
121
-
122
- The internal core is a $65\times65$ grid at 20-km sampling over $\pm640$ km.
123
- Its field is a learned reconstruction, not native 20-km observations.
124
- For each of twenty leads ($6,12,\ldots,120$ h), outputs include basin fields,
125
- regional pressure, core pressure/coordinates, storm centre, central pressure,
126
- domain masks and auxiliary maximum wind. This implementation does not output
127
- quadrant wind radii or a resolved surface-wind map.
128
-
129
- \section{Neural modules}
130
- The released non-tiny architecture has 21,452,595 trainable parameters.
131
  \subsection{Multiresolution evolution networks}
132
  The basin network receives 44 channels: eight fields, 32 memory features and
133
  four static channels. Encoder widths are $72,144,288,432$; residual-block depths
@@ -169,7 +206,7 @@ retrieves context; a zero-initialized $64\to32$ projection adds it to basin
169
  memory before the transition. No separate positional embedding is added:
170
  geographic coordinates are already among the token inputs.
171
 
172
- \section{Observation-conditioned initialization}
173
  Write $\sample(F,\xi)$ for bilinear field sampling. Sampling uses FP32 geometry,
174
  aligned corners and border padding for the basin; transported core anomalies
175
  use zero padding. All nine pressure histories are sampled onto the core grid
@@ -188,77 +225,73 @@ $T=\operatorname{clip}((640-\sqrt{x^2+y^2})/160,0,1)$.
188
  The anomaly is $a_0=c_0-\sample(g_{0,\mathrm{MSLP}},\xi_0)$.
189
  This compact prior is applied only at initialization. Future scalar pressure
190
  is not repeatedly inserted into the field.
191
-
192
- \section{One six-hour transition}
193
- \subsection{Basin transport}
194
- Learned softmax weights blend physical winds at 850, 500 and 200 hPa.
195
- Their initialization is $(0.4,0.4,0.2)$, not a fixed inference constraint.
196
- For network output $d$,
197
- \[
198
- U_k=\sum_\ell\alpha_\ell(u_\ell,v_\ell)_k+15\tanh(d_{8:10})
199
- \quad\mathrm{m\,s^{-1}}.
200
- \]
201
- Backtracing constructs departure coordinates $D(U_k)$ and
202
- \[
203
- g_{k+1}=\sample(g_k,D(U_k))+0.2\tanh(d_{0:8}).
204
- \]
205
- Six hours convert 1 m/s to 21.6 km. A latitude degree is treated as 111.2 km.
206
- A longitude degree is treated as $111.2\max(\cos\phi,0.2)$ km.
207
- The attention-conditioned memory is advected by the same departure grid and
208
- updated with new features. Tendencies are in normalized channel units.
209
- This semi-Lagrangian-style update is not a complete or exactly conservative
210
- atmospheric equation solver.
211
 
212
  \subsection{Moving-frame proposal}
213
- A $44\to64\to3$ GELU MLP reads sampled 32-channel memory, eight basin fields,
214
  previous motion divided by 100, and two issue-intensity masks. For output $b$,
215
- \[
216
  m_{k+1}=0.7(21.6)U_k(z_k)+0.3m_k+100\tanh(b_{0:2}).
217
- \]
218
- This displacement, in km per six hours, defines a new origin relative to
219
- $z_k$. It proposes where the pressure core moves, but is not the final reported
220
- centre. The third output updates auxiliary wind:
221
  $v_{k+1}=\max(0,v_k+5b_2)$.
222
 
223
- \subsection{Relative core transport}
224
- The core network reads the previous pressure core, its old sampled environment,
225
- core memory and geography at the proposed origin. Local velocity is
226
  $U^{\rm local}=\sample(U_k)+5\tanh(d^c_{1:3})$.
227
- Subtracting $m_{k+1}/21.6$ removes bulk frame translation. Backward coordinates
228
- also account for the offset between the previous associated centre and its
229
- previous grid origin. On this relative departure grid $D_c$,
230
- \begin{align*}
231
  a_{k+1}&=\sample(a_k,D_c)+0.1d^c_0,\\
232
  c_{k+1}&=\sample(g_{k+1,\mathrm{MSLP}},\xi_{k+1})+Ta_{k+1}.
233
- \end{align*}
234
- Unlike the basin tendency, $d^c_0$ is not bounded by $\tanh$.
235
- Core memory uses the same relative sampling followed by its gated update.
236
-
237
- \subsection{Centre association and central pressure}
238
- Convert core pressure to hPa: $P=\sigma_p c_{k+1}+\mu_p$.
239
- For cells within 300 km of the proposed origin,
240
- \[
241
- w_i=\operatorname{softmax}_i\left(-P_i/2-r_i^2/(2\,180^2)\right).
242
- \]
243
- Cells beyond 300 km are excluded. The weighted mean of cell latitude/longitude
244
- is $z_{k+1}$. Central pressure is the bilinear sample $\sample(P,z_{k+1})$,
245
- not necessarily the grid minimum. The spatial prior associates the forecast
246
- with this storm rather than an unrelated remote depression.
247
-
248
- \subsection{Geographic pressure output}
249
- The issue-anchored regional output is
250
- \[
251
  \widehat p_R=\sample(g_{k+1,\mathrm{MSLP}},R)
252
- +\sample(Ta_{k+1},R;\mathrm{zero\ padding}),
253
- \]
254
- then de-normalized to hPa. This basin-plus-core reconstruction is generated by
255
- the model itself; future observed fields and display-only vortices are not
256
- substituted. Its $0.25^\circ$ sampling is not proof of native resolution.
257
- Masks indicate valid basin coverage. Resampling and localized association mean
258
- the reported centre, central pressure and regional map minimum need not coincide
259
- exactly. A display must preserve those distinctions rather than relocate fields.
260
-
261
- \section{Autoregression and learning objectives}
262
  Inference initializes once and repeats the transition twenty times, using only
263
  fixed issue context and its own predicted future state. The implemented
264
  single-step loss is
@@ -281,7 +314,7 @@ together, without adding an independent central-pressure output head.
281
  Training-only normalization, storm-level split boundaries and causal input
282
  selection remain essential implementation constraints.
283
 
284
- \section{Single forecasts and 50-member means}
285
  Evaluation-mode inference is deterministic for fixed inputs and weights.
286
  The documented ensemble perturbs only normalized historical weather: independent
287
  Gaussian fields are spatially smoothed with kernels five (basin) and nine
@@ -301,38 +334,37 @@ These are perturbed-input realizations of one model, not 50 trained networks or
301
  learned latent samples. Spread is not automatically calibrated uncertainty.
302
  Member counts, perturbation settings and seed policy must accompany the outputs.
303
 
304
- \section{Source identity and use boundary}
305
- The public name is \textbf{Trackformer 1.2}. Its manifest retains the internal
306
- checkpoint identity, 1.2.73 epoch 4, and exact weight/source hashes.
307
- \texttt{model.py} implements environmental attention and added temporal losses;
308
- \texttt{baseline\_model.py} implements moving-core evolution and readout;
309
- \texttt{v165\_base.py} supplies sampling, evolution blocks and memories.
310
- All are in \texttt{models/trackformer\_1\_2\_field/}.
311
-
312
- Learned reconstructions, limited domain, input-source differences and retrospective
313
- availability constrain interpretation. The model remains a research artifact,
314
- not an operational warning system.
315
-
316
-
 
 
 
 
 
 
 
 
317
  \section{Development evaluation}
318
  \label{sec:evaluation}
319
- \subsection{Cohort and comparison boundary}
320
- The completed comparison comprises \textbf{270 issue cases from 90 storms},
321
- with three issues per storm and twenty six-hour leads per case. The 1.2 forecast
322
- is the mean of 50 causal historical-input perturbations; 1.1 uses its archived
323
- causal route pipeline. Cases, labels and forecast leads are matched, but the
324
- input pipelines differ. This is not a controlled architecture ablation.
325
- The cohort was repeatedly inspected and includes cases outside the intended
326
- Western Pacific domain. It is therefore development evidence, not an untouched
327
- or basin-specific validation claim.
328
-
329
- The original 1.1 archive stored absolute latitude/longitude under a local-km
330
- array name. The reproduced comparison converts those coordinates into the same
331
- issue-centred local projection used by 1.2 and checks a round trip to the
332
- original coordinates. Its corrected 902.3-km mean replaces the superseded
333
- 1,036.3-km score; original forecasts are preserved.
334
-
335
- \begin{figure}[htbp]
336
  \centering
337
  \begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
338
  \begin{scope}[shift={(0,0)}]
@@ -340,15 +372,15 @@ original coordinates. Its corrected 902.3-km mean replaces the superseded
340
  \node[rotate=90] at (-.75,1.75) {km; lower is better};
341
  \draw[black!10] (0,0.0000)--(5,0.0000);
342
  \node[left] at (0,0.0000) {0};
343
- \draw[black!10] (0,1.5909)--(5,1.5909);
344
- \node[left] at (0,1.5909) {500};
345
- \draw[black!10] (0,3.1818)--(5,3.1818);
346
- \node[left] at (0,3.1818) {1000};
347
- \fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.8710);
348
- \node[above] at (1.4,2.8710) {902.3};
349
  \node[below=3pt] at (1.4,0) {1.1};
350
- \fill[modelteal] (3.22,0) rectangle (3.98,2.2731);
351
- \node[above] at (3.6,2.2731) {714.4};
352
  \node[below=3pt] at (3.6,0) {1.2 mean of 50};
353
  \draw[black!55] (0,3.5)--(0,0)--(5,0);
354
  \end{scope}
@@ -363,11 +395,11 @@ original coordinates. Its corrected 902.3-km mean replaces the superseded
363
  \node[left] at (0,2.0000) {40};
364
  \draw[black!10] (0,3.0000)--(5,3.0000);
365
  \node[left] at (0,3.0000) {60};
366
- \fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.8110);
367
- \node[above] at (1.4,2.8110) {56.22};
368
  \node[below=3pt] at (1.4,0) {1.1};
369
- \fill[modelteal] (3.22,0) rectangle (3.98,2.2410);
370
- \node[above] at (3.6,2.2410) {44.82};
371
  \node[below=3pt] at (3.6,0) {1.2 mean of 50};
372
  \draw[black!55] (0,3.5)--(0,0)--(5,0);
373
  \end{scope}
@@ -401,49 +433,45 @@ original coordinates. Its corrected 902.3-km mean replaces the superseded
401
  \node[font=\small,rotate=90] at (-1.05,1.850) {Mean position error (km)};
402
  \begin{scope}
403
  \clip (0,0) rectangle (12.2,3.7);
404
- \draw[baselinegray,dashed,thick] (0.0000,0.1347)--(0.6421,0.2659)--(1.2842,0.4048)--(1.9263,0.5515)--(2.5684,0.7090)--(3.2105,0.8669)--(3.8526,1.0296)--(4.4947,1.1938)--(5.1368,1.3637)--(5.7789,1.5348)--(6.4211,1.7135)--(7.0632,1.8929)--(7.7053,2.0728)--(8.3474,2.2493)--(8.9895,2.4266)--(9.6316,2.6085)--(10.2737,2.7982)--(10.9158,2.9886)--(11.5579,3.1884)--(12.2000,3.3915);
405
- \draw[modelteal,solid,thick] (0.0000,0.1373)--(0.6421,0.2516)--(1.2842,0.3709)--(1.9263,0.4892)--(2.5684,0.6140)--(3.2105,0.7368)--(3.8526,0.8550)--(4.4947,0.9752)--(5.1368,1.0968)--(5.7789,1.2152)--(6.4211,1.3405)--(7.0632,1.4703)--(7.7053,1.6031)--(8.3474,1.7378)--(8.9895,1.8779)--(9.6316,2.0217)--(10.2737,2.1710)--(10.9158,2.3276)--(11.5579,2.4906)--(12.2000,2.6518);
406
  \end{scope}
407
  \draw[baselinegray,dashed,thick] (0.250,-1.23)--(0.800,-1.23);
408
  \node[anchor=west] at (0.900,-1.23) {Trackformer 1.1};
409
  \draw[modelteal,solid,thick] (6.350,-1.23)--(6.900,-1.23);
410
  \node[anchor=west] at (7.000,-1.23) {Trackformer 1.2: mean of 50};
411
  \end{tikzpicture}
412
- \caption{Matched development comparison. Top: all-lead position and heading
413
- errors. Bottom: mean position error at each lead. Distances use the saved
414
- local-km projection. Heading uses 5,382 common valid steps where truth and
415
- both forecasts move more than 1 km. All 270 cases remain in the position
416
- scores, including poor and out-of-domain cases. No uncertainty intervals are
417
- available in these saved results.}
418
  \label{fig:benchmark}
419
  \end{figure}
420
-
421
- \subsection{Position, direction and shape are different quantities}
422
- Mean local-coordinate position error improves by 20.8\%, from 902.3 to
423
- 714.4 km; +120 h error is 1,833.2 versus 1,433.4 km. The +6 h error is
424
- \emph{slightly worse} for 1.2 (74.2 versus 72.8 km), so the aggregate is not
425
- an improvement at every lead. Mean heading error is 56.22 versus 44.82 degrees.
426
- Centred route-shape similarity is 0.7273 versus 0.8249; normalized
427
- Fr\'echet path similarity is 0.4983 versus 0.5727. These metrics are not
428
- accuracy percentages and are not combined into a weighted rank.
429
-
430
- \textbf{Actual route alignment is the primary geographic criterion:}
431
- predicted and observed positions should coincide at the same valid times.
432
- The shape score removes translation and scale, so it can remain high for a
433
- geographically displaced forecast. It must not substitute for position error
434
- or an unshifted route overlay. The path score responds to displacement but
435
- is not a strict same-time distance.
436
-
437
- \clearpage
438
- \section{Selected Fung-wong case: alignment and intensity}
439
- The Fung-wong forecast is issued at 00 UTC on 7 November 2025, using the same
440
- saved 50-member policy. Figure~\ref{fig:fungwong} overlays predictions and exact
441
- six-hour observed coordinates without translating, rotating or scaling either
442
- route. Mean great-circle error is 130.9 km, +120 h error is 142.3 km and the
443
- maximum is 198.3 km: all twenty leads are within 200 km. These geographic
444
- errors are a separate diagnostic from the preceding local-coordinate scores.
445
-
446
- \begin{figure}[htbp]
447
  \centering
448
  \begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
449
  \node[font=\small\bfseries,align=center] at (6.100,7.800) {Fung-wong: geographic routes, +0 to +120 h};
@@ -569,10 +597,15 @@ The companion MP4 shows the actual common-grid model pressure fields and
569
  4-hPa isobars; none of the pressure fields is moved to match observations.}
570
  \label{fig:fungwong}
571
  \end{figure}
572
-
573
- \begin{figure}[htbp]
 
 
 
 
 
574
  \centering
575
- \begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
576
  \node[font=\small\bfseries,align=center] at (5.750,8.900) {Fung-wong: model-mean MSLP at +48 h};
577
  \draw[black!10] (0.676,0)--(0.676,8.5);
578
  \node[below=3pt] at (0.676,0) {124};
@@ -639,59 +672,36 @@ The companion MP4 shows the actual common-grid model pressure fields and
639
  \draw[routepink,solid,thick] (6.000,-1.23)--(6.550,-1.23);
640
  \node[anchor=west] at (6.650,-1.23) {Forecast route to +48 h};
641
  \end{tikzpicture}
642
- \caption{Unmodified saved physical ensemble-mean regional MSLP at +48 h
643
- (9 November 2025, 00 UTC), contoured every 4 hPa and labelled every 8 hPa.
644
- The source grid is sampled at $0.25^\circ$; this is not a claim of native
645
- resolution. Lines overlay the observed and forecast paths only through this
646
- valid time. The figure uses the portion inside the fixed regional patch;
647
- its western edge is $123^\circ$E. The map has not been shifted to match the
648
- observed centre, and its minimum need not equal member-mean central pressure.}
649
  \label{fig:pressuremap}
650
  \end{figure}
651
-
652
- \clearpage
653
- \subsection{Pressure and coverage limitations}
654
- Observed pressure is JMA best track as stored in IBTrACS \texttt{TOKYO\_PRES},
655
- not a JMA forecast. Native-detail input history was unavailable for this issue;
656
- the regional pressure field is a learned basin-plus-core reconstruction
657
- sampled every $0.25^\circ$, not a native-resolution forecast. Coverage masks
658
- and fixed regional boundaries must be retained. From +66 h the forecast
659
- centre leaves the fixed regional patch; the video then shows only the saved
660
- coarse basin field at that location. The moving-core readout still supplies
661
- central pressure; no detailed map core is invented outside coverage.
662
- The mean route and mean
663
- central pressure need not be the minimum of the common-grid mean field.
664
-
665
- No verified same-270 1.1 pressure-prediction array is available in the route
666
- artifacts, so a pressure-comparison bar is not invented. On the separate,
667
- repeatedly inspected TIP ten-start diagnostic, central-pressure MAE is worse
668
- for 1.2 (30.7 versus 20.1 hPa). Overlapping issues of one storm are not ten
669
- independent tests. The route improvement cannot be read as an all-metric win.
670
-
671
- \section{Reproducibility and limitations}
672
- The figures above embed numeric values from the published saved artifacts,
673
- so this source compiles without external images. Model code, checkpoint
674
- provenance, plot data, selection rules and reproduction utilities are available
675
- in the release repository\footnote{\url{https://github.com/yu314-coder/typhoon-predict}}.
676
- The 270-case scores are in \path{evaluation/trackformer_1_2_vs_1_1_270_metrics.json};
677
- Fung-wong coordinates, pressure, source hashes and member policy are in
678
- \path{evaluation/release_data/fung_wong_video.npz} and its JSON manifest.
679
- The \href{https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4}{Fung-wong MP4}
680
- is an illustration of saved predictions, not a new inference experiment.
681
-
682
- A separate expanded protocol freezes 1,473 daily issues from 270 distinct
683
- storms: average leads within each issue, issues within each storm, then storms
684
- equally. It is not the completed 270-case comparison reported here; its
685
- partial results are not used in these figures. Stronger generalization claims
686
- require a genuinely unused whole-storm holdout, fitting-only preprocessing,
687
- matched data availability and lead masks, and uncertainty estimates based on
688
- resampling whole storms rather than overlapping issues.
689
-
690
- \section{Conclusion}
691
- Trackformer 1.2 couples recurrent environmental evolution with a translating
692
- pressure anomaly and pressure-associated readout. Development results improve
693
- aggregate route measures relative to the archived 1.1 pipeline, while leaving
694
- large absolute errors, short-lead regressions and intensity failures. The
695
- architecture is a research model, not a substitute for official warnings;
696
- selected visually aligned routes do not establish operational reliability.
697
  \end{document}
 
1
+ \documentclass[10pt,letterpaper]{article}
2
+ \usepackage[left=0.82in,right=0.82in,top=0.78in,bottom=0.78in,headsep=0.22in]{geometry}
3
+ \usepackage[T1]{fontenc}
4
+ \usepackage{lmodern,amsmath,amssymb,booktabs,array,tabularx,microtype,xcolor}
5
+ \usepackage{tikz,caption,fancyhdr,float}
6
  \usetikzlibrary{arrows.meta,positioning,calc}
7
+ \definecolor{modelteal}{HTML}{007D78}
8
  \definecolor{baselinegray}{HTML}{65758B}
9
  \definecolor{ink}{HTML}{182C42}
10
  \definecolor{routepink}{HTML}{B31565}
11
+ \usepackage[colorlinks=true,linkcolor=modelteal,urlcolor=modelteal]{hyperref}
12
+ \hypersetup{pdftitle={Trackformer 1.2: Coupled Pressure-Field Forecasting},pdfauthor={Yu Yao-Hsing},pdfsubject={Architecture, forecast dynamics and development evaluation}}
13
  \pagestyle{fancy}
14
  \fancyhf{}
15
+ \fancyhead[L]{\sffamily\footnotesize\color{ink}TRACKFORMER / 1.2}
16
+ \fancyhead[R]{\sffamily\footnotesize\color{baselinegray}RESEARCH TECHNICAL REPORT}
17
+ \fancyfoot[L]{\sffamily\scriptsize\color{baselinegray}Architecture and development evaluation}
18
+ \fancyfoot[R]{\sffamily\small\thepage}
19
+ \renewcommand{\headrulewidth}{0.3pt}
20
  \setlength{\headheight}{14pt}
21
+ \captionsetup{font={small},labelfont={bf,sf,color=ink},labelsep=period,justification=justified,singlelinecheck=false,skip=7pt}
 
 
22
  \setlength{\parindent}{0pt}
23
+ \setlength{\parskip}{4pt plus 1pt}
24
+ \setlength{\intextsep}{10pt}
25
+ \setlength{\abovedisplayskip}{7pt}
26
+ \setlength{\belowdisplayskip}{7pt}
27
+ \setlength{\emergencystretch}{1.5em}
28
+ \renewcommand{\arraystretch}{1.18}
29
+ \widowpenalty=10000
30
+ \clubpenalty=10000
31
+ \raggedbottom
32
+ \makeatletter
33
+ \renewcommand\section{\@startsection{section}{1}{0pt}{12pt}{7pt}{\sffamily\bfseries\color{ink}\fontsize{14}{17}\selectfont}}
34
+ \renewcommand\subsection{\@startsection{subsection}{2}{0pt}{9pt}{4pt}{\sffamily\bfseries\color{ink}\fontsize{10.5}{13}\selectfont}}
35
+ \makeatother
36
  \newcommand{\sample}{\mathcal{S}}
37
+ \newcommand{\reportpage}{\clearpage}
38
+ \newcommand{\keyline}[2]{\textbf{#1}\quad #2\par}
 
39
  \begin{document}
40
+ \fontsize{10.3}{13}\selectfont
41
+ \thispagestyle{empty}
42
+ {\sffamily\small\color{modelteal}RESEARCH TECHNICAL REPORT\hfill 29 SEPTEMBER 2026\par}
43
+ \vspace{13pt}
44
+ {\sffamily\bfseries\color{ink}\fontsize{26}{29}\selectfont Trackformer 1.2\par}
45
+ \vspace{4pt}
46
+ {\sffamily\color{ink}\fontsize{17}{21}\selectfont Coupled pressure-field forecasting\par}
47
+ \vspace{6pt}
48
+ {\large Architecture, forecast dynamics and development evaluation\par}
49
+ \vspace{10pt}
50
+ {\sffamily\small Yu Yao-Hsing\hfill\href{https://github.com/yu314-coder/typhoon-predict}{Code, data and model release}\par}
51
+ \vspace{8pt}
52
+ {\color{modelteal}\hrule height 1pt}
53
+ \vspace{10pt}
54
+ {\sffamily\bfseries\color{ink}Abstract\par}
55
+ Trackformer 1.2 forecasts tropical-cyclone motion and central pressure by
56
+ jointly advancing a basin-scale atmospheric state and a translating pressure
57
+ core. Recurrent convolutional networks evolve the fields, while multiscale
58
+ attention supplies environmental context. A pressure-associated centre and
59
+ central-pressure sample are read from the evolving core at each six-hour
60
+ step. Twenty autoregressive transitions produce a five-day forecast.
61
+ This report specifies the released architecture, state updates, learning
62
+ objectives and 50-member input-perturbation policy. Across 1,473 daily forecasts
63
+ from 270 storms, equal-storm mean track error is 471.2 km for the 1.2 ensemble
64
+ and 798.4 km for the causal 1.1 route pipeline.
65
+ The pipelines differ; this is not an architecture ablation or an untouched-test
66
+ result. A selected Fung-wong case illustrates actual geographic alignment,
67
+ intensity evolution and model-generated pressure contours, including their
68
+ remaining errors and coverage limits.
69
+
70
+ \vspace{5pt}
71
+ \begin{tabularx}{\linewidth}{@{}XXX@{}}
72
+ \toprule
73
+ \sffamily\bfseries Forecast horizon & \sffamily\bfseries Representation & \sffamily\bfseries Evaluation policy\\
74
+ +6 to +120 h & Basin + moving pressure core & 50 input realizations\\
75
+ 20 six-hour transitions & 21.45 million parameters & One shared set of weights\\
76
+ \bottomrule
77
+ \end{tabularx}
78
+
79
+ \section{Overview and forecast state}
80
+ The defining constraint is a coupled readout: the route and central pressure
81
+ must follow the pressure core generated by the model. A learned,
82
+ steering-conditioned proposal moves the core's reference frame; a local
83
+ pressure association then determines the reported centre. The method is
84
+ therefore neither a pressure-only tracker nor a complete numerical atmospheric
85
+ solver. Its displayed pressure is a forecast field, not an independently drawn
86
+ vortex placed along a predicted route.
87
 
88
  At step $k$, the recurrent state is
89
  \[
90
  s_k=(g_k,h^g_k,c_k,h^c_k,a_k,o_k,z_k,m_k,v_k;G,R,I,M).
91
  \]
92
+ \begin{tabularx}{\linewidth}{@{}lX@{}}
93
+ \toprule
94
+ \textbf{State} & \textbf{Meaning}\\
95
+ $g,h^g$ & Eight normalized basin fields and basin memory.\\
96
+ $c,h^c,a$ & Core pressure, core memory and pressure anomaly relative to the basin.\\
97
+ $o,z,m,v$ & Core-grid origin, associated centre, six-hour motion and auxiliary wind.\\
98
+ $G,R,I,M$ & Fixed geography, regional coordinates, issue intensity and validity masks.\\
99
+ \bottomrule
100
+ \end{tabularx}
101
+
102
+ \vspace{7pt}
103
+ \keyline{Scope.}{A research model for the Western Pacific, not an operational
104
+ warning system. Selected examples do not establish generalization or safety.}
105
+ \reportpage
106
+ \section{Architecture and input contract}
107
+ The basin and moving core share a recurrent forecast cycle, but retain separate
108
+ spatial representations and memories. Figure~\ref{fig:architecture} shows the
109
+ dependency structure; the operator definitions follow in Section~\ref{sec:operators}.
110
+ \begin{figure}[H]
111
  \centering
112
  \begin{tikzpicture}[font=\small,>=Latex,
113
  box/.style={draw=ink!50,rounded corners=3pt,align=center,text width=4.6cm,minimum height=1.15cm,inner sep=6pt,fill=ink!3},
 
135
  \node[below=0.25cm of readout,align=center,font=\scriptsize,text=ink] {Shared transition repeated 20 times: +6, +12, $\ldots$, +120 h.\\
136
  50 perturbed input histories $\rightarrow$ 50 rollouts $\rightarrow$ separate route, pressure and common-grid field means.};
137
  \end{tikzpicture}
138
+ \caption{Coupled forecast architecture. Basin evolution supplies environmental
139
+ fields to the moving core; the pressure-associated readout follows the motion
140
+ proposal. All leads and ensemble members use the same trained network.}
 
 
141
  \label{fig:architecture}
142
  \end{figure}
143
 
144
+ \subsection{Causal inputs and spatial representations}
145
+ \begin{tabularx}{\linewidth}{@{}p{3.1cm}p{3.6cm}X@{}}
146
+ \toprule
147
+ \textbf{Input / state} & \textbf{Shape / sampling} & \textbf{Role and boundary}\\
148
+ \midrule
149
+ Basin history & $B\times9\times8\times25\times33$ & Nine analyses at $-48,-42,\ldots,0$ h; $0$--$60^\circ$N, $100$--$180^\circ$E, $2.5^\circ$ spacing.\\
150
+ Regional history & $B\times9\times1\times121\times121$ & Issue-anchored MSLP patch, $\pm15^\circ$, sampled every $0.25^\circ$; availability mask selects native or coarse initialization.\\
151
+ Moving core & $65\times65$; 20-km spacing & Learned pressure reconstruction over $\pm640$ km; grid spacing does not establish effective resolution.\\
152
+ Issue state & Centre, motion, intensity, masks & Latitude/longitude, six-hour east/north displacement, wind in knots, pressure in hPa.\\
153
+ Static channels & Four channels per grid & Latitude/90, longitude/180$-1$, land fraction and elevation/8000 m.\\
154
+ \bottomrule
155
+ \end{tabularx}
156
+
157
+ Basin channels are $[\mathrm{MSLP},Z_{500},u_{850},v_{850},u_{500},v_{500},u_{200},v_{200}]$
158
+ in hPa, metres and metres per second. Normalization uses fixed training-only
159
+ statistics, $\widetilde x_j=(x_j-\mu_j)/\sigma_j$. Regional coordinates remain
160
+ available when native history is missing. Only information available at or
161
+ before issue time may enter inference; retrospective archive coverage does
162
+ not establish operational availability. Outputs include basin and regional
163
+ pressure, core coordinates, centre, central pressure, masks and auxiliary wind;
164
+ there are no quadrant wind radii or resolved surface-wind maps.
165
+ \reportpage
166
+ \section{Neural operators}\label{sec:operators}
167
+ The released model has 21,452,595 trainable parameters.
 
 
 
 
 
 
 
 
 
168
  \subsection{Multiresolution evolution networks}
169
  The basin network receives 44 channels: eight fields, 32 memory features and
170
  four static channels. Encoder widths are $72,144,288,432$; residual-block depths
 
206
  memory before the transition. No separate positional embedding is added:
207
  geographic coordinates are already among the token inputs.
208
 
209
+ \subsection{Observation-conditioned initialization}
210
  Write $\sample(F,\xi)$ for bilinear field sampling. Sampling uses FP32 geometry,
211
  aligned corners and border padding for the basin; transported core anomalies
212
  use zero padding. All nine pressure histories are sampled onto the core grid
 
225
  The anomaly is $a_0=c_0-\sample(g_{0,\mathrm{MSLP}},\xi_0)$.
226
  This compact prior is applied only at initialization. Future scalar pressure
227
  is not repeatedly inserted into the field.
228
+ \reportpage
229
+ \section{One six-hour forecast transition}
230
+ Each step advances the basin first, proposes core motion, transports the core
231
+ relative to that motion, and reads out position and pressure. All twenty leads
232
+ share the same operators. Sampling $\sample(F,\xi)$ is bilinear; coordinates
233
+ and units are explicit below.
234
+
235
+ \subsection{Environmental transport}
236
+ Learned softmax weights blend 850-, 500- and 200-hPa winds; their initial values
237
+ are $(0.4,0.4,0.2)$. For basin-network output $d$, the corrected flow and next
238
+ normalized basin field are
239
+ \begin{align}
240
+ U_k &= \sum_\ell\alpha_\ell(u_\ell,v_\ell)_k+15\tanh(d_{8:10}),\label{eq:flow}\\
241
+ g_{k+1} &= \sample(g_k,D(U_k))+0.2\tanh(d_{0:8}).\label{eq:basin}
242
+ \end{align}
243
+ $U_k$ is in m\,s$^{-1}$ and $D$ is the backtraced departure grid. Six hours
244
+ convert 1 m\,s$^{-1}$ to 21.6 km. Coordinate conversion uses 111.2 km per
245
+ latitude degree and $111.2\max(\cos\phi,0.2)$ km per longitude degree.
246
+ The attention-conditioned memory follows the same advection and then a gated
247
+ update. This transport is not an exactly conservative atmospheric solver.
248
 
249
  \subsection{Moving-frame proposal}
250
+ A $44\to64\to3$ GELU MLP reads 32 sampled memory channels, eight basin fields,
251
  previous motion divided by 100, and two issue-intensity masks. For output $b$,
252
+ \begin{equation}
253
  m_{k+1}=0.7(21.6)U_k(z_k)+0.3m_k+100\tanh(b_{0:2}).
254
+ \end{equation}
255
+ This km-per-six-hour displacement defines the proposed origin relative to
256
+ $z_k$, not the final reported centre. Auxiliary wind becomes
 
257
  $v_{k+1}=\max(0,v_k+5b_2)$.
258
 
259
+ \subsection{Core transport in the moving frame}
260
+ The core network reads previous core pressure, the old sampled environment,
261
+ memory and geography at the proposed origin. Its local velocity is
262
  $U^{\rm local}=\sample(U_k)+5\tanh(d^c_{1:3})$.
263
+ Subtracting $m_{k+1}/21.6$ removes bulk frame motion. The relative departure
264
+ grid $D_c$ also accounts for the previous centre--origin offset:
265
+ \begin{align}
 
266
  a_{k+1}&=\sample(a_k,D_c)+0.1d^c_0,\\
267
  c_{k+1}&=\sample(g_{k+1,\mathrm{MSLP}},\xi_{k+1})+Ta_{k+1}.
268
+ \end{align}
269
+ The anomaly tendency $d^c_0$ is unbounded by $\tanh$. Core memory uses the same
270
+ relative sampling followed by its gated update.
271
+
272
+ \subsection{Pressure-associated readout}
273
+ Convert core pressure to hPa, $P=\sigma_p c_{k+1}+\mu_p$. For cells no more
274
+ than 300 km from the proposed origin, define
275
+ \begin{equation}
276
+ w_i=\operatorname{softmax}_i\!\left(-P_i/2-r_i^2/(2\,180^2)\right).
277
+ \end{equation}
278
+ The weighted mean of cell latitude/longitude gives $z_{k+1}$; central pressure
279
+ is $\sample(P,z_{k+1})$, not necessarily the grid minimum. Cells outside the
280
+ radius are excluded, reducing association with unrelated depressions.
281
+
282
+ \subsection{Geographic pressure reconstruction}
283
+ The issue-anchored regional field is
284
+ \begin{equation}
 
285
  \widehat p_R=\sample(g_{k+1,\mathrm{MSLP}},R)
286
+ +\sample(Ta_{k+1},R;\mathrm{zero\ padding}),
287
+ \end{equation}
288
+ then de-normalized to hPa. This is the model's basin-plus-core field; future
289
+ observations or display-only vortices are never substituted. Coverage masks
290
+ must remain visible. Resampling and local centre association mean the route,
291
+ central pressure and displayed field minimum need not coincide exactly.
292
+ \reportpage
293
+ \section{Training, ensembles and reproducibility}
294
+ \subsection{Autoregression and learning objectives}
 
295
  Inference initializes once and repeats the transition twenty times, using only
296
  fixed issue context and its own predicted future state. The implemented
297
  single-step loss is
 
314
  Training-only normalization, storm-level split boundaries and causal input
315
  selection remain essential implementation constraints.
316
 
317
+ \subsection{Single forecasts and 50-member means}
318
  Evaluation-mode inference is deterministic for fixed inputs and weights.
319
  The documented ensemble perturbs only normalized historical weather: independent
320
  Gaussian fields are spatially smoothed with kernels five (basin) and nine
 
334
  learned latent samples. Spread is not automatically calibrated uncertainty.
335
  Member counts, perturbation settings and seed policy must accompany the outputs.
336
 
337
+ \subsection{Release identity and reproducibility}
338
+ The public model is \textbf{Trackformer 1.2}; the manifest retains internal
339
+ checkpoint identity 1.2.73 epoch 4 and exact source/weight hashes. The following
340
+ artifacts are distributed in the
341
+ \href{https://github.com/yu314-coder/typhoon-predict}{release repository}:
342
+ \begin{tabularx}{\linewidth}{@{}p{4.2cm}X@{}}
343
+ \toprule
344
+ \textbf{Artifact} & \textbf{Purpose}\\
345
+ \midrule
346
+ \texttt{model.py} & Environmental attention and additional temporal losses.\\
347
+ \texttt{baseline\_model.py} & Moving-core evolution and pressure-associated readout.\\
348
+ \texttt{v165\_base.py} & Sampling, multiscale evolution blocks and recurrent memory.\\
349
+ \texttt{manifest.json} & Released grids, normalization and checkpoint/source identity.\\
350
+ Saved evaluation arrays & 270-storm final metrics; Fung-wong route, fields and member provenance.\\
351
+ \bottomrule
352
+ \end{tabularx}
353
+ Model sources are in \path{models/trackformer_1_2_field/}; evaluation arrays
354
+ are in \path{evaluation/}. This document embeds figure coordinates and needs
355
+ no external images. The released plotting utility regenerates the figure
356
+ data from saved arrays; it does not perform new inference.
357
+ \reportpage
358
  \section{Development evaluation}
359
  \label{sec:evaluation}
360
+ \subsection{Matched cases, different forecast pipelines}
361
+ The completed comparison uses \textbf{1,473 daily issues from 270 storms},
362
+ with 20 six-hour leads per issue. Scores average leads within daily issues,
363
+ issues within storms, then storms equally. The cohort was frozen before inference.
364
+ Trackformer 1.2 averages 50 causal input perturbations; 1.1 uses its causal
365
+ route pipeline. Cases and targets match, but inputs differ. This is broader
366
+ development evidence, not an untouched test or architecture ablation.
367
+ \begin{figure}[H]
 
 
 
 
 
 
 
 
 
368
  \centering
369
  \begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
370
  \begin{scope}[shift={(0,0)}]
 
372
  \node[rotate=90] at (-.75,1.75) {km; lower is better};
373
  \draw[black!10] (0,0.0000)--(5,0.0000);
374
  \node[left] at (0,0.0000) {0};
375
+ \draw[black!10] (0,1.7500)--(5,1.7500);
376
+ \node[left] at (0,1.7500) {500};
377
+ \draw[black!10] (0,3.5000)--(5,3.5000);
378
+ \node[left] at (0,3.5000) {1000};
379
+ \fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.7944);
380
+ \node[above] at (1.4,2.7944) {798.4};
381
  \node[below=3pt] at (1.4,0) {1.1};
382
+ \fill[modelteal] (3.22,0) rectangle (3.98,1.6492);
383
+ \node[above] at (3.6,1.6492) {471.2};
384
  \node[below=3pt] at (3.6,0) {1.2 mean of 50};
385
  \draw[black!55] (0,3.5)--(0,0)--(5,0);
386
  \end{scope}
 
395
  \node[left] at (0,2.0000) {40};
396
  \draw[black!10] (0,3.0000)--(5,3.0000);
397
  \node[left] at (0,3.0000) {60};
398
+ \fill[baselinegray] (1.02,0) rectangle (1.7799999999999998,2.5790);
399
+ \node[above] at (1.4,2.5790) {51.58};
400
  \node[below=3pt] at (1.4,0) {1.1};
401
+ \fill[modelteal] (3.22,0) rectangle (3.98,1.7480);
402
+ \node[above] at (3.6,1.7480) {34.96};
403
  \node[below=3pt] at (3.6,0) {1.2 mean of 50};
404
  \draw[black!55] (0,3.5)--(0,0)--(5,0);
405
  \end{scope}
 
433
  \node[font=\small,rotate=90] at (-1.05,1.850) {Mean position error (km)};
434
  \begin{scope}
435
  \clip (0,0) rectangle (12.2,3.7);
436
+ \draw[baselinegray,dashed,thick] (0.0000,0.1223)--(0.6421,0.2469)--(1.2842,0.3744)--(1.9263,0.5056)--(2.5684,0.6418)--(3.2105,0.7768)--(3.8526,0.9131)--(4.4947,1.0510)--(5.1368,1.1916)--(5.7789,1.3369)--(6.4211,1.4867)--(7.0632,1.6402)--(7.7053,1.8006)--(8.3474,1.9621)--(8.9895,2.1304)--(9.6316,2.3032)--(10.2737,2.4856)--(10.9158,2.6699)--(11.5579,2.8565)--(12.2000,3.0458);
437
+ \draw[modelteal,solid,thick] (0.0000,0.0893)--(0.6421,0.1686)--(1.2842,0.2389)--(1.9263,0.3043)--(2.5684,0.3683)--(3.2105,0.4346)--(3.8526,0.5041)--(4.4947,0.5762)--(5.1368,0.6539)--(5.7789,0.7389)--(6.4211,0.8312)--(7.0632,0.9296)--(7.7053,1.0329)--(8.3474,1.1399)--(8.9895,1.2537)--(9.6316,1.3707)--(10.2737,1.4963)--(10.9158,1.6281)--(11.5579,1.7666)--(12.2000,1.9085);
438
  \end{scope}
439
  \draw[baselinegray,dashed,thick] (0.250,-1.23)--(0.800,-1.23);
440
  \node[anchor=west] at (0.900,-1.23) {Trackformer 1.1};
441
  \draw[modelteal,solid,thick] (6.350,-1.23)--(6.900,-1.23);
442
  \node[anchor=west] at (7.000,-1.23) {Trackformer 1.2: mean of 50};
443
  \end{tikzpicture}
444
+ \caption{Completed equal-storm development comparison, with every daily issue
445
+ retained. Positions use a common issue-centred local-km projection. Heading
446
+ uses common steps where truth and both models move more than 1 km.
447
+ No uncertainty intervals are claimed.}
 
 
448
  \label{fig:benchmark}
449
  \end{figure}
450
+ \subsection{Interpretation: alignment, not shape alone}
451
+ Mean position error improves by 41.0\%, from 798.4 to 471.2 km; +120 h error
452
+ falls from 1,646.4 to 1,031.6 km. At +6 h, error is 48.3 versus 66.1 km.
453
+ Centred shape similarity is 0.8837 versus 0.7544 for 1.1; normalized
454
+ Fr\'echet path similarity is 0.6560 versus 0.5345. These are separate
455
+ measures, not accuracy percentages or evidence of a win on every storm.
456
+
457
+ Actual alignment requires predicted and observed positions to coincide at the
458
+ same valid times. Removing translation and scale can make displaced routes
459
+ look similar; the shape score cannot replace geographic error or an unshifted
460
+ overlay. Path similarity responds to displacement but is not a strict
461
+ same-time distance.
462
+
463
+ \keyline{Pressure coverage.}{The 1.2 central-pressure MAE is 12.62 hPa on
464
+ 40 storms with valid labels; basin MSLP MAE is 2.72 hPa on 270 storms.
465
+ No matched 1.1 pressure comparison is available. All 1,473 saved forecast
466
+ hashes were verified; the earlier 270-case/90-storm study remains separate.}
467
+ \reportpage
468
+ \section{Fung-wong: route alignment and intensity}
469
+ The selected forecast begins at 00 UTC on 7 November 2025 and uses the saved
470
+ 50-member policy. Figure~\ref{fig:fungwong} compares exact six-hour positions
471
+ without translating, rotating or scaling either route. Great-circle errors
472
+ are 130.9 km on average, 142.3 km at +120 h and 198.3 km at maximum; all 20
473
+ leads are within 200 km. These are separate from the benchmark's local-km scores.
474
+ \begin{figure}[H]
 
 
475
  \centering
476
  \begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]
477
  \node[font=\small\bfseries,align=center] at (6.100,7.800) {Fung-wong: geographic routes, +0 to +120 h};
 
597
  4-hPa isobars; none of the pressure fields is moved to match observations.}
598
  \label{fig:fungwong}
599
  \end{figure}
600
+ The broad curves align, but position, timing and intensity differences remain.
601
+ This user-selected development example is not representative performance.
602
+ The \href{https://yu314-coder.github.io/typhoon-predict/trackformer_1_2_fung_wong.mp4}{companion MP4}
603
+ animates the saved fields; no additional inference or route adjustment is used.
604
+ \reportpage
605
+ \section{Pressure-field interpretation and limitations}
606
+ \begin{figure}[H]
607
  \centering
608
+ \begin{tikzpicture}[scale=0.9,x=1cm,y=1cm,font=\scriptsize]
609
  \node[font=\small\bfseries,align=center] at (5.750,8.900) {Fung-wong: model-mean MSLP at +48 h};
610
  \draw[black!10] (0.676,0)--(0.676,8.5);
611
  \node[below=3pt] at (0.676,0) {124};
 
672
  \draw[routepink,solid,thick] (6.000,-1.23)--(6.550,-1.23);
673
  \node[anchor=west] at (6.650,-1.23) {Forecast route to +48 h};
674
  \end{tikzpicture}
675
+ \caption{Saved model-mean regional MSLP at +48 h (9 November 2025, 00 UTC), with 4-hPa isobars labelled every 8 hPa. Routes extend only to this valid time. The western boundary is $123^\circ$E. The field is not shifted to the observed centre; its minimum need not equal member-mean central pressure.}
 
 
 
 
 
 
676
  \label{fig:pressuremap}
677
  \end{figure}
678
+ \subsection{Intensity labels and map coverage}
679
+ Pressure truth is JMA best track from IBTrACS \texttt{TOKYO\_PRES}, not a JMA
680
+ forecast. Native-detail history was unavailable; $0.25^\circ$ output sampling
681
+ does not establish native resolution. From +66 h the centre leaves the fixed
682
+ regional patch. The video then shows only the saved coarse basin field there,
683
+ while the moving-core readout supplies central pressure. No detailed core is
684
+ invented beyond coverage.
685
+
686
+ There is no verified matched 1.1 pressure array for this benchmark. On the
687
+ repeatedly inspected TIP ten-start diagnostic, 1.2 pressure MAE is worse
688
+ (30.7 versus 20.1 hPa). Those overlapping starts are not independent tests;
689
+ route gains do not establish an all-metric improvement.
690
+
691
+ \subsection{What remains unverified}
692
+ The completed cohort includes 40 recent and 230 historical storms (1980--1999),
693
+ outside this checkpoint's fitting and validation years but not certified unused
694
+ across prior experiments. Complete five-day truth excludes short remaining
695
+ lifetimes. Historical hindcasts use a model trained on later years and
696
+ retrospective analyses. Generalization still requires an unused whole-storm
697
+ holdout, training-only preprocessing and whole-storm uncertainty resampling.
698
+
699
+ \subsection{Conclusion}
700
+ Trackformer 1.2 links recurrent environmental evolution to a translating
701
+ pressure anomaly and coupled geographic readout. Its development results
702
+ improve aggregate route measures over the archived 1.1 pipeline, but leave
703
+ large absolute errors and intensity failures; the older cohort also showed
704
+ a short-lead regression. The
705
+ released model is a reproducible research artifact, not a substitute for
706
+ official warnings.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
707
  \end{document}
release_tools/build_paper_figures.py CHANGED
@@ -50,11 +50,11 @@ def curves(x,y1,y2,xlim,ylim,xticks,yticks,width,height,title,xlabel,ylabel,labe
50
  return '\n'.join(lines)
51
 
52
  def generate():
53
- m=json.loads((ROOT/'evaluation/trackformer_1_2_vs_1_1_270_metrics.json').read_text())['metrics']
54
  with np.load(ROOT/'evaluation/release_data/fung_wong_video.npz') as z:a={k:z[k] for k in z.files}
55
  lead=np.arange(6,121,6)
56
  bars=[r'\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]']
57
- for offset,title,values,top,ticks,unit in [(0,'Mean position error',[902.3,714.4],1100,[0,500,1000],'km; lower is better'),(7.2,'Six-hour heading error',[56.22,44.82],70,[0,20,40,60],'degrees; lower is better')]:
58
  bars.append(rf'\begin{{scope}}[shift={{({offset},0)}}]')
59
  bars.append(rf'\node[font=\small\bfseries] at (2.5,4.05) {{{title}}};')
60
  bars.append(rf'\node[rotate=90] at (-.75,1.75) {{{unit}}};')
@@ -96,7 +96,10 @@ def main():
96
  text=match.group()
97
  for label,picture in figures.items():
98
  if '\\label{'+label+'}' in text:
99
- return '\\begin{figure}[htbp]\n\\centering\n'+picture+'\n'+text[text.index('\\caption'):]
 
 
 
100
  return text
101
  new=re.sub(r'\\begin\{figure\}.*?\\end\{figure\}',replace,new,flags=re.S)
102
  print('*** Begin Patch\n*** Update File: '+str(path))
 
50
  return '\n'.join(lines)
51
 
52
  def generate():
53
+ m=json.loads((ROOT/'evaluation/daily_storm_final.json').read_text())['aggregate_equal_storm']
54
  with np.load(ROOT/'evaluation/release_data/fung_wong_video.npz') as z:a={k:z[k] for k in z.files}
55
  lead=np.arange(6,121,6)
56
  bars=[r'\begin{tikzpicture}[x=1cm,y=1cm,font=\scriptsize]']
57
+ for offset,title,values,top,ticks,unit in [(0,'Mean position error',[round(m[k]['mean_track_error_km'],1) for k in ['1.1','1.2']],1000,[0,500,1000],'km; lower is better'),(7.2,'Six-hour heading error',[round(m[k]['direction_error_deg'],2) for k in ['1.1','1.2']],70,[0,20,40,60],'degrees; lower is better')]:
58
  bars.append(rf'\begin{{scope}}[shift={{({offset},0)}}]')
59
  bars.append(rf'\node[font=\small\bfseries] at (2.5,4.05) {{{title}}};')
60
  bars.append(rf'\node[rotate=90] at (-.75,1.75) {{{unit}}};')
 
96
  text=match.group()
97
  for label,picture in figures.items():
98
  if '\\label{'+label+'}' in text:
99
+ if label == 'fig:pressuremap':
100
+ picture = picture.replace(r'\begin{tikzpicture}[x=1cm', r'\begin{tikzpicture}[scale=0.9,x=1cm')
101
+ opening = text[:text.index('\\begin{tikzpicture}')]
102
+ return opening+picture+'\n'+text[text.index('\\caption'):]
103
  return text
104
  new=re.sub(r'\\begin\{figure\}.*?\\end\{figure\}',replace,new,flags=re.S)
105
  print('*** Begin Patch\n*** Update File: '+str(path))
release_tools/plot_daily_storm_final.py ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Plot the completed equal-storm development benchmark; no new inference."""
2
+ from pathlib import Path
3
+ import json
4
+ import matplotlib
5
+ matplotlib.use('Agg')
6
+ import matplotlib.pyplot as plt
7
+
8
+ ROOT = Path(__file__).resolve().parents[1]
9
+ report = json.loads((ROOT / 'evaluation/daily_storm_final.json').read_text())
10
+ assert report['status'] == 'complete' and report['completed_storms'] == 270
11
+ metrics = report['aggregate_equal_storm']
12
+ panels = [
13
+ ('mean_track_error_km', 'Mean position error ↓', 'km', 1),
14
+ ('track_error_120h_km', '+120 h position error ↓', 'km', 1),
15
+ ('direction_error_deg', 'Six-hour heading error ↓', 'degrees', 2),
16
+ ('shape_similarity', 'Centred route shape ↑', 'similarity (not accuracy)', 4),
17
+ ('path_similarity', 'Geographic path similarity ↑', 'similarity (not accuracy)', 4),
18
+ ('frechet_distance_km', 'Fréchet distance ↓', 'km', 1),
19
+ ]
20
+ plt.rcParams.update({'font.family': 'DejaVu Sans', 'font.size': 10, 'axes.spines.top': False, 'axes.spines.right': False})
21
+ fig, axes = plt.subplots(2, 3, figsize=(13, 8))
22
+ for ax, (key, title, unit, digits) in zip(axes.flat, panels):
23
+ values = [metrics[m][key] for m in ['1.1', '1.2']]
24
+ bars = ax.bar(['1.1', '1.2 · mean of 50'], values, width=.55, color=['#65758b', '#007d78'])
25
+ ax.bar_label(bars, labels=[f'{v:,.{digits}f}' for v in values], padding=6, fontweight='bold')
26
+ ax.set_title(title, loc='left', pad=14, fontweight='bold')
27
+ ax.set_ylabel(unit)
28
+ ax.set_ylim(0, 1.06 if 'similarity' in key else max(values)*1.22)
29
+ ax.set_axisbelow(True)
30
+ ax.grid(axis='y', alpha=.16)
31
+ fig.suptitle('Trackformer 1.2 vs 1.1 | 270 storms · 1,473 daily forecasts', x=.07, ha='left', fontsize=17, fontweight='bold')
32
+ fig.text(.07, .922, 'Complete +6 to +120 h rollouts · each storm has equal weight · broader development evidence', color='#526174')
33
+ fig.text(.07, .022, 'Different input pipelines; not an architecture ablation or certified untouched test. No matched 1.1 pressure comparison.', fontsize=9, color='#526174')
34
+ fig.subplots_adjust(top=.84, bottom=.11, left=.07, right=.98, hspace=.42, wspace=.35)
35
+ fig.savefig(ROOT / 'evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png', dpi=180, facecolor='white')