euler314 commited on
Commit
7cbcfd5
·
verified ·
1 Parent(s): c0b3ee5

Add audited best route examples and frozen 270-storm daily benchmark protocol

Browse files
README.md CHANGED
@@ -1,10 +1,10 @@
1
  ---
2
  license: mit
3
  tags:
4
- - tropical-cyclone
5
- - weather-forecasting
6
- - sea-level-pressure
7
- - research
8
  ---
9
 
10
  # Trackformer 1.2
@@ -55,12 +55,16 @@ The saved 1.1 archive had a coordinate-label bug: its `v11_local` array held abs
55
 
56
  ### What the routes look like
57
 
58
- These are cases 1, 136 and 270, selected by their row positions before inspecting their forecast quality. Every curve starts at the same issue-time origin. The examples show both useful motion and substantial remaining errors, including missed turns.
59
 
60
- ![Observed, 1.1 and 1.2 routes on three fixed benchmark examples](paper/trackformer_1_2_vs_1_1_route_examples.png)
61
 
62
  The 1.2 mean has better aggregate direction, shape and position scores on this previously inspected development cohort. That does not establish an overall win on every storm or pressure metric. A same-270 1.1 central-pressure prediction array was not verified in these route artifacts, so it is not assigned a pressure-error bar. The separate 1.2 report gives 15.73 hPa central-pressure MAE and 2.54 hPa area-weighted regional MSLP MAE. TIP remains a separate diagnostic: its pressure error was worse for 1.2 (30.7 vs 20.1 hPa). See [metric definitions and limitations](docs/trackformer_1_2_evaluation.md).
63
 
 
 
 
 
64
  ## Recent pressure forecast with isobars
65
 
66
  The archived **Surigae forecast issued 27 September 2026 at 12:00 UTC** shows the model's regional pressure field at +6, +24 and +36 h. Thin lines are isobars every **4 hPa**, with selected labels every 12 hPa; magenta shows the forecast track and centre. Coastlines provide geographic context. These panels use the actual saved model values.
 
1
  ---
2
  license: mit
3
  tags:
4
+ - tropical-cyclone
5
+ - weather-forecasting
6
+ - sea-level-pressure
7
+ - research
8
  ---
9
 
10
  # Trackformer 1.2
 
55
 
56
  ### What the routes look like
57
 
58
+ These are **six selected best-performing examples from six distinct storms**, not a representative sample. Among in-domain cases with at least 300 km of observed travel, shape similarity ≥0.90 and direction error ≤30° on at least 18 common valid leads, we select the lowest 1.2 mean track errors, with one case per storm. Every curve starts at the same issue-time origin. The [selection manifest and all 270 case scores](paper/trackformer_1_2_showcase_selection.json) make the choice auditable; the aggregate bars above still include every case, including poor forecasts.
59
 
60
+ ![Observed, 1.1 and 1.2 routes on six selected best-performing examples](paper/trackformer_1_2_vs_1_1_route_examples.png)
61
 
62
  The 1.2 mean has better aggregate direction, shape and position scores on this previously inspected development cohort. That does not establish an overall win on every storm or pressure metric. A same-270 1.1 central-pressure prediction array was not verified in these route artifacts, so it is not assigned a pressure-error bar. The separate 1.2 report gives 15.73 hPa central-pressure MAE and 2.54 hPa area-weighted regional MSLP MAE. TIP remains a separate diagnostic: its pressure error was worse for 1.2 (30.7 vs 20.1 hPa). See [metric definitions and limitations](docs/trackformer_1_2_evaluation.md).
63
 
64
+ ### Expanded daily-issue benchmark: 270 distinct typhoons
65
+
66
+ A new **1,473-case / 270-storm** benchmark is in progress; it is not the completed 270-case result above. Each storm contributes at most one forecast per UTC day, through +120 h. Daily errors are averaged within each storm, then the 270 storm scores receive equal weight. Selection was frozen before new inference, without filtering on forecast quality. See the [protocol](docs/daily_storm_benchmark.md) and [frozen cohort](paper/release_data/daily_storm_cohort.json). Results are pending; no improvement is claimed from the partial run.
67
+
68
  ## Recent pressure forecast with isobars
69
 
70
  The archived **Surigae forecast issued 27 September 2026 at 12:00 UTC** shows the model's regional pressure field at +6, +24 and +36 h. Thin lines are isobars every **4 hPa**, with selected labels every 12 hPa; magenta shows the forecast track and centre. Coastlines provide geographic context. These panels use the actual saved model values.
RELEASE_NOTES_TRACKFORMER_1_2.md CHANGED
@@ -6,6 +6,8 @@ The attached archive includes inference-only weights, exact source modules, an i
6
 
7
  On the 270-case development cohort, 1.1 versus 1.2 mean-of-50 scores are: direction error **56.22° vs 44.82°**, route-shape similarity **0.7273 vs 0.8249**, path similarity **0.4983 vs 0.5727**, and mean track error **902.3 vs 714.4 km**. Direction uses 5,382 common valid steps. The old 1.1 report had stored latitude/longitude under a local-kilometre key; the new comparison repairs that coordinate interpretation and records the correction. Original artifacts remain available. These results are development comparisons, and no matched 1.1 pressure-field bar is claimed.
8
 
 
 
9
  The Surigae illustration is a single forecast issued 2026-09-27 12 UTC, with 4 hPa isobars at +6/+24/+36 h. It is separate from the 50-member historical benchmark. Reproduction data and metric definitions are included. The model weights are unchanged in this documentation revision.
10
 
11
  This is not an operational warning service. Do not use for safety-critical decisions. A genuinely untouched storm-level holdout is still required before generalization claims.
 
6
 
7
  On the 270-case development cohort, 1.1 versus 1.2 mean-of-50 scores are: direction error **56.22° vs 44.82°**, route-shape similarity **0.7273 vs 0.8249**, path similarity **0.4983 vs 0.5727**, and mean track error **902.3 vs 714.4 km**. Direction uses 5,382 common valid steps. The old 1.1 report had stored latitude/longitude under a local-kilometre key; the new comparison repairs that coordinate interpretation and records the correction. Original artifacts remain available. These results are development comparisons, and no matched 1.1 pressure-field bar is claimed.
8
 
9
+ The route gallery now shows six explicitly selected best-performing examples from distinct storms, with the selection rule and all case scores supplied. It is not representative evidence; the aggregate bars retain all 270 cases. A separate larger benchmark has frozen **270 distinct storms / 1,473 daily issues**, with daily scores averaged per storm and storms weighted equally. Its forecasts are being computed; results are pending.
10
+
11
  The Surigae illustration is a single forecast issued 2026-09-27 12 UTC, with 4 hPa isobars at +6/+24/+36 h. It is separate from the 50-member historical benchmark. Reproduction data and metric definitions are included. The model weights are unchanged in this documentation revision.
12
 
13
  This is not an operational warning service. Do not use for safety-critical decisions. A genuinely untouched storm-level holdout is still required before generalization claims.
docs/daily_storm_benchmark.md ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Daily forecasts, equal-storm benchmark
2
+
3
+ Status: inference in progress. This document specifies the frozen evaluation, not its final results.
4
+
5
+ ## What counts as a case and a score
6
+
7
+ One case is one typhoon at one issue time on one UTC day. Choose the earliest eligible issue on that day, once per typhoon. Each forecast predicts all 20 six-hour leads, +6 through +120 h, from information at or before issue time.
8
+
9
+ For position error, daily score = mean of the 20 lead errors; storm score = mean of that storm's daily scores; benchmark score = mean of the 270 storm scores. A storm with ten available daily starts has exactly the same final weight as a storm with three. The per-lead comparison also averages daily values within storm before averaging storms.
10
+
11
+ Direction error, route-shape similarity, Fréchet distance and path similarity remain separate metrics. Their definitions match the [evaluation notes](trackformer_1_2_evaluation.md); direction uses only common non-stationary leads. Missing direction scores are reported, not replaced by zero. There is no invented combined weighted quality mark.
12
+
13
+ ## Frozen cohort
14
+
15
+ - 270 distinct storms, 1,473 daily issue cases, selected from 523 locally eligible storms.
16
+ - All 40 eligible recent storms beginning in 2024 or later, plus 230 historical storms selected at evenly spaced positions in the chronological 1980–1999 storm list.
17
+ - The entire storm must lie outside the selected 1.2 checkpoint's 2000–2021 fitting and 2022–2023 validation years, with the history/target boundary also outside those years.
18
+ - Issue centre within 0–60°N / 100–180°E; exact six-hour weather history through issue; complete route labels through +120 h; matching pressure-atlas coverage for scoring.
19
+ - Use all eligible daily starts for each chosen storm. Never drop a case because its forecast looks poor.
20
+ - [Cohort manifest](../paper/release_data/daily_storm_cohort.json) freezes rows, dates, storm IDs, source hashes and checkpoint identity before inference. Forecast artifacts are new; some recent storms overlap earlier development comparisons.
21
+
22
+ Requiring complete five-day truth excludes short remaining lifetimes. This is a declared coverage limitation, not an all-storm operational sample. Historical cases are retrospective hindcasts from a model trained on later years; reanalysis and best-track issue states do not establish real-time data availability. This cohort has not been certified untouched across all earlier experiments.
23
+
24
+ ## Models and input boundaries
25
+
26
+ Trackformer **1.2** uses the unchanged released checkpoint and 50 independently seeded, smooth input-perturbation members. Seeds are `2043 + source_row * 100 + member_id`, with member IDs 0–49. Normalized perturbation amplitudes remain 0.025 for basin history and 0.015 for regional history. These are not learned latent samples or 50 separate trained models. Member routes and central pressures are averaged; physical fields are averaged on fixed common geographic grids. Missing native regional inputs retain the model's missing-detail mask.
27
+
28
+ Trackformer **1.1** uses the released causal weighted route builder with issue, −12 h and −24 h analyses and historical motion. Its absolute route is converted into the same issue-centred local-km coordinate system as 1.2. The pipelines differ, so this is a released-system comparison, not an architecture-only ablation.
29
+
30
+ Future weather/track/pressure are used for eligibility coverage and scoring only, never inference. No retraining, coefficient tuning, new weather downloads or checkpoint substitution occurs in this run.
31
+
32
+ ## Pressure and failures
33
+
34
+ For 1.2, report central-pressure MAE over valid labels and cosine-area-weighted basin MSLP MAE on the native basin grid, daily then equally per storm. Keep pressure label coverage visible. No matched 1.1 pressure score is fabricated; this run does not establish superiority on pressure. Regional/core field scores need their own valid coverage and are not implied by basin-wide error.
35
+
36
+ Completed daily forecasts are saved atomically with hashes. A nonfinite forecast or source mismatch stops the run visibly rather than silently deleting the case. Resume skips completed cases without rerunning them. Partial aggregate scores include only fully completed storms, identify their counts, and must not be presented as the final 270-storm result. Final uncertainty should resample whole storms, not overlapping daily issues independently.
docs/trackformer_1_2_evaluation.md CHANGED
@@ -17,7 +17,7 @@ The main chart compares 270 issue cases from 90 storms and all 20 leads (+6 thro
17
  - **Path similarity:** calculate discrete Fréchet distance on the ordered 20-point forecast/truth curves, then `exp(-distance / observed path length)`, and average over cases. The observed length includes the step from the issue origin. Means: **0.498254 for 1.1; 0.572741 for 1.2**. This responds to geographic displacement as well as the curve.
18
  - **Position error:** mean Euclidean error in the saved benchmark's local-kilometre projection, **902.3001 vs 714.4444 km**. At +120 h it is **1,833.2481 vs 1,433.3863 km**. These are local-coordinate distances, not newly calculated geodesic distances.
19
 
20
- No combined weighted ranking is introduced. Cases 1, 136 and 270 are shown for visual comparison, chosen by row position without filtering on forecast quality. Their axes retain actual kilometre displacements from the shared issue origin.
21
 
22
  The full 270 cohort includes cases beyond the intended Western Pacific model domain. Its aggregate is a legacy development comparison, not a basin-specific validation claim. The previously inspected 133-case subset is separate and is not substituted for the requested 270 cases.
23
 
@@ -43,7 +43,11 @@ The public repository includes the small numerical plot inputs under `paper/rele
43
  python release_tools/plot_release_270.py
44
  ```
45
 
46
- This regenerates the four-panel benchmark, per-lead curve, fixed route examples, isobar figure and metrics JSON. `--source-root` is optional and is only used to rebuild the small plot inputs from the original local research archives. No inference, fitting or new weather retrieval is performed by this plotting script.
 
 
 
 
47
 
48
  ## Validation boundary
49
 
 
17
  - **Path similarity:** calculate discrete Fréchet distance on the ordered 20-point forecast/truth curves, then `exp(-distance / observed path length)`, and average over cases. The observed length includes the step from the issue origin. Means: **0.498254 for 1.1; 0.572741 for 1.2**. This responds to geographic displacement as well as the curve.
18
  - **Position error:** mean Euclidean error in the saved benchmark's local-kilometre projection, **902.3001 vs 714.4444 km**. At +120 h it is **1,833.2481 vs 1,433.3863 km**. These are local-coordinate distances, not newly calculated geodesic distances.
19
 
20
+ No combined weighted ranking is introduced. The route gallery shows six deliberately selected best-performing examples, one per distinct storm, not a representative sample. Eligibility requires an issue inside 0–60°N / 100–180°E, observed travel ≥300 km, 1.2 shape similarity ≥0.90, and direction error ≤30° on ≥18 common valid leads. Eligible cases are sorted by ascending 1.2 mean track error, then the first six distinct storms are selected. `paper/trackformer_1_2_showcase_selection.json` records the rule, chosen cases and scores for all 270 cases. Axes retain actual kilometre displacements; no route is moved or rotated to improve the visual comparison. Full-cohort bars are unchanged.
21
 
22
  The full 270 cohort includes cases beyond the intended Western Pacific model domain. Its aggregate is a legacy development comparison, not a basin-specific validation claim. The previously inspected 133-case subset is separate and is not substituted for the requested 270 cases.
23
 
 
43
  python release_tools/plot_release_270.py
44
  ```
45
 
46
+ This regenerates the four-panel benchmark, per-lead curve, selected-best route examples, full selection audit, isobar figure and metrics JSON. `--source-root` is optional and is only used to rebuild the small plot inputs from the original local research archives. No inference, fitting or new weather retrieval is performed by this plotting script.
47
+
48
+ ## Expanded benchmark in progress
49
+
50
+ The [daily-storm protocol](daily_storm_benchmark.md) freezes 270 distinct Western Pacific storms and 1,473 daily issue cases. It averages leads within each daily issue, daily issues within each storm, and then storms equally. This new run does not reuse the old 270 cases as though they were 270 storms. Results remain pending until the whole frozen cohort is evaluated.
51
 
52
  ## Validation boundary
53
 
paper/release_data/daily_storm_cohort.json ADDED
The diff for this file is too large to render. See raw diff
 
paper/trackformer_1_2_showcase_selection.json ADDED
The diff for this file is too large to render. See raw diff
 
paper/trackformer_1_2_vs_1_1_route_examples.png CHANGED

Git LFS Details

  • SHA256: d1b5df488f6ee79706ee0b4f09aab97c775e155ca768df43266d33dd95e504a1
  • Pointer size: 131 Bytes
  • Size of remote file: 222 kB

Git LFS Details

  • SHA256: 463df86fa2e6f31f0ac14a87b7110cbd437bf0298b6ba3e90a874f891c72c5fa
  • Pointer size: 131 Bytes
  • Size of remote file: 379 kB
release_tools/plot_release_270.py CHANGED
@@ -120,11 +120,50 @@ def similarity_metrics(routes, truth):
120
  return result
121
 
122
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
123
  def plot_benchmark(data, output):
124
  meta = json.loads((data / "cohort_270.json").read_text())
125
  with np.load(data / "routes_270.npz", allow_pickle=False) as z:
126
  truth = z["truth_local"].astype("float64")
127
  routes = [z[k].astype("float64") for k in ("v11_local", "v12_local")]
 
128
  errors = [np.linalg.norm(p - truth, axis=-1) for p in routes]
129
  similarity = similarity_metrics(routes, truth)
130
  assert all(e.shape == (270, 20) and np.isfinite(e).all() for e in errors)
@@ -171,21 +210,25 @@ def plot_benchmark(data, output):
171
  "position": "Euclidean displacement error in the existing benchmark local-km coordinate system"}
172
  meta["mean_track_error_reduction_percent"] = float(100*(1-errors[1].mean()/errors[0].mean()))
173
  (output / "trackformer_1_2_vs_1_1_270_metrics.json").write_text(json.dumps(meta, indent=2) + "\n")
174
- fig, axes = plt.subplots(1, 3, figsize=(13, 5), layout="constrained")
175
- for ax, case_index in zip(axes, (0, 135, 269)):
 
 
 
176
  for route, color, label in [(truth, "#182c42", "Observed"), (routes[0], colors[0], "1.1"), (routes[1], colors[1], "1.2 · mean of 50")]:
177
  points = np.vstack([np.zeros((1,2)), route[case_index]])
178
  ax.plot(points[:,0], points[:,1], color=color, label=label, linewidth=2,
179
  linestyle="--" if label=="1.1" else "-", marker="o", markersize=2.5)
180
- ax.annotate("+120 h", points[-1], color=color, fontsize=8, xytext=(3,3), textcoords="offset points")
181
  c = meta["cases"][case_index]
182
- ax.set_title(f"Case {case_index+1} · {c['storm_id']}\n{c['issue_time_utc'][:16].replace('T',' ')} UTC", fontsize=10)
 
183
  ax.set(xlabel="East displacement (km)", ylabel="North displacement (km)")
184
  ax.set_aspect("equal", adjustable="datalim")
185
  ax.grid(alpha=.18)
186
- axes[0].legend(frameon=False, fontsize=8)
187
- fig.suptitle("What the tracks look like · three fixed examples from the 270-case cohort", fontsize=15, fontweight="bold")
188
- fig.text(.5, -.08, "Cases 1, 136 and 270 chosen by row position, without filtering by forecast quality. All start at the same issue-time origin.", ha="center", fontsize=9)
189
  fig.savefig(output / "trackformer_1_2_vs_1_1_route_examples.png", dpi=180, bbox_inches="tight")
190
  plt.close(fig)
191
 
 
120
  return result
121
 
122
 
123
+ def select_showcase(routes, truth, meta, base_lat, base_lon):
124
+ steps = [np.diff(np.concatenate([np.zeros((len(p), 1, 2)), p], axis=1), axis=1) for p in [truth, *routes]]
125
+ valid = np.logical_and.reduce([np.linalg.norm(s, axis=-1) > 1 for s in steps])
126
+ all_cases = []
127
+ for i, case in enumerate(meta["cases"]):
128
+ row = {"case_index": i, **case, "base_lat": float(base_lat[i]), "base_lon": float(base_lon[i]),
129
+ "truth_path_length_km": float(np.linalg.norm(steps[0][i], axis=-1).sum()), "models": {}}
130
+ for key, route, step in zip(("1.1", "1.2"), routes, steps[1:]):
131
+ a = (route[i] - route[i].mean(axis=0)).ravel()
132
+ b = (truth[i] - truth[i].mean(axis=0)).ravel()
133
+ cosine = np.dot(a, b) / max(np.linalg.norm(a)*np.linalg.norm(b), 1e-8)
134
+ delta = np.arctan2(step[i,:,1], step[i,:,0]) - np.arctan2(steps[0][i,:,1], steps[0][i,:,0])
135
+ angles = np.degrees(np.abs(np.arctan2(np.sin(delta), np.cos(delta))))
136
+ error = np.linalg.norm(route[i]-truth[i], axis=-1)
137
+ row["models"][key] = {"mean_track_error_km": float(error.mean()),
138
+ "track_error_120h_km": float(error[-1]),
139
+ "shape_similarity": float(np.clip((1+cosine)/2, 0, 1)),
140
+ "direction_error_deg": float(angles[valid[i]].mean()) if valid[i].any() else None,
141
+ "direction_valid_steps": int(valid[i].sum())}
142
+ all_cases.append(row)
143
+ eligible = [c for c in all_cases if 0 <= c["base_lat"] <= 60 and 100 <= c["base_lon"] <= 180
144
+ and c["truth_path_length_km"] >= 300 and c["models"]["1.2"]["shape_similarity"] >= .9
145
+ and c["models"]["1.2"]["direction_valid_steps"] >= 18
146
+ and c["models"]["1.2"]["direction_error_deg"] <= 30]
147
+ chosen, storms = [], set()
148
+ for case in sorted(eligible, key=lambda c: (c["models"]["1.2"]["mean_track_error_km"], c["case_index"])):
149
+ if case["storm_id"] not in storms:
150
+ chosen.append(case)
151
+ storms.add(case["storm_id"])
152
+ if len(chosen) == 6:
153
+ break
154
+ if len(chosen) != 6:
155
+ raise ValueError("Fewer than six distinct storms meet the declared showcase thresholds")
156
+ return {"selection": "Selected best-performing examples, not a representative sample",
157
+ "rule": "Within 0-60N / 100-180E; observed path length >=300 km; 1.2 shape >=0.90; direction error <=30 degrees on >=18 common valid leads; ascending 1.2 mean track error; one case per storm; first six distinct storms",
158
+ "eligible_case_count": len(eligible), "selected": chosen, "all_case_metrics": all_cases}
159
+
160
+
161
  def plot_benchmark(data, output):
162
  meta = json.loads((data / "cohort_270.json").read_text())
163
  with np.load(data / "routes_270.npz", allow_pickle=False) as z:
164
  truth = z["truth_local"].astype("float64")
165
  routes = [z[k].astype("float64") for k in ("v11_local", "v12_local")]
166
+ base_lat, base_lon = z["base_lat"], z["base_lon"]
167
  errors = [np.linalg.norm(p - truth, axis=-1) for p in routes]
168
  similarity = similarity_metrics(routes, truth)
169
  assert all(e.shape == (270, 20) and np.isfinite(e).all() for e in errors)
 
210
  "position": "Euclidean displacement error in the existing benchmark local-km coordinate system"}
211
  meta["mean_track_error_reduction_percent"] = float(100*(1-errors[1].mean()/errors[0].mean()))
212
  (output / "trackformer_1_2_vs_1_1_270_metrics.json").write_text(json.dumps(meta, indent=2) + "\n")
213
+ showcase = select_showcase(routes, truth, meta, base_lat, base_lon)
214
+ (output / "trackformer_1_2_showcase_selection.json").write_text(json.dumps(showcase, indent=2) + "\n")
215
+ fig, axes = plt.subplots(2, 3, figsize=(15, 10), layout="constrained")
216
+ for ax, selected in zip(axes.ravel(), showcase["selected"]):
217
+ case_index = selected["case_index"]
218
  for route, color, label in [(truth, "#182c42", "Observed"), (routes[0], colors[0], "1.1"), (routes[1], colors[1], "1.2 · mean of 50")]:
219
  points = np.vstack([np.zeros((1,2)), route[case_index]])
220
  ax.plot(points[:,0], points[:,1], color=color, label=label, linewidth=2,
221
  linestyle="--" if label=="1.1" else "-", marker="o", markersize=2.5)
222
+ ax.scatter(*points[-1], color=color, s=24, zorder=5)
223
  c = meta["cases"][case_index]
224
+ scores = selected["models"]["1.2"]
225
+ ax.set_title(f"{c['storm_id']} · {c['issue_time_utc'][:10]}\n1.2: {scores['mean_track_error_km']:.0f} km · shape {scores['shape_similarity']:.3f} · direction {scores['direction_error_deg']:.1f}°", fontsize=10)
226
  ax.set(xlabel="East displacement (km)", ylabel="North displacement (km)")
227
  ax.set_aspect("equal", adjustable="datalim")
228
  ax.grid(alpha=.18)
229
+ axes[0,0].legend(frameon=False, fontsize=9)
230
+ fig.suptitle("Trackformer 1.2 · selected best-performing route examples\nSix distinct storms from the 270-case benchmark", fontsize=16, fontweight="bold")
231
+ fig.text(.5, -.045, "Selected for low 1.2 track error, high shape similarity and low direction error. These are showcase examples, not typical performance.", ha="center", fontsize=10)
232
  fig.savefig(output / "trackformer_1_2_vs_1_1_route_examples.png", dpi=180, bbox_inches="tight")
233
  plt.close(fig)
234