euler314 commited on
Commit
1e1dffa
·
verified ·
1 Parent(s): 63bb05a

Publish Trackformer 1.2 announcement assets and matched DeepMind benchmark protocol

Browse files
.gitattributes CHANGED
@@ -195,3 +195,4 @@ evaluation/storm_structure/meranti_wind_radius.png filter=lfs diff=lfs merge=lfs
195
  evaluation/storm_structure/soudelor_wind_radius.png filter=lfs diff=lfs merge=lfs -text
196
  evaluation/release_data/pressure_core/mangkhut/common-pressure.npz filter=lfs diff=lfs merge=lfs -text
197
  evaluation/released_daily/pressure_comparison.png filter=lfs diff=lfs merge=lfs -text
 
 
195
  evaluation/storm_structure/soudelor_wind_radius.png filter=lfs diff=lfs merge=lfs -text
196
  evaluation/release_data/pressure_core/mangkhut/common-pressure.npz filter=lfs diff=lfs merge=lfs -text
197
  evaluation/released_daily/pressure_comparison.png filter=lfs diff=lfs merge=lfs -text
198
+ evaluation/released_daily/model_1_2_benchmark.png filter=lfs diff=lfs merge=lfs -text
docs/deepmind_daily_benchmark.md ADDED
@@ -0,0 +1,43 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Matched WeatherNext Cyclones Mini daily benchmark
2
+
3
+ ## Model and frozen plan
4
+
5
+ The reference is Google's official [WeatherNext Cyclones Mini](https://github.com/google-deepmind/weathernext), the released `<2024` checkpoint trained through 2023—not a custom imitation or full-sized WeatherNext. It runs one seeded member using the existing JAX CPU runtime on the Mac. This is a new benchmark; scores remain pending until saved forecasts and the full verification receipt exist.
6
+
7
+ - Cohort: **1,473 daily starts / 270 Western Pacific storms**, identical to the frozen Trackformer 1.1/1.2 route comparison.
8
+ - Forecast: twenty six-hour steps, **+6 to +120 h**, with the exact observed issue-time centre as the initialized-storm anchor.
9
+ - Inputs: exact ERA5 analyses at **issue−6 h and issue**, calendar forcings only after issue. No future weather, storm labels or official forecast tracks enter the worker.
10
+ - Weights SHA-256: `a1bb151457077248d70a458b1a7b19deacd2926fbcb46eee4ff2691444e3596e`.
11
+ - Cohort SHA-256: `965184f1e5ab3fcb52e304f39b38583a14f753870aa2295f2d01942a441dcfd6`.
12
+ - Official code commit: `89c4b2a77a1c57b328b909c575550fd2e5aadc9c`.
13
+
14
+ ## Like-for-like scoring
15
+
16
+ All three models use the frozen starts, exact future label times and original issue-relative kilometre projection. Routes are unshifted. Direction is recomputed on **common moving steps across truth and all three models**, not compared using different masks.
17
+
18
+ Pressure uses the original 134 eligible native-intensity starts / 40 storms, restricted to common valid leads across three predictions and the selected reference. JMA and USA remain separate. MAE and centred time-curve similarity are separate; missing, nonphysical or failed outputs are not zeros. Flat or shorter-than-six-point curves have unavailable shape similarity.
19
+
20
+ Average valid leads within an issue, days within a storm, then storms equally. Partial aggregates contain only fully completed storms and explicit coverage. Recent 2024+ and historical 1980–1999 groups are separate because the latter overlap WeatherNext fitting years. Different input pipelines and the one-member Mini versus 50-input-member 1.2 policy make this an output comparison, not an equal-compute architecture ablation or a certified unused test.
21
+
22
+ The official six-hour direct tracker uses a frozen initialized-storm continuity policy: no cyclogenesis, no dissipation pruning and no nearby-cyclone pruning. This is disclosed rather than called the operational default. Native 1° WP MSLP grids are saved in physical hPa; cyclone-head central pressure is not the basin-grid minimum.
23
+
24
+ ## Execution and verification
25
+
26
+ The [resumable runner](../release_tools/deepmind_daily_benchmark.py) freezes source hashes, the cohort, original baseline forecast hashes and checkpoint identity. Workers save exact input-state hashes, twenty route points, central-pressure values/masks and twenty native WP grids. Future truth is read only **after** inference for scoring; Trackformer forecasts and weights are unchanged.
27
+
28
+ Local output: `/Volumes/D/typhoon_predict/output/deepmind-daily-20261002`. `progress.json` records on-disk coverage; `logs/` holds diagnostics. `verification.json` is produced only after all 1,473 cases pass the final audit. These are local results, not automatically published Site API data. No score is uploaded before review.
29
+
30
+ Existing Mac runtime and weights are reused without package installation. Earlier CPU rollouts suggest **several days** for the full run; retrieval and compilation affect the total. All generated states, caches and outputs stay on `/Volumes/D`. Cloud backfill, training and Site deployments are untouched.
31
+
32
+ ```bash
33
+ # Existing compatible WeatherNext environment and project assets required.
34
+ python release_tools/deepmind_daily_benchmark.py prepare \
35
+ --project /Volumes/D/typhoon_predict \
36
+ --output /Volumes/D/typhoon_predict/output/deepmind-daily-20261002
37
+
38
+ python release_tools/deepmind_daily_benchmark.py launch \
39
+ --project /Volumes/D/typhoon_predict \
40
+ --output /Volumes/D/typhoon_predict/output/deepmind-daily-20261002
41
+ ```
42
+
43
+ `prepare` is only for a fresh directory. `launch` refuses an active run; `run` resumes completed cases without repetition. Source retrieval has bounded retries and failures remain visible. Final completion requires all planned cases, exact leads, valid routes/fields, unchanged forecast hashes and the verification receipt—not a process exit or count alone.
docs/showcase_archive.md ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Forecast examples and diagnostics
2
+
3
+ The [model announcement](../README.md) features the geographically registered moving-core Mangkhut (2018) forecast. Its 50-member arrays and replay/playback receipts are linked there. These are development illustrations, not representative or untouched-test evidence.
4
+
5
+ ## Additional films
6
+
7
+ These earlier 50-member films retain a fixed issue-relative regional pressure patch. When the centre leaves it, only the saved coarse basin field covers that position. They use 4 hPa isobars and 23-second playback with a held ending; do not confuse them with the current moving-core Mangkhut rendering. Their original model outputs remain preserved.
8
+
9
+ - [Fung-wong (2025), 7 November 00 UTC](https://huggingface.co/euler314/typhoon-predict/resolve/87a6e366b42bb4cc0edc95d2c50c55fca21a2c93/docs/trackformer_1_2_fung_wong.mp4) · [provenance](../evaluation/release_data/fung_wong_video.json)
10
+ - [Soudelor (2015), 5 August 00 UTC](https://huggingface.co/euler314/typhoon-predict/resolve/87a6e366b42bb4cc0edc95d2c50c55fca21a2c93/docs/trackformer_1_2_soudelor.mp4) · [provenance](../evaluation/release_data/soudelor_video.json)
11
+ - [Meranti (2016), 10 September 00 UTC](https://huggingface.co/euler314/typhoon-predict/resolve/87a6e366b42bb4cc0edc95d2c50c55fca21a2c93/docs/trackformer_1_2_meranti.mp4) · [provenance](../evaluation/release_data/meranti_video.json)
12
+
13
+ ## Wind and structure
14
+
15
+ [Four-storm wind comparison](../evaluation/storm_structure/four_storm_wind_candidate.png) · [audit index](../evaluation/storm_structure/index.json) · [experimental structure method](../release_tools/STORM_STRUCTURE_DIAGNOSTICS.md)
16
+
17
+ USA one-minute and JMA ten-minute winds are separate references. Native 1.1 RMW and quadrant R34/R50/R64 are scored against matching USA definitions. Trackformer 1.2 has an auxiliary maximum-wind scalar but no native radius head. Its pressure-derived radius candidate is not an official-equivalent prediction; masked/persistent components are not learned radius forecasts.
18
+
19
+ | Storm | Exact-time comparisons | Data and assumptions |
20
+ | --- | --- | --- |
21
+ | Fung-wong | [Plot](../evaluation/storm_structure/fung_wong_wind_radius.png) | [Record](../evaluation/storm_structure/fung_wong_structure.json) |
22
+ | Soudelor | [Plot](../evaluation/storm_structure/soudelor_wind_radius.png) | [Record](../evaluation/storm_structure/soudelor_structure.json) |
23
+ | Mangkhut | [Plot](../evaluation/storm_structure/mangkhut_wind_radius.png) | [Record](../evaluation/storm_structure/mangkhut_structure.json) |
24
+ | Meranti | [Plot](../evaluation/storm_structure/meranti_wind_radius.png) | [Record](../evaluation/storm_structure/meranti_structure.json) |
25
+
26
+ ## Full evaluation
27
+
28
+ [Daily protocol](daily_storm_benchmark.md) · [all daily scores](../evaluation/daily_storm_final.json) · [intensity/radius protocol](intensity_benchmark.md) · [native intensity results](../evaluation/intensity/intensity_final.json) · [matched pressure curves](../evaluation/released_daily/pressure_comparison.png)
29
+
30
+ Route uses 1,473 starts / 270 storms; the matched native pressure comparison uses 134 starts / 40 storms. JMA pressure MAE falls from 13.53 to 12.84 hPa, with paired uncertainty including zero. JMA curve similarity is 0.7074 / 0.7118 (1.1 / 1.2), while USA curve similarity is 0.7237 / 0.6637 on 131 eligible days—a regression. A centred shape score discards level and amplitude and must be read with physical hPa timelines, not used to conceal pressure errors.
docs/trackformer_1_2_architecture.svg ADDED
evaluation/released_daily/model_1_2_benchmark.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source": "evaluation/released_daily/released_daily_benchmark.json",
3
+ "source_sha256": "79f67a5f4490150cf0ab24d73e5d3cdedb4884d37e86560cbadde1d4ecdafc43",
4
+ "chart": "evaluation/released_daily/model_1_2_benchmark.png",
5
+ "chart_sha256": "3bd866eede33fc93d36a0a5373a3b1637bf0f56860938c3d5691b9b7dfe770b4",
6
+ "metrics": [
7
+ "pressure_JMA_hpa",
8
+ "mean_track_error_km",
9
+ "direction_error_deg"
10
+ ],
11
+ "values": {
12
+ "pressure_JMA_hpa": {
13
+ "1.1": 13.53233861138585,
14
+ "1.2": 12.837077366789085
15
+ },
16
+ "mean_track_error_km": {
17
+ "1.1": 798.4143395733709,
18
+ "1.2": 471.2045276140213
19
+ },
20
+ "direction_error_deg": {
21
+ "1.1": 51.58419207122625,
22
+ "1.2": 34.95994745459279
23
+ }
24
+ },
25
+ "deepmind_score": null,
26
+ "zero_fill": false,
27
+ "model_forecasts_modified": false
28
+ }
evaluation/released_daily/model_1_2_benchmark.png ADDED

Git LFS Details

  • SHA256: 3bd866eede33fc93d36a0a5373a3b1637bf0f56860938c3d5691b9b7dfe770b4
  • Pointer size: 131 Bytes
  • Size of remote file: 160 kB
models/trackformer_1_2_field/README.md CHANGED
@@ -1,6 +1,6 @@
1
  # Trackformer 1.2 field-model inference
2
 
3
- **Documentation revision: 1 October 2026.** This describes the current GitHub/Hugging Face wrapper. The 29 September release tarball remains an earlier snapshot; its optional wind/radius diagnostics may differ. Released neural weights, architecture, causal input schema and route/pressure predictions are unchanged.
4
 
5
  This folder contains the exact source, input contract and inference-only weights for the research model in the [main README](../../README.md). `model.py`, `baseline_model.py` and `v165_base.py` are source-identical to the selected training implementation. The weight file is distributed via the [GitHub release](https://github.com/yu314-coder/typhoon-predict/releases/tag/trackformer-1.2) and [Hugging Face](https://huggingface.co/euler314/typhoon-predict/tree/main/models/trackformer_1_2_field), not Git.
6
 
@@ -45,6 +45,6 @@ No model weights, learned architecture or existing benchmark results changed.
45
 
46
  The published 50-member mean is a separate evaluation policy: 50 deterministic seeds perturb only normalized historical basin/regional fields with smooth zero-centred noise, then average routes, central pressures and common-grid fields. `predict.py` computes one clean forecast; do not label it a 50-member mean.
47
 
48
- **Matched development evaluation:** the released 1.1 and 1.2 route comparison uses the same frozen **1,473 daily starts / 270 equal-weight storms**, +6 to +120 h. The completed native-pressure comparison has **134 common starts / 40 storms**; the other 1,339 starts lack valid frozen 1.1 intensity inputs and are not zero-scored. The Site defaults to central-pressure MAE: **13.53 / 12.84 hPa** against JMA and **12.84 / 12.55 hPa** against USA (1.1 / 1.2), slight mean improvements with paired whole-storm uncertainty including zero. The optional curve-similarity diagnostic is **(1 + centred cosine) / 2 at exact common valid times**, without time shifting or warping. JMA similarity is **0.7074 / 0.7118**; USA is **0.7237 / 0.6637** on 131 eligible non-flat curves. Removing level/amplitude makes this a shape diagnostic, not proof of correct pressure levels. Actual unshifted hPa timelines, agency masks and uncertainty remain separate. See the [shared snapshot](../../evaluation/released_daily/released_daily_benchmark.json), [verification receipt](../../evaluation/released_daily/released_daily_verification.json), [pressure protocol](../../docs/intensity_benchmark.md) and [public benchmark API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/api/benchmarks/released). DeepMind Mini has not run on this daily cohort; no old score is transferred into it.
49
 
50
- **Pressure-display coverage:** `regional_mslp_hpa` is the original fixed issue-relative composite, not an automatically extended moving-core map. When the forecast centre leaves that patch, a scalar central-pressure readout must not be inserted into the image to make it appear consistent. The corrected Mangkhut film separately exports the actual moving model cores and registers member fields before averaging; [the current model card](https://huggingface.co/euler314/typhoon-predict) links the verified pressure arrays and shows which other films retain their older fixed patches. The automatic History archive and its separate core recovery expose exact geography and coverage through [the public API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/history-api).
 
1
  # Trackformer 1.2 field-model inference
2
 
3
+ The released pressure-field model predicts Western Pacific storm tracks, central pressure and spatial MSLP through +120 h. This is the inference reference; see the [model announcement](../../README.md) for capabilities, the pressure-map showcase and development results.
4
 
5
  This folder contains the exact source, input contract and inference-only weights for the research model in the [main README](../../README.md). `model.py`, `baseline_model.py` and `v165_base.py` are source-identical to the selected training implementation. The weight file is distributed via the [GitHub release](https://github.com/yu314-coder/typhoon-predict/releases/tag/trackformer-1.2) and [Hugging Face](https://huggingface.co/euler314/typhoon-predict/tree/main/models/trackformer_1_2_field), not Git.
6
 
 
45
 
46
  The published 50-member mean is a separate evaluation policy: 50 deterministic seeds perturb only normalized historical basin/regional fields with smooth zero-centred noise, then average routes, central pressures and common-grid fields. `predict.py` computes one clean forecast; do not label it a 50-member mean.
47
 
48
+ **Matched development evaluation:** the released 1.1 and 1.2 route comparison uses the same frozen **1,473 daily starts / 270 equal-weight storms**, +6 to +120 h. The completed native-pressure comparison has **134 common starts / 40 storms**; the other 1,339 starts lack valid frozen 1.1 intensity inputs and are not zero-scored. The Site defaults to central-pressure MAE: **13.53 / 12.84 hPa** against JMA and **12.84 / 12.55 hPa** against USA (1.1 / 1.2), slight mean improvements with paired whole-storm uncertainty including zero. The optional curve-similarity diagnostic is **(1 + centred cosine) / 2 at exact common valid times**, without time shifting or warping. JMA similarity is **0.7074 / 0.7118**; USA is **0.7237 / 0.6637** on 131 eligible non-flat curves. Removing level/amplitude makes this a shape diagnostic, not proof of correct pressure levels. Actual unshifted hPa timelines, agency masks and uncertainty remain separate. See the [shared snapshot](../../evaluation/released_daily/released_daily_benchmark.json), [verification receipt](../../evaluation/released_daily/released_daily_verification.json), [pressure protocol](../../docs/intensity_benchmark.md) and [public benchmark API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/api/benchmarks/released). A new matched WeatherNext Cyclones Mini run is in progress on these exact daily starts; results remain pending. See the [frozen protocol](../../docs/deepmind_daily_benchmark.md); no old scores are transferred.
49
 
50
+ **Pressure-display coverage:** `regional_mslp_hpa` is the original fixed issue-relative composite, not an automatically extended moving-core map. When the forecast centre leaves that patch, a scalar central-pressure readout must not be inserted into the image to make it appear consistent. The Mangkhut film separately exports the actual moving model cores and registers member fields before averaging; [the current model card](https://huggingface.co/euler314/typhoon-predict) links the verified pressure arrays and links older fixed-patch films in a separate archive. The automatic History archive and its separate core recovery expose exact geography and coverage through [the public API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/history-api).
release_tools/deepmind_daily_benchmark.py ADDED
@@ -0,0 +1,483 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Frozen WeatherNext Cyclones Mini daily benchmark. No fitting or old scores.
2
+
3
+ The supervisor is resumable and CPU-only on the Mac. Each isolated worker runs
4
+ the official released predictor and tracker, using only two causal ERA5 states
5
+ and the frozen issue-time storm anchor. Truth is loaded afterwards for scoring.
6
+ All output, source caches and compilation caches remain on /Volumes/D.
7
+ """
8
+ from __future__ import annotations
9
+
10
+ import argparse
11
+ from collections import defaultdict
12
+ from datetime import datetime, timezone
13
+ import fcntl
14
+ import hashlib
15
+ import json
16
+ import os
17
+ from pathlib import Path
18
+ import subprocess
19
+ import sys
20
+ import time
21
+
22
+ import numpy as np
23
+
24
+ COHORT = '965184f1e5ab3fcb52e304f39b38583a14f753870aa2295f2d01942a441dcfd6'
25
+ WEIGHTS = 'a1bb151457077248d70a458b1a7b19deacd2926fbcb46eee4ff2691444e3596e'
26
+ CHECKPOINT = 'f194a23d3f91ea76ad776dfad942fabd669eeae8b3fd665815463095367e9ee0'
27
+ STEP_NS = 6 * 3600 * 10**9
28
+ HERE = Path(__file__).resolve()
29
+
30
+
31
+ def sha(path):
32
+ with Path(path).open('rb') as stream:
33
+ return hashlib.file_digest(stream, 'sha256').hexdigest()
34
+
35
+
36
+ def read(path):
37
+ return json.loads(Path(path).read_text())
38
+
39
+
40
+ def write(path, value):
41
+ path = Path(path)
42
+ temporary = path.with_suffix(path.suffix + '.tmp')
43
+ temporary.write_text(json.dumps(value, indent=2, allow_nan=False) + '\n')
44
+ temporary.replace(path)
45
+
46
+
47
+ def stamp():
48
+ return datetime.now(timezone.utc).isoformat()
49
+
50
+
51
+ def assets(project):
52
+ return {
53
+ 'weights': project/'v226/weathernext_global/weights/WeatherNextCyclones_Mini_<2024.npz',
54
+ 'shape_template': project/'v226/weathernext_global/data/hres_2024-10-07_00z_1deg_steps20.nc',
55
+ 'initial_state_reader': project/'v226_build_arco_weathernext_state.py',
56
+ 'inference_adapter': project/'v226_weathernext_cyclones_baseline.py',
57
+ 'scoring_definition': project/'benchmark_daily_storms_1_2.py',
58
+ 'runner': HERE,
59
+ }
60
+
61
+
62
+ def prepare(project, out):
63
+ if (out/'protocol.json').exists():
64
+ raise ValueError('Protocol already frozen; run or status, not prepare again')
65
+ original = project/'benchmark_daily_storms_1_2/cohort.json'
66
+ plan = read(original)
67
+ if sha(original) != COHORT or len(plan['cases']) != 1473 or plan['storm_count'] != 270:
68
+ raise ValueError('Wrong frozen daily cohort')
69
+ if plan['models']['1.2']['checkpoint_sha256'] != CHECKPOINT:
70
+ raise ValueError('Wrong Trackformer reference')
71
+ sources = {key: {'path': str(path), 'sha256': sha(path)} for key, path in assets(project).items()}
72
+ if sources['weights']['sha256'] != WEIGHTS:
73
+ raise ValueError('Wrong official WeatherNext Mini checkpoint')
74
+ rows = []
75
+ for case in plan['cases']:
76
+ path = project/'benchmark_daily_storms_1_2/cases'/f"{case['case_index']:05d}.npz"
77
+ metadata = read(path.with_suffix('.json'))
78
+ if any(metadata[key] != value for key, value in case.items()) or sha(path) != metadata['forecast_sha256']:
79
+ raise ValueError('Original daily forecast differs')
80
+ rows.append(dict(case, original_forecast_sha256=metadata['forecast_sha256']))
81
+ out.mkdir(parents=True, exist_ok=True)
82
+ vendor = subprocess.check_output(['git', '-C', str(project/'vendor/weathernext'), 'rev-parse', 'HEAD'], text=True).strip()
83
+ write(out/'protocol.json', {
84
+ 'schema': 'weathernext-mini-daily-equal-storm-v1', 'frozen_at_utc': stamp(),
85
+ 'cohort_sha256': COHORT, 'target_daily_cases': 1473, 'target_storms': 270,
86
+ 'model': 'Google DeepMind WeatherNext Cyclones Mini <2024', 'members': 1,
87
+ 'seed': 0, 'trained_through': 2023, 'resolution_degrees': 1,
88
+ 'official_weights_url': 'https://storage.googleapis.com/dm_graphcast/weathernext2/params/WeatherNextCyclones_Mini_%3C2024.npz',
89
+ 'vendor_commit': vendor, 'sources': sources, 'cases': rows,
90
+ 'leads_hours': list(range(6, 121, 6)), 'backend': 'JAX CPU',
91
+ 'causal_policy': 'Exactly issue -6h and issue ERA5 analyses; frozen observed issue centre only. Calendar forcings only after issue. No future weather, truth or official forecast route in inference.',
92
+ 'tracker_policy': 'Official direct_tracker_6h_v1 with initialized-storm continuity and no cyclogenesis; dissipation and nearby pruning disabled, disclosed rather than official operational tracker policy.',
93
+ 'aggregation': 'Mean valid leads per day, days within storm, then equal storm weights. All three direction scores recomputed on common moving steps. Pressure uses the original matched native intensity cases, common leads and separate JMA/USA labels.',
94
+ 'failure_policy': 'Failures stay visible and unscored. No zero fills, source substitutions, calibration or case selection by performance. Full-cohort completion requires all 1473 forecasts.',
95
+ 'interpretation': 'Development comparison, different input pipelines and member policies. Historical 1980-1999 cases overlap WeatherNext fitting years; recent 2024+ cases are separately reported. Not a certified untouched test.',
96
+ })
97
+ print(json.dumps({'status': 'protocol_frozen', 'cases': len(rows), 'storms': 270, 'protocol_sha256': sha(out/'protocol.json')}), flush=True)
98
+
99
+
100
+ def validate_state(path, case):
101
+ import xarray as xr
102
+ with xr.open_dataset(path, engine='h5netcdf') as state:
103
+ times = state.datetime.values.astype('datetime64[ns]').astype('int64')
104
+ expected = int(case['issue_ns']) + np.array([-STEP_NS, 0])
105
+ if times.shape != (2,) or not np.array_equal(times, expected):
106
+ raise ValueError('Weather initial state is not the exact causal two-analysis history')
107
+ for key in ('future_weather_fields_used', 'official_forecast_fields_used'):
108
+ if str(state.attrs.get(key)).lower() != 'false':
109
+ raise ValueError('Weather source does not declare causal analysis inputs')
110
+
111
+
112
+ def state_path(project, out, case):
113
+ tag = case['issue_time_utc'][:16].replace('-', '').replace(':', '').replace('T', '_')
114
+ filename = f'arco_era5_1deg_{tag}.nc'
115
+ for directory in (project/'benchmark_weathernext_cyclones_mini_same270/arco_states',
116
+ project/'benchmark_tip_ten/deepmind_states', out/'states'):
117
+ candidate = directory/filename
118
+ if candidate.exists():
119
+ validate_state(candidate, case)
120
+ return candidate
121
+ dest = out/'states'
122
+ subprocess.run([sys.executable, str(project/'v226_build_arco_weathernext_state.py'),
123
+ '--init', case['issue_time_utc'], '--output-dir', str(dest), '--skip-land-cache',
124
+ '--download-workers', '2', '--download-retries', '2'], check=True, cwd=project)
125
+ candidate = dest/filename
126
+ validate_state(candidate, case)
127
+ return candidate
128
+
129
+
130
+ def worker(project, out, index):
131
+ """Inference has no access to any future labels or original forecast outputs."""
132
+ import dataclasses
133
+ import haiku as hk
134
+ import jax
135
+ import pandas as pd
136
+ import xarray as xr
137
+ import xarray_jax
138
+ sys.path.insert(0, str(project))
139
+ import v226_weathernext_cyclones_baseline as adapter
140
+ from weathernext.utils import checkpoint, data_utils, fiddle_config_io, rollout
141
+ from weathernext.weathernext2 import fgn
142
+ from weathernext.cyclones import direct_tracker, direct_tracker_6h_v1_config
143
+ protocol = read(out/'protocol.json')
144
+ case = protocol['cases'][index]
145
+ initial = state_path(project, out, case)
146
+ config = fiddle_config_io.get_fiddle_config_by_name(adapter.CONFIG_NAME)
147
+ if jax.default_backend() != 'cpu':
148
+ raise ValueError('This frozen run requires its declared CPU backend')
149
+ jax.config.update('jax_compilation_cache_dir', str(out/'jax-cache'))
150
+ jax.config.update('jax_persistent_cache_min_compile_time_secs', 0)
151
+ attention = adapter.configure_attention(config, 'cpu')
152
+ paths = assets(project)
153
+ with paths['weights'].open('rb') as stream:
154
+ ckpt = checkpoint.load(stream, fgn.CheckPoint)
155
+ template = xr.load_dataset(paths['shape_template'], engine='h5netcdf').compute()
156
+ template, issue = adapter.inject_initial_state(template, initial, config.task)
157
+ if int(np.datetime64(issue, 'ns').astype('int64')) != case['issue_ns']:
158
+ raise ValueError('Issue does not match frozen cohort')
159
+ if set(config.task.forcing_variables) != adapter.ALLOWED_FORCINGS:
160
+ raise ValueError('Unexpected future forcing')
161
+ inputs, targets, forcings = data_utils.extract_inputs_targets_forcings(template,
162
+ target_lead_times=slice('6h', '120h'), **dataclasses.asdict(config.task))
163
+ if not np.array_equal(inputs.time.values.astype('timedelta64[h]').astype(int), [-6, 0]):
164
+ raise ValueError('Non-causal predictor input times')
165
+ prediction_config = fgn.PredictorConfig(task=config.task,
166
+ predictor_constructor=config.predictor_constructor,
167
+ predictor_kwargs=config.predictor_kwargs, predictor_wrappers=config.predictor_wrappers[:-1])
168
+
169
+ @hk.transform
170
+ def forward(history, shape, forcing):
171
+ return fgn.construct_predictor(prediction_config)(history, targets_template=shape, forcings=forcing)
172
+
173
+ compiled = jax.jit(lambda rng, history, shape, forcing: forward.apply(ckpt.params, rng, history, shape, forcing))
174
+ mapped = xarray_jax.pmap(compiled, dim='sample')
175
+ rngs = np.stack([jax.random.fold_in(jax.random.PRNGKey(protocol['seed']), 0)])
176
+ chunks = []
177
+ started = time.monotonic()
178
+ for i, chunk in enumerate(rollout.chunked_prediction_generator_multiple_runs(
179
+ predictor_fn=mapped, rngs=rngs, inputs=inputs,
180
+ targets_template=targets*np.nan, forcings=forcings,
181
+ num_steps_per_chunk=1, num_samples=1, pmap_devices=jax.local_devices())):
182
+ if i >= 20:
183
+ break
184
+ chunks.append(jax.device_get(chunk))
185
+ print(json.dumps({'case': index, 'step': i+1, 'steps': 20, 'elapsed_seconds': round(time.monotonic()-started, 1)}), flush=True)
186
+ predictions = xr.combine_by_coords(chunks).isel(time=slice(0, 20), batch=0, sample=0)
187
+ if predictions.sizes.get('time') != 20:
188
+ raise ValueError('Incomplete model rollout')
189
+ config_tracker = direct_tracker_6h_v1_config.get_config()
190
+ kwargs = dict(config_tracker.tracker_kwargs,
191
+ dissipation_min_mean_probability_threshold=None, prune_nearby_cyclones=False)
192
+ tracker = config_tracker.tracker_constructor(**kwargs)
193
+ scalar = predictions['cyclone_exists_gaussian_unit_mode']
194
+ for key in direct_tracker.IBTRACS_VARIABLE_NAME_TO_GRIDDED_VARIABLE_NAME.values():
195
+ if key not in predictions:
196
+ predictions[key] = xr.full_like(scalar, np.nan)
197
+ grid = tracker.preprocess_gridded_ds(predictions).expand_dims(forecast_datetime=[np.datetime64(issue)])
198
+ grid = grid.assign_coords(lead_time_secs=grid.time.astype('timedelta64[s]').astype(int),
199
+ date_time=np.datetime64(issue)+grid.time).as_numpy()
200
+ anchor = pd.DataFrame([dict(track_id=case['storm_id'], lat=case['base_lat'],
201
+ lon=case['base_lon'], valid_time=issue, init_time=issue)])
202
+ tracks = tracker(gridded_ds=grid, initial_storms_df=anchor, do_cyclogenesis=False)
203
+ tracks = tracks[tracks.track_id == case['storm_id']].copy()
204
+ wanted = pd.to_datetime(case['issue_ns'] + np.arange(1, 21)*STEP_NS)
205
+ rows = tracks.set_index('valid_time').reindex(wanted)
206
+ route = rows[['lat', 'lon']].to_numpy(dtype='float64')
207
+ route[:, 1] %= 360
208
+ if route.shape != (20, 2) or not np.isfinite(route).all():
209
+ raise ValueError('Tracker did not produce twenty finite exact-time points')
210
+ pressure = rows['minimum_sea_level_pressure_hpa'].to_numpy(dtype='float64')
211
+ pressure_valid = np.isfinite(pressure) & (pressure >= 800) & (pressure <= 1100)
212
+ native = predictions.mean_sea_level_pressure.sel(lat=slice(0, 60), lon=slice(100, 180)) / 100
213
+ values = native.transpose('time', 'lat', 'lon').values.astype('float32')
214
+ if values.shape != (20, 61, 81) or not np.isfinite(values).all() or values.min() < 800 or values.max() > 1100:
215
+ raise ValueError('Nonphysical or incomplete native Western Pacific pressure fields')
216
+ dest = out/'cases'/f'{index:05d}.npz'
217
+ temporary = dest.with_suffix('.npz.tmp')
218
+ with temporary.open('wb') as stream:
219
+ np.savez_compressed(stream, route_lat_lon=route,
220
+ central_pressure_hpa=np.where(pressure_valid, pressure, np.nan), pressure_valid=pressure_valid,
221
+ wp_pressure_hpa=values, latitude=native.lat.values, longitude=native.lon.values,
222
+ lead_hours=np.arange(6, 121, 6), valid_time_ns=case['issue_ns']+np.arange(1, 21)*STEP_NS)
223
+ temporary.replace(dest)
224
+ write(dest.with_suffix('.json'), dict(case, status='complete', members=1,
225
+ model=protocol['model'], checkpoint_sha256=WEIGHTS, protocol_sha256=sha(out/'protocol.json'),
226
+ forecast_sha256=sha(dest), input_state_sha256=sha(initial), input_state_path=str(initial),
227
+ input_time_ns=[case['issue_ns']-STEP_NS, case['issue_ns']], backend='cpu',
228
+ attention_override=attention, runtime_seconds=time.monotonic()-started, completed_at_utc=stamp(),
229
+ future_weather_used=False, future_storm_labels_used=False))
230
+
231
+
232
+ def route_metrics(routes, truth):
233
+ steps = [np.diff(np.vstack([np.zeros((1, 2)), p]), axis=0) for p in [*routes.values(), truth]]
234
+ valid = np.logical_and.reduce([np.linalg.norm(p, axis=-1) > 1 for p in steps])
235
+ result = {}
236
+ for (key, path), step in zip(routes.items(), steps[:-1]):
237
+ error = np.linalg.norm(path-truth, axis=-1)
238
+ angle = np.arctan2(step[:, 1], step[:, 0])-np.arctan2(steps[-1][:, 1], steps[-1][:, 0])
239
+ angle = np.degrees(np.abs(np.arctan2(np.sin(angle), np.cos(angle))))
240
+ a, b = (path-path.mean(0)).ravel(), (truth-truth.mean(0)).ravel()
241
+ similarity = (1+np.dot(a, b)/max(np.linalg.norm(a)*np.linalg.norm(b), 1e-8))/2
242
+ result[key] = dict(mean_track_error_km=float(error.mean()), track_error_120h_km=float(error[-1]),
243
+ track_error_by_lead_km=error.tolist(), direction_error_deg=float(angle[valid].mean()) if valid.any() else None,
244
+ direction_valid_steps=int(valid.sum()), shape_similarity=float(np.clip(similarity, 0, 1)))
245
+ return result
246
+
247
+
248
+ def local(route, lat, lon):
249
+ return np.stack([((route[:, 1]-lon+180)%360-180)*111.2*max(np.cos(np.deg2rad(lat)), .2),
250
+ (route[:, 0]-lat)*111.2], -1)
251
+
252
+
253
+ def curve_similarity(forecast, truth):
254
+ if len(truth) < 6:
255
+ return None
256
+ a, b = forecast-forecast.mean(), truth-truth.mean()
257
+ norm = np.linalg.norm(a)*np.linalg.norm(b)
258
+ return float(np.clip((1+np.dot(a, b)/norm)/2, 0, 1)) if norm > 1e-8 else None
259
+
260
+
261
+ def score(project, out, case):
262
+ index = case['case_index']
263
+ path = out/'cases'/f'{index:05d}.npz'
264
+ record = read(path.with_suffix('.json'))
265
+ if record['forecast_sha256'] != sha(path) or record['protocol_sha256'] != sha(out/'protocol.json'):
266
+ raise ValueError('Saved candidate identity changed')
267
+ original = project/'benchmark_daily_storms_1_2/cases'/path.name
268
+ if sha(original) != case['original_forecast_sha256']:
269
+ raise ValueError('Frozen baseline forecast changed')
270
+ with np.load(path, allow_pickle=False) as z, np.load(original, allow_pickle=False) as reference:
271
+ candidate = local(z['route_lat_lon'], case['base_lat'], case['base_lon'])
272
+ metrics = route_metrics({'1.1': reference['v11_local'], '1.2': reference['v12_local'], 'deepmind': candidate}, reference['truth_local'])
273
+ pressure = z['central_pressure_hpa'].copy()
274
+ source = read(project/'output/intensity-v12-v11-20261001/cases'/f'{index:05d}.json')
275
+ pressure_scores = {}
276
+ if source['v11'] is not None:
277
+ curves = {'1.1': np.array([p['central_pressure_hpa'] for p in source['v11']], dtype=float),
278
+ '1.2': np.array(source['v12_pressure_hpa'], dtype=float), 'deepmind': pressure}
279
+ for agency in ('JMA', 'USA'):
280
+ truth = np.array(source['truth']['pressure_'+agency+'_hpa'], dtype=float)
281
+ mask = pressure_mask(truth, curves)
282
+ pressure_scores[agency] = {'valid_leads': int(mask.sum()), 'models': {
283
+ key: {'mae_hpa': float(np.abs(p[mask]-truth[mask]).mean()) if mask.any() else None,
284
+ 'curve_similarity': curve_similarity(p[mask], truth[mask]),
285
+ 'error_by_lead_hpa': [float(abs(p[i]-truth[i])) if mask[i] else None for i in range(20)]}
286
+ for key, p in curves.items()}}
287
+ record.update(route_metrics=metrics, pressure_metrics=pressure_scores)
288
+ write(path.with_suffix('.json'), record)
289
+ return record
290
+
291
+
292
+ def pressure_mask(truth, curves):
293
+ arrays = [np.asarray(p, dtype=float) for p in [truth, *curves.values()]]
294
+ if any(p.shape != (20,) for p in arrays):
295
+ raise ValueError('Pressure comparison requires twenty matching exact leads')
296
+ return np.logical_and.reduce([np.isfinite(p) & (p >= 800) & (p <= 1100) for p in arrays])
297
+
298
+
299
+ def aggregate(groups):
300
+ scores = {}
301
+ for model in ('1.1', '1.2', 'deepmind'):
302
+ scores[model] = {}
303
+ for metric in ('mean_track_error_km', 'track_error_120h_km', 'direction_error_deg', 'shape_similarity'):
304
+ storms = []
305
+ for days in groups.values():
306
+ values = [d['route_metrics'][model][metric] for d in days if d['route_metrics'][model][metric] is not None]
307
+ if values:
308
+ storms.append(float(np.mean(values)))
309
+ scores[model][metric] = {'value': float(np.mean(storms)) if storms else None, 'storms': len(storms)}
310
+ scores[model]['track_error_by_lead_km'] = np.mean([
311
+ np.mean([d['route_metrics'][model]['track_error_by_lead_km'] for d in days], axis=0)
312
+ for days in groups.values()], axis=0).tolist() if groups else [None]*20
313
+ pressure = {}
314
+ for agency in ('JMA', 'USA'):
315
+ eligible = {sid: [d['pressure_metrics'][agency] for d in days
316
+ if agency in d['pressure_metrics'] and d['pressure_metrics'][agency]['valid_leads']]
317
+ for sid, days in groups.items()}
318
+ eligible = {sid: days for sid, days in eligible.items() if days}
319
+ pressure[agency] = {'daily_cases': sum(map(len, eligible.values())), 'storms': len(eligible),
320
+ 'common_lead_points': sum(d['valid_leads'] for days in eligible.values() for d in days), 'models': {}}
321
+ for model in ('1.1', '1.2', 'deepmind'):
322
+ values = {}
323
+ for metric in ('mae_hpa', 'curve_similarity'):
324
+ storm_values = []
325
+ case_count = 0
326
+ for days in eligible.values():
327
+ valid = [d['models'][model][metric] for d in days if d['models'][model][metric] is not None]
328
+ if valid:
329
+ storm_values.append(float(np.mean(valid)))
330
+ case_count += len(valid)
331
+ values[metric] = {'value': float(np.mean(storm_values)) if storm_values else None,
332
+ 'storms': len(storm_values), 'daily_cases': case_count}
333
+ pressure[agency]['models'][model] = values
334
+ return {'route': scores, 'pressure': pressure}
335
+
336
+
337
+ def status(project, out, active=None):
338
+ protocol = read(out/'protocol.json')
339
+ groups = defaultdict(list)
340
+ expected = defaultdict(int)
341
+ failures = []
342
+ for case in protocol['cases']:
343
+ expected[case['storm_id']] += 1
344
+ path = out/'cases'/f"{case['case_index']:05d}.json"
345
+ completed = False
346
+ if path.exists():
347
+ record = read(path)
348
+ if 'route_metrics' in record:
349
+ groups[case['storm_id']].append(record)
350
+ completed = True
351
+ failure = out/'failures'/f"{case['case_index']:05d}.json"
352
+ if failure.exists() and not completed:
353
+ failures.append(read(failure))
354
+ full = {key: days for key, days in groups.items() if len(days) == expected[key]}
355
+ done = sum(map(len, groups.values()))
356
+ receipt = read(out/'verification.json') if (out/'verification.json').exists() else {}
357
+ verified = (receipt.get('verified_daily_cases') == 1473 and receipt.get('protocol_sha256') == sha(out/'protocol.json'))
358
+ report = {'state': 'complete_verified' if verified else 'forecasts_finished_audit_pending' if done == 1473 else 'running_partial' if active is not False else 'finished_with_unavailable',
359
+ 'updated_at_utc': stamp(), 'cohort_sha256': COHORT, 'protocol_sha256': sha(out/'protocol.json'),
360
+ 'completed_daily_cases': done, 'target_daily_cases': 1473, 'completed_storms': len(full),
361
+ 'target_storms': 270, 'failed_daily_cases': len(failures), 'failures': failures,
362
+ 'checkpoint_sha256': WEIGHTS, 'members': 1, 'backend': 'JAX CPU',
363
+ 'aggregate_equal_complete_storm': aggregate(full), 'final_results': verified,
364
+ 'periods': {period: aggregate({sid: days for sid, days in full.items() if days[0]['period'] == period})
365
+ for period in sorted({c['period'] for c in protocol['cases']})},
366
+ 'original_forecasts_modified': False, 'original_model_weights_modified': False,
367
+ 'missing_results_scored_as_zero': 0, 'pressure_note': 'Per-case three-way common-support pressure MAE and curve similarity in cases/*.json; separate agency masks. Final publication requires coverage audit, including pressure-specific coverage.'}
368
+ write(out/'progress.json', report)
369
+ return report
370
+
371
+
372
+ def verify(project, out):
373
+ protocol = read(out/'protocol.json')
374
+ verified = []
375
+ for case in protocol['cases']:
376
+ path = out/'cases'/f"{case['case_index']:05d}.npz"
377
+ record = read(path.with_suffix('.json'))
378
+ if any(record.get(k) != v for k, v in case.items()):
379
+ raise ValueError('Forecast case does not match the frozen definition')
380
+ if record.get('checkpoint_sha256') != WEIGHTS or record.get('members') != 1:
381
+ raise ValueError('Forecast model/member identity differs')
382
+ if record.get('input_time_ns') != [case['issue_ns']-STEP_NS, case['issue_ns']]:
383
+ raise ValueError('Weather history is noncausal')
384
+ with np.load(path, allow_pickle=False) as z:
385
+ if not np.array_equal(z['lead_hours'], np.arange(6, 121, 6)) or not np.array_equal(
386
+ z['valid_time_ns'], case['issue_ns']+np.arange(1, 21)*STEP_NS):
387
+ raise ValueError('Forecast leads are not exact +6 to +120h')
388
+ if z['route_lat_lon'].shape != (20, 2) or not np.isfinite(z['route_lat_lon']).all():
389
+ raise ValueError('Incomplete route')
390
+ p = z['wp_pressure_hpa']
391
+ if p.shape != (20, 61, 81) or not np.isfinite(p).all() or p.min() < 800 or p.max() > 1100:
392
+ raise ValueError('Incomplete physical native pressure fields')
393
+ score(project, out, case) # Recheck hashes and recompute common-support metrics.
394
+ verified.append(case['case_index'])
395
+ if len(set(verified)) != 1473:
396
+ raise ValueError('Audit did not cover every planned case')
397
+ write(out/'verification.json', {'status': 'complete_verified', 'verified_at_utc': stamp(),
398
+ 'protocol_sha256': sha(out/'protocol.json'), 'cohort_sha256': COHORT,
399
+ 'checkpoint_sha256': WEIGHTS, 'verified_daily_cases': 1473, 'verified_storms': 270,
400
+ 'case_indices': verified, 'checks': ['exact frozen issue identity', 'forecast and baseline SHA256',
401
+ 'one official Mini member', 'causal two-analysis history', 'twenty exact leads',
402
+ 'finite unshifted routes', 'twenty native physical WP pressure grids',
403
+ 'recomputed matched three-way route and pressure masks']})
404
+
405
+
406
+ def run(project, out, limit=0):
407
+ out.mkdir(parents=True, exist_ok=True)
408
+ lock = (out/'run.lock').open('a')
409
+ fcntl.flock(lock, fcntl.LOCK_EX | fcntl.LOCK_NB)
410
+ protocol = read(out/'protocol.json')
411
+ for row in protocol['sources'].values():
412
+ if sha(row['path']) != row['sha256']:
413
+ raise ValueError('Frozen source changed: '+row['path'])
414
+ for name in ('cases', 'failures', 'logs', 'states', 'jax-cache'):
415
+ (out/name).mkdir(exist_ok=True)
416
+ write(out/'runtime.json', {'pid': os.getpid(), 'started_at_utc': stamp(), 'backend': 'JAX CPU',
417
+ 'python': sys.executable, 'protocol_sha256': sha(out/'protocol.json')})
418
+ count = 0
419
+ status(project, out, True)
420
+ for case in protocol['cases']:
421
+ index = case['case_index']
422
+ saved = out/'cases'/f'{index:05d}.json'
423
+ if saved.exists() and 'route_metrics' in read(saved):
424
+ continue
425
+ if limit and count >= limit:
426
+ break
427
+ count += 1
428
+ try:
429
+ if not saved.exists():
430
+ with (out/'logs'/f'{index:05d}.log').open('a') as log:
431
+ subprocess.run([sys.executable, str(HERE), 'worker', '--project', str(project),
432
+ '--output', str(out), '--case-index', str(index)], cwd=project,
433
+ env=dict(os.environ, JAX_PLATFORMS='cpu', PYTHONUNBUFFERED='1',
434
+ XLA_PYTHON_CLIENT_PREALLOCATE='false'), stdout=log, stderr=subprocess.STDOUT, check=True)
435
+ score(project, out, case)
436
+ print(json.dumps({'case_complete': index, 'storm_id': case['storm_id'], 'issue': case['issue_time_utc']}), flush=True)
437
+ except Exception as error:
438
+ write(out/'failures'/f'{index:05d}.json', dict(case_index=index, storm_id=case['storm_id'],
439
+ issue_time_utc=case['issue_time_utc'], error=str(error), at_utc=stamp(), result=None))
440
+ print(json.dumps({'case_failed': index, 'error': str(error)}), flush=True)
441
+ status(project, out, True)
442
+ report = status(project, out, False)
443
+ if report['completed_daily_cases'] == 1473:
444
+ verify(project, out)
445
+ report = status(project, out, False)
446
+ print(json.dumps({k: v for k, v in report.items() if k not in ('failures', 'aggregate_equal_complete_storm')}), flush=True)
447
+
448
+
449
+ def launch(project, out):
450
+ # Lock is acquired by run; launch never stops or replaces an existing job.
451
+ with (out/'run.lock').open('a') as lock:
452
+ fcntl.flock(lock, fcntl.LOCK_EX | fcntl.LOCK_NB)
453
+ log = (out/'run.log').open('a')
454
+ command = [sys.executable, str(HERE), 'run', '--project', str(project), '--output', str(out)]
455
+ if sys.platform == 'darwin':
456
+ command = ['/usr/bin/caffeinate', '-i', '-s', *command]
457
+ child = subprocess.Popen(command, cwd=project, stdin=subprocess.DEVNULL,
458
+ stdout=log, stderr=subprocess.STDOUT, start_new_session=True,
459
+ env=dict(os.environ, JAX_PLATFORMS='cpu', PYTHONUNBUFFERED='1', XLA_PYTHON_CLIENT_PREALLOCATE='false'))
460
+ write(out/'launcher.json', {'pid': child.pid, 'launched_at_utc': stamp(), 'command': command})
461
+ print(json.dumps({'launched_pid': child.pid, 'output': str(out), 'backend': 'JAX CPU'}), flush=True)
462
+
463
+
464
+ def main():
465
+ parser = argparse.ArgumentParser(description=__doc__)
466
+ parser.add_argument('mode', choices=('prepare', 'run', 'worker', 'status', 'launch'))
467
+ parser.add_argument('--project', type=Path, required=True)
468
+ parser.add_argument('--output', type=Path, required=True)
469
+ parser.add_argument('--case-index', type=int, default=0)
470
+ parser.add_argument('--limit', type=int, default=0)
471
+ args = parser.parse_args()
472
+ project, out = args.project.resolve(), args.output.resolve()
473
+ if not out.is_relative_to(Path('/Volumes/D')) or not project.is_relative_to(Path('/Volumes/D')):
474
+ raise ValueError('Project and all generated artifacts must remain on D')
475
+ if args.mode == 'prepare': prepare(project, out)
476
+ elif args.mode == 'worker': worker(project, out, args.case_index)
477
+ elif args.mode == 'run': run(project, out, args.limit)
478
+ elif args.mode == 'launch': launch(project, out)
479
+ else: print(json.dumps(status(project, out), indent=2))
480
+
481
+
482
+ if __name__ == '__main__':
483
+ main()
release_tools/plot_model_announcement.py ADDED
@@ -0,0 +1,67 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Three release metrics from frozen scores; pending models have no fake bars."""
2
+ import hashlib
3
+ import json
4
+ from pathlib import Path
5
+ import matplotlib
6
+ matplotlib.use('Agg')
7
+ import matplotlib.pyplot as plt
8
+
9
+ ROOT = Path(__file__).resolve().parents[1]
10
+ SOURCE = ROOT/'evaluation/released_daily/released_daily_benchmark.json'
11
+ DEST = ROOT/'evaluation/released_daily/model_1_2_benchmark.png'
12
+
13
+
14
+ def main():
15
+ source = json.loads(SOURCE.read_text())
16
+ if source['status'] != 'complete_verified' or source['cohort']['daily_issues'] != 1473:
17
+ raise ValueError('Chart requires verified daily-cohort results')
18
+ metrics = {m['key']: m for m in source['metrics']}
19
+ plt.rcParams.update({'font.family': 'DejaVu Sans', 'font.size': 11,
20
+ 'axes.spines.top': False, 'axes.spines.right': False,
21
+ 'axes.spines.left': False, 'axes.spines.bottom': False})
22
+ fig, axes = plt.subplots(1, 3, figsize=(13.8, 5.7))
23
+ fig.set_facecolor('#f6f8fc')
24
+ specs = [('pressure_JMA_hpa', 'Pressure intensity', 'Central-pressure MAE · hPa', 18, 2),
25
+ ('mean_track_error_km', 'Track position', 'Mean position error · km', 1050, 1),
26
+ ('direction_error_deg', 'Track direction', 'Six-hour heading error · degrees', 65, 2)]
27
+ for ax, (key, title, label, maximum, digits) in zip(axes, specs):
28
+ row = metrics[key]
29
+ values = [row['values'][k] for k in ('1.1', '1.2')]
30
+ ax.set_facecolor('#f6f8fc')
31
+ ax.set_axisbelow(True)
32
+ ax.grid(axis='y', color='#dfe5ee', linewidth=.8)
33
+ ax.bar([0, 1], values, width=.53, color=['#9aa9bc', '#3c67d6'], zorder=3)
34
+ if key == 'pressure_JMA_hpa':
35
+ intervals = [row['intervals'][k] for k in ('1.1', '1.2')]
36
+ ax.errorbar([0, 1], values, yerr=[[v-low for v, (low, high) in zip(values, intervals)],
37
+ [high-v for v, (low, high) in zip(values, intervals)]],
38
+ fmt='none', color='#253853', capsize=5, linewidth=1.2, zorder=4)
39
+ for x, value in enumerate(values):
40
+ top = row['intervals'][('1.1', '1.2')[x]][1] if key == 'pressure_JMA_hpa' else value
41
+ ax.text(x, top+maximum*.038, f'{value:,.{digits}f}', ha='center',
42
+ fontsize=17, fontweight='bold', color='#21334c')
43
+ ax.set_ylim(0, maximum)
44
+ ax.set_xticks([0, 1], ['1.1', '1.2 · mean of 50'])
45
+ ax.tick_params(axis='both', length=0, labelcolor='#53647c')
46
+ ax.set_title(title, loc='left', fontsize=16, fontweight='bold', color='#21334c', pad=38)
47
+ ax.text(0, 1.03, label, transform=ax.transAxes, color='#53647c', fontsize=10)
48
+ coverage = row['coverage']
49
+ ax.text(.5, -.16, f"{coverage['daily_issues']:,} daily starts / {coverage['storms']} storms",
50
+ transform=ax.transAxes, ha='center', color='#53647c', fontsize=10)
51
+ fig.text(.055, .945, 'Trackformer 1.2', fontsize=23, fontweight='bold', color='#182b45')
52
+ fig.text(.055, .895, 'Matched five-day forecasts · equal weight per storm · lower is better', fontsize=12, color='#53647c')
53
+ fig.text(.055, .095, 'Pressure: JMA reference; bars show 95% whole-storm intervals. Paired improvement is uncertain.', fontsize=10, color='#53647c')
54
+ fig.text(.055, .057, 'DeepMind Mini: matched run in progress — no score yet. Different pipelines/member policies; development evidence.', fontsize=10, color='#53647c')
55
+ fig.subplots_adjust(left=.055, right=.975, top=.735, bottom=.25, wspace=.36)
56
+ fig.savefig(DEST, dpi=180, facecolor=fig.get_facecolor())
57
+ plt.close(fig)
58
+ receipt = {'source': SOURCE.relative_to(ROOT).as_posix(), 'source_sha256': hashlib.sha256(SOURCE.read_bytes()).hexdigest(),
59
+ 'chart': DEST.relative_to(ROOT).as_posix(), 'chart_sha256': hashlib.sha256(DEST.read_bytes()).hexdigest(),
60
+ 'metrics': [s[0] for s in specs], 'values': {k: metrics[k]['values'] for k, *_ in specs},
61
+ 'deepmind_score': None, 'zero_fill': False, 'model_forecasts_modified': False}
62
+ DEST.with_suffix('.json').write_text(json.dumps(receipt, indent=2)+'\n')
63
+ print(json.dumps(receipt, indent=2))
64
+
65
+
66
+ if __name__ == '__main__':
67
+ main()
release_tools/sync_public_model_cards.py CHANGED
@@ -1,6 +1,6 @@
1
  """Render the complete HF card from the canonical GitHub README.
2
 
3
- Preserve remote metadata, four inline films and frozen neural weights. Default
4
  mode only prepares artifacts on D; --publish explicitly updates the HF card,
5
  matching existing inference wrapper/docs and this small reproduction utility.
6
  No forecasts are rerun, no media are regenerated and no neural modules change.
@@ -20,11 +20,10 @@ REPO = 'euler314/typhoon-predict'
20
  MEDIA_REVISION = '87a6e366b42bb4cc0edc95d2c50c55fca21a2c93'
21
  CACHE = '/Volumes/D/typhoon_predict/.cache/huggingface-intensity'
22
  HF = 'https://huggingface.co/' + REPO
23
- GALLERY = '| Soudelor (2015) | Mangkhut (2018) | Meranti (2016) |\n'
24
- FILMS = ('mangkhut', 'fung_wong', 'soudelor', 'meranti')
25
  CURRENT_FIGURES = frozenset((
26
- 'evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png',
27
- 'evaluation/released_daily/pressure_comparison.png',
28
  ))
29
  SYNC_FILES = (
30
  'models/trackformer_1_2_field/README.md',
@@ -36,9 +35,15 @@ SYNC_FILES = (
36
  'release_tools/build_release_benchmark.py',
37
  'release_tools/plot_release_pressure_benchmark.py',
38
  'release_tools/plot_daily_storm_final.py',
 
 
 
39
  'docs/daily_storm_benchmark.md',
40
  'docs/intensity_benchmark.md',
41
  'docs/trackformer_1_2_evaluation.md',
 
 
 
42
  'evaluation/daily_storm_final.json',
43
  'evaluation/released_daily/released_daily_benchmark.json',
44
  'evaluation/released_daily/released_daily_verification.json',
@@ -46,8 +51,8 @@ SYNC_FILES = (
46
  'evaluation/intensity/verification.json',
47
  'evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png',
48
  'evaluation/released_daily/pressure_comparison.png',
49
- 'paper/trackformer.tex',
50
- 'paper/trackformer.pdf',
51
  )
52
 
53
 
@@ -73,38 +78,14 @@ def render_card(original, github, figure_revision='main'):
73
  metadata = re.match(r'\A---\n.*?\n---\n', original, re.S)
74
  if not metadata:
75
  raise ValueError('Preserve existing HF YAML metadata; missing header')
76
- if github.count('## Corrected pressure forecast — Mangkhut (2018)\n') != 1:
77
- raise ValueError('Missing or ambiguous corrected-film section')
78
- if github.count(GALLERY) != 1:
79
- raise ValueError('Missing or ambiguous historical video gallery')
80
  card = github
81
- for stem, alt in (
82
- ('mangkhut', 'Corrected Mangkhut moving-core pressure forecast — 1 October 2026 revision'),
83
- ('fung_wong', 'Play the Trackformer 1.2 Fung-wong pressure forecast MP4'),
84
- ):
85
  old = '[![' + alt + '](' + media_url(stem, True) + ')](' + media_url(stem) + ')'
86
  if card.count(old) != 1:
87
  raise ValueError('Missing or ambiguous primary preview: ' + stem)
88
  card = card.replace(old, player(stem, alt), 1)
89
- start = card.index(GALLERY)
90
- end = card.index('\n\n', start)
91
- original_gallery = card[start:end]
92
- if len(original_gallery.splitlines()) != 5:
93
- raise ValueError('Unexpected gallery; preserve existing documentation')
94
- replacements = []
95
- for stem, name, issue in (
96
- ('soudelor', 'Soudelor (2015)', '5 August 2015, 00 UTC'),
97
- ('mangkhut', 'Mangkhut (2018)', '11 September 2018, 00 UTC'),
98
- ('meranti', 'Meranti (2016)', '10 September 2016, 00 UTC'),
99
- ):
100
- replacements.append('### ' + name + ' — issue ' + issue)
101
- if stem == 'mangkhut':
102
- replacements.append('The corrected 20-second player appears near the top of this card; its moving-core repair and audit are described there.')
103
- else:
104
- replacements.append(player(stem, name + ' original fixed-patch film'))
105
- replacements.append('[Play / download ' + name.split(' (')[0] + ' MP4](' + media_url(stem) + ') · [Provenance](evaluation/release_data/' + stem + '_video.json)')
106
- card = card[:start] + '\n\n'.join(replacements) + card[end:]
107
-
108
  # Relative GitHub links need HF-specific URLs. Keep evolving documentation
109
  # on main; pin unchanged scientific image bytes to the verified asset commit.
110
  def link(match):
@@ -121,12 +102,12 @@ def render_card(original, github, figure_revision='main'):
121
  return image + '[' + label + '](' + target + ')'
122
  card = re.sub(r'(!?)\[([^\[\]\n]*)\]\(([^)\s]+)\)', link, card)
123
  card = metadata.group(0) + '\n' + card
124
- if card.count('<video ') != 4:
125
- raise ValueError('Expected exactly one native player for each of four films')
126
  for stem in FILMS:
127
  if card.count('src="' + media_url(stem) + '"') != 1:
128
  raise ValueError('Missing, repeated or stale player: ' + stem)
129
- if '/resolve/main/docs/trackformer_1_2_' in card:
130
  raise ValueError('Unversioned movie source')
131
  return card
132
 
@@ -171,7 +152,7 @@ def run(output, publish=False):
171
  'github_readme_sha256': sha(ROOT / 'README.md'),
172
  'original_hf_readme_sha256': sha(original),
173
  'files': {name: sha(path) for name, path in files.items()},
174
- 'media_revision': MEDIA_REVISION, 'native_video_players': 4,
175
  'inference_weights_sha256': weights_before,
176
  'weights_changed': False, 'neural_modules_changed': False,
177
  'forecast_arrays_changed': False, 'videos_changed': False,
@@ -184,13 +165,13 @@ def run(output, publish=False):
184
  # immutable commit in the card. Never pin new bytes to the old movie
185
  # revision or let a missing future image render as a stale figure.
186
  assets = api.create_commit(repo_id=REPO, repo_type='model', parent_commit=parent,
187
- commit_message='Publish verified matched pressure-curve benchmark and revised technical paper',
188
  operations=[CommitOperationAdd(path_in_repo=name, path_or_fileobj=str(ROOT/name)) for name in SYNC_FILES])
189
  card_path.write_text(render_card(original.read_text(), github, assets.oid))
190
  receipt['files']['README.md'] = sha(card_path)
191
  receipt['figure_revision'] = assets.oid
192
  committed = api.create_commit(repo_id=REPO, repo_type='model', parent_commit=assets.oid,
193
- commit_message='Refresh model card with exact-time pressure graph similarity and pinned figures',
194
  operations=[CommitOperationAdd(path_in_repo='README.md', path_or_fileobj=str(card_path))])
195
  for name, expected in receipt['files'].items():
196
  saved = hf_hub_download(REPO, name, revision=committed.oid, cache_dir=CACHE)
 
1
  """Render the complete HF card from the canonical GitHub README.
2
 
3
+ Preserve remote metadata, the featured film and frozen neural weights. Default
4
  mode only prepares artifacts on D; --publish explicitly updates the HF card,
5
  matching existing inference wrapper/docs and this small reproduction utility.
6
  No forecasts are rerun, no media are regenerated and no neural modules change.
 
20
  MEDIA_REVISION = '87a6e366b42bb4cc0edc95d2c50c55fca21a2c93'
21
  CACHE = '/Volumes/D/typhoon_predict/.cache/huggingface-intensity'
22
  HF = 'https://huggingface.co/' + REPO
23
+ FILMS = ('mangkhut',)
 
24
  CURRENT_FIGURES = frozenset((
25
+ 'evaluation/released_daily/model_1_2_benchmark.png',
26
+ 'docs/trackformer_1_2_architecture.svg',
27
  ))
28
  SYNC_FILES = (
29
  'models/trackformer_1_2_field/README.md',
 
35
  'release_tools/build_release_benchmark.py',
36
  'release_tools/plot_release_pressure_benchmark.py',
37
  'release_tools/plot_daily_storm_final.py',
38
+ 'release_tools/plot_model_announcement.py',
39
+ 'release_tools/deepmind_daily_benchmark.py',
40
+ 'release_tools/test_deepmind_daily_benchmark.py',
41
  'docs/daily_storm_benchmark.md',
42
  'docs/intensity_benchmark.md',
43
  'docs/trackformer_1_2_evaluation.md',
44
+ 'docs/deepmind_daily_benchmark.md',
45
+ 'docs/showcase_archive.md',
46
+ 'docs/trackformer_1_2_architecture.svg',
47
  'evaluation/daily_storm_final.json',
48
  'evaluation/released_daily/released_daily_benchmark.json',
49
  'evaluation/released_daily/released_daily_verification.json',
 
51
  'evaluation/intensity/verification.json',
52
  'evaluation/trackformer_1_2_vs_1_1_270_storms_bars.png',
53
  'evaluation/released_daily/pressure_comparison.png',
54
+ 'evaluation/released_daily/model_1_2_benchmark.png',
55
+ 'evaluation/released_daily/model_1_2_benchmark.json',
56
  )
57
 
58
 
 
78
  metadata = re.match(r'\A---\n.*?\n---\n', original, re.S)
79
  if not metadata:
80
  raise ValueError('Preserve existing HF YAML metadata; missing header')
81
+ if github.count('## See the forecast: Mangkhut (2018)\n') != 1:
82
+ raise ValueError('Missing or ambiguous featured-film section')
 
 
83
  card = github
84
+ for stem, alt in (('mangkhut', 'Trackformer 1.2 Mangkhut pressure and route forecast'),):
 
 
 
85
  old = '[![' + alt + '](' + media_url(stem, True) + ')](' + media_url(stem) + ')'
86
  if card.count(old) != 1:
87
  raise ValueError('Missing or ambiguous primary preview: ' + stem)
88
  card = card.replace(old, player(stem, alt), 1)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
89
  # Relative GitHub links need HF-specific URLs. Keep evolving documentation
90
  # on main; pin unchanged scientific image bytes to the verified asset commit.
91
  def link(match):
 
102
  return image + '[' + label + '](' + target + ')'
103
  card = re.sub(r'(!?)\[([^\[\]\n]*)\]\(([^)\s]+)\)', link, card)
104
  card = metadata.group(0) + '\n' + card
105
+ if card.count('<video ') != 1:
106
+ raise ValueError('Expected exactly one featured native player')
107
  for stem in FILMS:
108
  if card.count('src="' + media_url(stem) + '"') != 1:
109
  raise ValueError('Missing, repeated or stale player: ' + stem)
110
+ if re.search(r'/resolve/main/docs/trackformer_1_2_[^"\s)]+\.mp4', card):
111
  raise ValueError('Unversioned movie source')
112
  return card
113
 
 
152
  'github_readme_sha256': sha(ROOT / 'README.md'),
153
  'original_hf_readme_sha256': sha(original),
154
  'files': {name: sha(path) for name, path in files.items()},
155
+ 'media_revision': MEDIA_REVISION, 'native_video_players': 1,
156
  'inference_weights_sha256': weights_before,
157
  'weights_changed': False, 'neural_modules_changed': False,
158
  'forecast_arrays_changed': False, 'videos_changed': False,
 
165
  # immutable commit in the card. Never pin new bytes to the old movie
166
  # revision or let a missing future image render as a stale figure.
167
  assets = api.create_commit(repo_id=REPO, repo_type='model', parent_commit=parent,
168
+ commit_message='Publish Trackformer 1.2 announcement assets and matched DeepMind benchmark protocol',
169
  operations=[CommitOperationAdd(path_in_repo=name, path_or_fileobj=str(ROOT/name)) for name in SYNC_FILES])
170
  card_path.write_text(render_card(original.read_text(), github, assets.oid))
171
  receipt['files']['README.md'] = sha(card_path)
172
  receipt['figure_revision'] = assets.oid
173
  committed = api.create_commit(repo_id=REPO, repo_type='model', parent_commit=assets.oid,
174
+ commit_message='Introduce Trackformer 1.2 with a focused showcase and verified daily metrics',
175
  operations=[CommitOperationAdd(path_in_repo='README.md', path_or_fileobj=str(card_path))])
176
  for name, expected in receipt['files'].items():
177
  saved = hf_hub_download(REPO, name, revision=committed.oid, cache_dir=CACHE)
release_tools/test_deepmind_daily_benchmark.py ADDED
@@ -0,0 +1,60 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Small offline checks; never import the weather model or run inference."""
2
+ import ast
3
+ from pathlib import Path
4
+ import unittest
5
+ import numpy as np
6
+ from deepmind_daily_benchmark import aggregate, curve_similarity, local, pressure_mask, route_metrics
7
+
8
+
9
+ class DeepMindBenchmarkTest(unittest.TestCase):
10
+ def test_identical_routes_have_zero_error_and_matching_shape(self):
11
+ truth = np.stack([np.arange(1, 21)*3, np.arange(1, 21)*4], -1)
12
+ metrics = route_metrics({k: truth.copy() for k in ('1.1', '1.2', 'deepmind')}, truth)
13
+ for result in metrics.values():
14
+ self.assertAlmostEqual(result['mean_track_error_km'], 0)
15
+ self.assertAlmostEqual(result['direction_error_deg'], 0)
16
+ self.assertAlmostEqual(result['shape_similarity'], 1)
17
+
18
+ def test_direction_mask_is_common_to_all_three_models(self):
19
+ path = np.stack([np.arange(1, 21)*10, np.zeros(20)], -1)
20
+ stationary = np.zeros((20, 2))
21
+ results = route_metrics({'1.1': path, '1.2': path, 'deepmind': stationary}, path)
22
+ self.assertTrue(all(r['direction_valid_steps'] == 0 and r['direction_error_deg'] is None for r in results.values()))
23
+
24
+ def test_curve_shape_is_not_pressure_level_error(self):
25
+ truth = np.arange(20, dtype=float)+950
26
+ self.assertAlmostEqual(curve_similarity(truth+20, truth), 1)
27
+ self.assertIsNone(curve_similarity(np.ones(20), truth))
28
+ self.assertIsNone(curve_similarity(truth[:5], truth[:5]))
29
+
30
+ def test_missing_and_invalid_pressure_never_become_zero_errors(self):
31
+ truth = np.full(20, 950.)
32
+ one, two, candidate = truth.copy(), truth.copy(), truth.copy()
33
+ one[0], two[1], candidate[2] = np.nan, 0, 700
34
+ mask = pressure_mask(truth, {'1.1': one, '1.2': two, 'deepmind': candidate})
35
+ self.assertEqual(mask.sum(), 17)
36
+ self.assertFalse(mask[:3].any())
37
+ with self.assertRaises(ValueError):
38
+ pressure_mask(truth, {'model': truth[:19]})
39
+
40
+ def test_longitude_wrap_keeps_geographic_alignment(self):
41
+ result = local(np.array([[0., 1.]]), 0., 359.)
42
+ self.assertAlmostEqual(result[0, 0], 222.4)
43
+
44
+ def test_empty_coverage_is_null_not_a_zero_score(self):
45
+ results = aggregate({})
46
+ self.assertIsNone(results['route']['deepmind']['mean_track_error_km']['value'])
47
+ self.assertIsNone(results['pressure']['JMA']['models']['deepmind']['mae_hpa']['value'])
48
+
49
+ def test_worker_never_reads_scoring_labels(self):
50
+ source = Path(__file__).with_name('deepmind_daily_benchmark.py').read_text()
51
+ worker = next(n for n in ast.parse(source).body if isinstance(n, ast.FunctionDef) and n.name == 'worker')
52
+ segment = ast.get_source_segment(source, worker)
53
+ self.assertNotIn('truth_local', segment)
54
+ self.assertNotIn('intensity-v12-v11', segment)
55
+ self.assertNotIn('np.load', segment)
56
+ self.assertIn('targets*np.nan', segment)
57
+
58
+
59
+ if __name__ == '__main__':
60
+ unittest.main()
release_tools/test_public_model_cards.py CHANGED
@@ -13,22 +13,22 @@ class PublicModelCardsTest(unittest.TestCase):
13
  def test_full_card_and_remote_metadata_are_preserved(self):
14
  rendered = render_card(self.original, self.source)
15
  self.assertTrue(rendered.startswith(self.header))
16
- for text in ('README revision: 1 October 2026',
17
- 'Daily-issue benchmark: 1,473 forecasts',
18
- 'API contract **1.1**', 'still partial', '25.03 kt'):
19
  self.assertIn(text, rendered)
20
- self.assertIn('N/A — no radius head', rendered)
 
21
  self.assertNotIn('Old model card', rendered)
22
 
23
- def test_exactly_four_pinned_players_and_corrected_primary(self):
24
  rendered = render_card(self.original, self.source)
25
- self.assertEqual(rendered.count('<video '), 4)
26
- for stem in ('fung_wong', 'soudelor', 'mangkhut', 'meranti'):
27
- self.assertEqual(rendered.count('src="' + media_url(stem) + '"'), 1)
28
- self.assertLess(rendered.index('src="' + media_url('mangkhut') + '"'), rendered.index('## What changed'))
29
- self.assertNotIn('/resolve/main/docs/trackformer_', rendered)
30
  self.assertIn(MEDIA_REVISION, rendered)
31
- self.assertIn('Original fixed-patch film', rendered)
32
 
33
  def test_no_github_relative_links_survive(self):
34
  rendered = render_card(self.original, self.source)
@@ -36,8 +36,8 @@ class PublicModelCardsTest(unittest.TestCase):
36
  self.assertTrue(url.startswith(('https:', 'http:', '#', 'mailto:')), url)
37
 
38
  def test_missing_anchor_fails_instead_of_partial_stale_update(self):
39
- with self.assertRaisesRegex(ValueError, 'corrected-film'):
40
- render_card(self.original, self.source.replace('## Corrected pressure forecast', '## Wrong heading'))
41
 
42
  def test_missing_metadata_fails_closed(self):
43
  with self.assertRaisesRegex(ValueError, 'YAML metadata'):
@@ -47,19 +47,22 @@ class PublicModelCardsTest(unittest.TestCase):
47
  with self.assertRaisesRegex(ValueError, 'relative public link'):
48
  render_card(self.original, self.source + '\n[No secrets](../../private-file.txt)\n')
49
 
50
- def test_updated_pressure_graph_and_paper_are_published_together(self):
51
  revision = 'a' * 40
52
  rendered = render_card(self.original, self.source, revision)
53
- for text in ('Pressure graph similarity', '0.7074', '0.7118', '0.7237', '0.6637', '131 days', 'time-warped', 'slight mean improvement', '13.53 to 12.84'):
 
54
  self.assertIn(text, rendered)
55
  for path in CURRENT_FIGURES:
56
  self.assertIn('/resolve/' + revision + '/' + path, rendered)
57
  self.assertIn(path, SYNC_FILES)
58
- for path in ('paper/trackformer.tex', 'paper/trackformer.pdf',
59
  'evaluation/released_daily/released_daily_benchmark.json',
60
  'evaluation/released_daily/released_daily_verification.json'):
61
  self.assertIn(path, SYNC_FILES)
62
  self.assertNotIn('models/trackformer_1_2_field/weights.pt', SYNC_FILES)
 
 
63
  for path in SYNC_FILES:
64
  self.assertTrue((ROOT / path).is_file(), path)
65
 
 
13
  def test_full_card_and_remote_metadata_are_preserved(self):
14
  rendered = render_card(self.original, self.source)
15
  self.assertTrue(rendered.startswith(self.header))
16
+ for text in ('# Introducing Trackformer 1.2', '1,473 days / 270 storms',
17
+ '134 days / 40 storms', '13.53', '12.84',
18
+ 'scores are **pending**', 'no native wind-radius forecast head'):
19
  self.assertIn(text, rendered)
20
+ self.assertNotIn('README revision', rendered)
21
+ self.assertNotIn('Corrected pressure forecast', rendered)
22
  self.assertNotIn('Old model card', rendered)
23
 
24
+ def test_one_pinned_featured_player(self):
25
  rendered = render_card(self.original, self.source)
26
+ self.assertEqual(rendered.count('<video '), 1)
27
+ self.assertEqual(rendered.count('src="' + media_url('mangkhut') + '"'), 1)
28
+ self.assertLess(rendered.index('src="' + media_url('mangkhut') + '"'), rendered.index('## What improves'))
29
+ self.assertIsNone(re.search(r'/resolve/main/docs/trackformer_[^"\s)]+\.mp4', rendered))
 
30
  self.assertIn(MEDIA_REVISION, rendered)
31
+ self.assertIn('docs/showcase_archive.md', rendered)
32
 
33
  def test_no_github_relative_links_survive(self):
34
  rendered = render_card(self.original, self.source)
 
36
  self.assertTrue(url.startswith(('https:', 'http:', '#', 'mailto:')), url)
37
 
38
  def test_missing_anchor_fails_instead_of_partial_stale_update(self):
39
+ with self.assertRaisesRegex(ValueError, 'featured-film'):
40
+ render_card(self.original, self.source.replace('## See the forecast:', '## Wrong heading:'))
41
 
42
  def test_missing_metadata_fails_closed(self):
43
  with self.assertRaisesRegex(ValueError, 'YAML metadata'):
 
47
  with self.assertRaisesRegex(ValueError, 'relative public link'):
48
  render_card(self.original, self.source + '\n[No secrets](../../private-file.txt)\n')
49
 
50
+ def test_new_figures_are_pinned_without_changing_weights_or_paper(self):
51
  revision = 'a' * 40
52
  rendered = render_card(self.original, self.source, revision)
53
+ for text in ('41.0%', '95% interval includes no improvement', 'same frozen starts',
54
+ 'Missing labels', 'certified untouched holdout', '50-member mean'):
55
  self.assertIn(text, rendered)
56
  for path in CURRENT_FIGURES:
57
  self.assertIn('/resolve/' + revision + '/' + path, rendered)
58
  self.assertIn(path, SYNC_FILES)
59
+ for path in ('docs/deepmind_daily_benchmark.md', 'docs/showcase_archive.md',
60
  'evaluation/released_daily/released_daily_benchmark.json',
61
  'evaluation/released_daily/released_daily_verification.json'):
62
  self.assertIn(path, SYNC_FILES)
63
  self.assertNotIn('models/trackformer_1_2_field/weights.pt', SYNC_FILES)
64
+ self.assertNotIn('paper/trackformer.tex', SYNC_FILES)
65
+ self.assertNotIn('paper/trackformer.pdf', SYNC_FILES)
66
  for path in SYNC_FILES:
67
  self.assertTrue((ROOT / path).is_file(), path)
68