Publish completed WeatherNext Cyclones Mini daily comparison and audited three-model chart
Browse files
RELEASE_NOTES_TRACKFORMER_1_2.md
CHANGED
|
@@ -28,14 +28,16 @@ Inference weight SHA-256: `db49f36e85a3766defc4c172746897a1f783705d1ce8e6f9dfb8e
|
|
| 28 |
|
| 29 |
## Matched development results
|
| 30 |
|
| 31 |
-
| Measure | 1.1 | 1.2 mean of 50 | Matched coverage |
|
| 32 |
-
| --- | ---: | ---: | --- |
|
| 33 |
-
| Mean track error | 798.4 km | 471.2 km | 1,473 daily starts / 270 storms |
|
| 34 |
-
| Direction error | 51.58° | 34.96° | Same daily starts |
|
| 35 |
-
| Central-pressure MAE against JMA | 13.53 hPa | 12.84 hPa | 134 common starts / 40 storms |
|
| 36 |
|
| 37 |
Storms receive equal weight after valid leads and daily starts are averaged. The mean pressure reduction is small and its paired whole-storm uncertainty includes no improvement. These repeatedly inspected results are development evidence, not a certified untouched holdout. [Verified common-support metrics](https://github.com/yu314-coder/typhoon-predict/blob/main/evaluation/released_daily/released_daily_benchmark.json).
|
| 38 |
|
|
|
|
|
|
|
| 39 |
The [model announcement](https://github.com/yu314-coder/typhoon-predict) includes the selected Mangkhut pressure-map animation and architecture. Selected examples are not representative skill. Auxiliary wind and pressure-derived radius diagnostics remain unvalidated; there is no native wind-radius forecast head in 1.2.
|
| 40 |
|
| 41 |
This is a research model, not an operational warning service or a safety-critical forecast. Trackformer 1.1 remains a separate release. No historical archive, media, benchmark predictions or training run is modified by this exporter update.
|
|
|
|
| 28 |
|
| 29 |
## Matched development results
|
| 30 |
|
| 31 |
+
| Measure | 1.1 | 1.2 mean of 50 | DeepMind Mini, one member | Matched coverage |
|
| 32 |
+
| --- | ---: | ---: | ---: | --- |
|
| 33 |
+
| Mean track error | 798.4 km | 471.2 km | 498.4 km | 1,473 daily starts / 270 storms |
|
| 34 |
+
| Direction error | 51.58° | 34.96° | 39.76° | Same daily starts |
|
| 35 |
+
| Central-pressure MAE against JMA | 13.53 hPa | 12.84 hPa | 10.20 hPa | 134 common starts / 40 storms |
|
| 36 |
|
| 37 |
Storms receive equal weight after valid leads and daily starts are averaged. The mean pressure reduction is small and its paired whole-storm uncertainty includes no improvement. These repeatedly inspected results are development evidence, not a certified untouched holdout. [Verified common-support metrics](https://github.com/yu314-coder/typhoon-predict/blob/main/evaluation/released_daily/released_daily_benchmark.json).
|
| 38 |
|
| 39 |
+
**The DeepMind comparison is complete and verified:** the official **WeatherNext Cyclones Mini `<2024`** checkpoint (1° resolution, trained through 2023), using **WeatherNext software v0.3.0**, completed all 1,473 daily starts / 270 storms on an RTX 3070 using CUDA. Mini's recent-only mean track error is **288.5 km**, better than 1.2's **477.2 km**. Historical fitting-year overlap, different weather inputs and unequal member counts prevent a general superiority claim. [Completed results and model identity](https://github.com/yu314-coder/typhoon-predict/blob/main/docs/deepmind_daily_benchmark.md) · [CUDA completion receipt](https://github.com/yu314-coder/typhoon-predict/blob/main/evaluation/deepmind_daily/verification.json) · [Independent publication audit](https://github.com/yu314-coder/typhoon-predict/blob/main/evaluation/deepmind_daily/publication_audit.json).
|
| 40 |
+
|
| 41 |
The [model announcement](https://github.com/yu314-coder/typhoon-predict) includes the selected Mangkhut pressure-map animation and architecture. Selected examples are not representative skill. Auxiliary wind and pressure-derived radius diagnostics remain unvalidated; there is no native wind-radius forecast head in 1.2.
|
| 42 |
|
| 43 |
This is a research model, not an operational warning service or a safety-critical forecast. Trackformer 1.1 remains a separate release. No historical archive, media, benchmark predictions or training run is modified by this exporter update.
|
docs/deepmind_daily_benchmark.md
CHANGED
|
@@ -69,3 +69,5 @@ python release_tools/import_deepmind_release_results.py \
|
|
| 69 |
```
|
| 70 |
|
| 71 |
Completed scores are also exposed by the existing [public benchmark API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/api/benchmarks/released). Documentation publication does not replace immutable forecasts, change either model's weights, restart the benchmark, modify active training, or redeploy the website. The original two-model verification receipt remains unchanged; the new publication audit identifies the merged snapshot's exact hash.
|
|
|
|
|
|
|
|
|
| 69 |
```
|
| 70 |
|
| 71 |
Completed scores are also exposed by the existing [public benchmark API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/api/benchmarks/released). Documentation publication does not replace immutable forecasts, change either model's weights, restart the benchmark, modify active training, or redeploy the website. The original two-model verification receipt remains unchanged; the new publication audit identifies the merged snapshot's exact hash.
|
| 72 |
+
|
| 73 |
+
The returned worker's `benchmark.json` is retained byte-for-byte for its audit hashes. Its `published_to_site: false` and `integration_note` describe the handoff before import, not the current benchmark status. The completed CUDA receipt, independent publication audit and shared public snapshot above are the current completion/publication evidence; those historical worker fields do not mean the comparison is unfinished.
|
docs/trackformer_1_2_evaluation.md
CHANGED
|
@@ -20,6 +20,12 @@ The original track run's 1.2-only central-pressure MAE is **12.62 hPa on 40 stor
|
|
| 20 |
|
| 21 |
The other **1,339 starts lack valid frozen 1.1 issue-time intensity inputs**, so this is not 270-storm pressure coverage. The original pressure-only mask is not substituted into the new common-mask comparison. A basin-wide average does not establish cyclone-core accuracy. The frozen cohort contains 40 recent and 230 historical storms from 1980–1999; retrospective analyses, complete-five-day eligibility and prior model selection limit interpretation.
|
| 22 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
## Recent pressure-line example
|
| 24 |
|
| 25 |
### Fung-wong MP4: selected route-and-pressure example
|
|
|
|
| 20 |
|
| 21 |
The other **1,339 starts lack valid frozen 1.1 issue-time intensity inputs**, so this is not 270-storm pressure coverage. The original pressure-only mask is not substituted into the new common-mask comparison. A basin-wide average does not establish cyclone-core accuracy. The frozen cohort contains 40 recent and 230 historical storms from 1980–1999; retrospective analyses, complete-five-day eligibility and prior model selection limit interpretation.
|
| 22 |
|
| 23 |
+
## Completed DeepMind comparison
|
| 24 |
+
|
| 25 |
+
Google's official **WeatherNext Cyclones Mini `<2024`**, using **WeatherNext software v0.3.0**, completed the same **1,473 daily starts / 270 storms** on an RTX 3070 with CUDA. Mini is one seeded member, not the full-sized WeatherNext model; 1.2 remains a 50-member mean. The common-support results are **498.4 km** mean track error, **39.76°** direction error and **10.20 hPa** JMA central-pressure MAE on **134 starts / 40 storms**. Its JMA pressure-curve similarity is **0.8088**.
|
| 26 |
+
|
| 27 |
+
Mini's recent-only mean track error is **288.5 km** versus **477.2 km** for 1.2. Historical fitting-year overlap and different input/member policies limit the combined-cohort comparison; no general superiority claim is justified. The [completed protocol and period breakdowns](deepmind_daily_benchmark.md), [CUDA receipt](../evaluation/deepmind_daily/verification.json) and [independent three-model publication audit](../evaluation/deepmind_daily/publication_audit.json) establish completion. The older two-model receipt remains the original Trackformer audit, not the hash receipt for the subsequently merged snapshot. No missing or unsupported score is treated as zero.
|
| 28 |
+
|
| 29 |
## Recent pressure-line example
|
| 30 |
|
| 31 |
### Fung-wong MP4: selected route-and-pressure example
|
evaluation/README.md
CHANGED
|
@@ -1,12 +1,18 @@
|
|
| 1 |
# Evaluation assets
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
- [
|
| 6 |
-
- [
|
| 7 |
-
- [
|
| 8 |
-
- [
|
| 9 |
|
| 10 |
-
|
| 11 |
|
| 12 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Evaluation assets
|
| 2 |
|
| 3 |
+
The current **completed three-model comparison** scores Trackformer 1.1, Trackformer 1.2 and Google's **WeatherNext Cyclones Mini `<2024` (software v0.3.0)** on **1,473 daily starts / 270 Western Pacific storms**, through +120 h. Mini's RTX 3070 CUDA run is complete and verified; Trackformer 1.2 is a mean of 50 members and Mini is one member.
|
| 4 |
|
| 5 |
+
- [Announcement chart](released_daily/model_1_2_benchmark.png) and [shared three-model scores](released_daily/released_daily_benchmark.json).
|
| 6 |
+
- [Completed DeepMind results](deepmind_daily/benchmark.json), [CUDA completion receipt](deepmind_daily/verification.json) and [independent publication audit](deepmind_daily/publication_audit.json).
|
| 7 |
+
- [Daily-case protocol](../docs/daily_storm_benchmark.md), [DeepMind model identity and period breakdowns](../docs/deepmind_daily_benchmark.md) and [Trackformer evaluation definitions](../docs/trackformer_1_2_evaluation.md).
|
| 8 |
+
- [Matched pressure, wind and radius evaluation](../docs/intensity_benchmark.md): pressure uses 134 shared daily starts / 40 storms, not all 270 storms. Unsupported scores stay unavailable, never zero.
|
| 9 |
|
| 10 |
+
Scores average valid leads within each day, days within each storm, then storms equally. The historical 1980–1999 group overlaps Mini's fitting years; recent 2024+ results are reported separately. This is development evidence, not a certified untouched holdout or an equal-compute architecture ablation.
|
| 11 |
|
| 12 |
+
For geographically aligned routes, compare forecasts and observations at the **same valid times on the same map**. Centred shape similarity alone cannot establish alignment. Never translate, rotate or rescale a forecast to make it match the observed route.
|
| 13 |
+
|
| 14 |
+
Reproduce the current three-model announcement chart from saved results with `python release_tools/plot_model_announcement.py`. This renders audited scores without new inference or weather retrieval; NumPy and Matplotlib are required. The paper is in [`../paper/`](../paper/), and saved pressure arrays and preview provenance are under [`release_data/`](release_data/).
|
| 15 |
+
|
| 16 |
+
## Archived diagnostics and selected examples
|
| 17 |
+
|
| 18 |
+
Older figures, the [TIP diagnostic](trackformer_1_2_vs_1_1_tip_metrics.json) and [showcase selection](trackformer_1_2_showcase_selection.json) are preserved for provenance, not presented as the current benchmark. See the [showcase archive](../docs/showcase_archive.md) for selected pressure-map films and their limitations.
|
models/trackformer_1_2_field/README.md
CHANGED
|
@@ -78,6 +78,8 @@ No model weights, learned architecture or existing benchmark results changed.
|
|
| 78 |
|
| 79 |
The published 50-member mean is a separate evaluation policy: 50 deterministic seeds perturb only normalized historical basin/regional fields with smooth zero-centred noise, then average routes, central pressures and common-grid fields. `predict.py` computes one clean forecast; do not label it a 50-member mean.
|
| 80 |
|
| 81 |
-
**Matched development evaluation:** the released 1.1 and 1.2 route comparison uses the same frozen **1,473 daily starts / 270 equal-weight storms**, +6 to +120 h. The completed native-pressure comparison has **134 common starts / 40 storms**; the other 1,339 starts lack valid frozen 1.1 intensity inputs and are not zero-scored. The Site defaults to central-pressure MAE: **13.53 / 12.84 hPa** against JMA and **12.84 / 12.55 hPa** against USA (1.1 / 1.2), slight mean improvements with paired whole-storm uncertainty including zero. The optional curve-similarity diagnostic is **(1 + centred cosine) / 2 at exact common valid times**, without time shifting or warping. JMA similarity is **0.7074 / 0.7118**; USA is **0.7237 / 0.6637** on 131 eligible non-flat curves. Removing level/amplitude makes this a shape diagnostic, not proof of correct pressure levels. Actual unshifted hPa timelines, agency masks and uncertainty remain separate. See the [shared snapshot](../../evaluation/released_daily/released_daily_benchmark.json), [
|
|
|
|
|
|
|
| 82 |
|
| 83 |
**Pressure-display coverage:** `regional_mslp_hpa` remains the original fixed issue-relative composite. Use the separately exported `core_mslp_hpa` and its per-lead coordinates/mask when the storm moves away from that fixed patch. Do not extend it by inventing a vortex or moving it onto a route. If the moving core leaves supported basin geography, the corresponding masked cells remain unavailable. For an ensemble, register each physical member field onto a common geographic grid **before** averaging; do not average moving-frame array indices. The Mangkhut film uses that separate 50-member policy. Older fixed-patch films retain their original data in [the showcase archive](../../docs/showcase_archive.md). The automatic History archive and its separate core recovery expose geography and coverage through [the public API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/history-api).
|
|
|
|
| 78 |
|
| 79 |
The published 50-member mean is a separate evaluation policy: 50 deterministic seeds perturb only normalized historical basin/regional fields with smooth zero-centred noise, then average routes, central pressures and common-grid fields. `predict.py` computes one clean forecast; do not label it a 50-member mean.
|
| 80 |
|
| 81 |
+
**Matched development evaluation:** the released 1.1 and 1.2 route comparison uses the same frozen **1,473 daily starts / 270 equal-weight storms**, +6 to +120 h. The completed native-pressure comparison has **134 common starts / 40 storms**; the other 1,339 starts lack valid frozen 1.1 intensity inputs and are not zero-scored. The Site defaults to central-pressure MAE: **13.53 / 12.84 hPa** against JMA and **12.84 / 12.55 hPa** against USA (1.1 / 1.2), slight mean improvements with paired whole-storm uncertainty including zero. The optional curve-similarity diagnostic is **(1 + centred cosine) / 2 at exact common valid times**, without time shifting or warping. JMA similarity is **0.7074 / 0.7118**; USA is **0.7237 / 0.6637** on 131 eligible non-flat curves. Removing level/amplitude makes this a shape diagnostic, not proof of correct pressure levels. Actual unshifted hPa timelines, agency masks and uncertainty remain separate. See the [shared snapshot](../../evaluation/released_daily/released_daily_benchmark.json), [original Trackformer audit](../../evaluation/released_daily/released_daily_verification.json), [pressure protocol](../../docs/intensity_benchmark.md) and [public benchmark API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/api/benchmarks/released).
|
| 82 |
+
|
| 83 |
+
**Completed DeepMind comparison:** Google's official **WeatherNext Cyclones Mini `<2024`**, run with **WeatherNext software v0.3.0**, completed all **1,473 daily starts / 270 storms** on an **RTX 3070 using CUDA**. This is the 1° Mini checkpoint trained through 2023, not the full-sized WeatherNext model. Mini uses one seeded member; Trackformer 1.2 uses a 50-member mean. On shared support, Mini's mean track error is **498.4 km**, direction error **39.76°**, and JMA central-pressure MAE **10.20 hPa** on **134 starts / 40 storms**. Mini's recent-only track error is **288.5 km**, better than 1.2's **477.2 km**; the combined historical/recent result is not a general superiority claim. Missing outputs are unavailable, never zero errors. [Completed results and model identity](../../docs/deepmind_daily_benchmark.md) · [CUDA completion receipt](../../evaluation/deepmind_daily/verification.json) · [Independent three-model publication audit](../../evaluation/deepmind_daily/publication_audit.json). No older benchmark's DeepMind scores were substituted.
|
| 84 |
|
| 85 |
**Pressure-display coverage:** `regional_mslp_hpa` remains the original fixed issue-relative composite. Use the separately exported `core_mslp_hpa` and its per-lead coordinates/mask when the storm moves away from that fixed patch. Do not extend it by inventing a vortex or moving it onto a route. If the moving core leaves supported basin geography, the corresponding masked cells remain unavailable. For an ensemble, register each physical member field onto a common geographic grid **before** averaging; do not average moving-frame array indices. The Mangkhut film uses that separate 50-member policy. Older fixed-patch films retain their original data in [the showcase archive](../../docs/showcase_archive.md). The automatic History archive and its separate core recovery expose geography and coverage through [the public API](https://trackformer-weatherlab.rudin-euler-8253.chatgpt.site/history-api).
|
release_tools/sync_public_model_cards.py
CHANGED
|
@@ -50,6 +50,7 @@ SYNC_FILES = (
|
|
| 50 |
'docs/trackformer_1_2_evaluation.md',
|
| 51 |
'docs/deepmind_daily_benchmark.md',
|
| 52 |
'docs/showcase_archive.md',
|
|
|
|
| 53 |
'docs/trackformer_1_2_architecture.svg',
|
| 54 |
'docs/trackformer_1_2_mangkhut.gif',
|
| 55 |
'docs/trackformer_1_2_mangkhut.gif.json',
|
|
|
|
| 50 |
'docs/trackformer_1_2_evaluation.md',
|
| 51 |
'docs/deepmind_daily_benchmark.md',
|
| 52 |
'docs/showcase_archive.md',
|
| 53 |
+
'evaluation/README.md',
|
| 54 |
'docs/trackformer_1_2_architecture.svg',
|
| 55 |
'docs/trackformer_1_2_mangkhut.gif',
|
| 56 |
'docs/trackformer_1_2_mangkhut.gif.json',
|
release_tools/test_public_model_cards.py
CHANGED
|
@@ -1,4 +1,5 @@
|
|
| 1 |
"""Portable regression checks for complete shared README publication."""
|
|
|
|
| 2 |
import re
|
| 3 |
import unittest
|
| 4 |
from sync_public_model_cards import CURRENT_FIGURES, SYNC_FILES, MEDIA_REVISION, ROOT, media_url, render_card
|
|
@@ -79,6 +80,83 @@ class PublicModelCardsTest(unittest.TestCase):
|
|
| 79 |
for path in SYNC_FILES:
|
| 80 |
self.assertTrue((ROOT / path).is_file(), path)
|
| 81 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
if __name__ == '__main__':
|
| 84 |
unittest.main()
|
|
|
|
| 1 |
"""Portable regression checks for complete shared README publication."""
|
| 2 |
+
import json
|
| 3 |
import re
|
| 4 |
import unittest
|
| 5 |
from sync_public_model_cards import CURRENT_FIGURES, SYNC_FILES, MEDIA_REVISION, ROOT, media_url, render_card
|
|
|
|
| 80 |
for path in SYNC_FILES:
|
| 81 |
self.assertTrue((ROOT / path).is_file(), path)
|
| 82 |
|
| 83 |
+
def test_linked_current_docs_cannot_call_completed_run_pending(self):
|
| 84 |
+
for name in ('README.md', 'models/trackformer_1_2_field/README.md',
|
| 85 |
+
'evaluation/README.md', 'docs/deepmind_daily_benchmark.md',
|
| 86 |
+
'docs/daily_storm_benchmark.md', 'docs/trackformer_1_2_evaluation.md',
|
| 87 |
+
'RELEASE_NOTES_TRACKFORMER_1_2.md'):
|
| 88 |
+
with self.subTest(path=name):
|
| 89 |
+
source = (ROOT / name).read_text()
|
| 90 |
+
self.assertIn('Mini', source)
|
| 91 |
+
self.assertIn('1,473', source)
|
| 92 |
+
self.assertIn('270', source)
|
| 93 |
+
for stale in ('results remain pending', 'run is in progress',
|
| 94 |
+
'scores are **pending**', 'deepmind comparison has not been run'):
|
| 95 |
+
self.assertNotIn(stale, source.lower())
|
| 96 |
+
if name != 'README.md':
|
| 97 |
+
self.assertIn(name, SYNC_FILES)
|
| 98 |
+
|
| 99 |
+
def assert_metric_row(self, source, label, values, places):
|
| 100 |
+
rows = [line for line in source.splitlines() if line.startswith('| ')
|
| 101 |
+
and line.split('|')[1].strip().startswith(label)]
|
| 102 |
+
self.assertEqual(len(rows), 1, label)
|
| 103 |
+
cells = [cell.strip() for cell in rows[0].strip().strip('|').split('|')]
|
| 104 |
+
self.assertEqual(len(cells), 5, label)
|
| 105 |
+
for index, key in enumerate(('1.1', '1.2', 'deepmind'), 1):
|
| 106 |
+
match = re.search(r'\d+(?:\.\d+)?', cells[index].replace('**', '').replace(',', ''))
|
| 107 |
+
self.assertIsNotNone(match, (label, key))
|
| 108 |
+
self.assertEqual(match.group(), f'{values[key]:.{places}f}', (label, key))
|
| 109 |
+
|
| 110 |
+
def test_readme_and_release_tables_match_completed_metrics_not_estimates(self):
|
| 111 |
+
exported = json.loads((ROOT / 'evaluation/deepmind_daily/benchmark.json').read_text())
|
| 112 |
+
receipt = json.loads((ROOT / 'evaluation/deepmind_daily/verification.json').read_text())
|
| 113 |
+
metrics = {row['key']: row['values'] for row in exported['metrics']}
|
| 114 |
+
self.assertEqual(exported['status'], 'complete_verified')
|
| 115 |
+
self.assertEqual(receipt['verified_daily_cases'], 1473)
|
| 116 |
+
self.assertEqual(receipt['verified_storms'], 270)
|
| 117 |
+
specs = (
|
| 118 |
+
('Mean track error, +6 to +120 h', 'mean_track_error_km', 1),
|
| 119 |
+
('Six-hour track-direction error', 'direction_error_deg', 2),
|
| 120 |
+
('Central-pressure MAE · JMA', 'pressure_JMA_hpa', 2),
|
| 121 |
+
('Centred route-shape similarity', 'shape_similarity', 4),
|
| 122 |
+
('Pressure-curve similarity · JMA', 'pressure_JMA_curve_similarity', 4),
|
| 123 |
+
)
|
| 124 |
+
for source in (self.source, render_card(self.original, self.source)):
|
| 125 |
+
for label, key, places in specs:
|
| 126 |
+
with self.subTest(label=label):
|
| 127 |
+
self.assert_metric_row(source, label, metrics[key], places)
|
| 128 |
+
release = (ROOT / 'RELEASE_NOTES_TRACKFORMER_1_2.md').read_text()
|
| 129 |
+
for label, key, places in (
|
| 130 |
+
('Mean track error', 'mean_track_error_km', 1),
|
| 131 |
+
('Direction error', 'direction_error_deg', 2),
|
| 132 |
+
('Central-pressure MAE against JMA', 'pressure_JMA_hpa', 2),
|
| 133 |
+
):
|
| 134 |
+
self.assert_metric_row(release, label, metrics[key], places)
|
| 135 |
+
|
| 136 |
+
def test_recent_group_is_not_confused_with_combined_cohort(self):
|
| 137 |
+
exported = json.loads((ROOT / 'evaluation/deepmind_daily/benchmark.json').read_text())
|
| 138 |
+
recent = exported['periods']['recent_2024_onward']['route']
|
| 139 |
+
section = self.source.split('| Recent storms beginning in 2024+', 1)[1]
|
| 140 |
+
for label, key, places in (('Mean track error', 'mean_track_error_km', 1),
|
| 141 |
+
('Six-hour direction error', 'direction_error_deg', 2)):
|
| 142 |
+
values = {model: recent[model][key]['value'] for model in ('1.1', '1.2', 'deepmind')}
|
| 143 |
+
# The recent-only table has four columns, unlike the main coverage table.
|
| 144 |
+
rows = [line for line in section.splitlines() if line.startswith('| ')
|
| 145 |
+
and line.split('|')[1].strip().startswith(label)]
|
| 146 |
+
self.assertEqual(len(rows), 1)
|
| 147 |
+
row = rows[0].rstrip('|').rstrip() + ' | recent-only |'
|
| 148 |
+
self.assert_metric_row(row, label, values, places)
|
| 149 |
+
|
| 150 |
+
def test_evaluation_index_and_worker_handoff_metadata_are_explained(self):
|
| 151 |
+
index = (ROOT / 'evaluation/README.md').read_text()
|
| 152 |
+
self.assertIn('completed three-model comparison', index)
|
| 153 |
+
self.assertNotIn('270 issue times from 90 storms', index)
|
| 154 |
+
self.assertNotIn('results are not implied', index)
|
| 155 |
+
self.assertIn('evaluation/README.md', SYNC_FILES)
|
| 156 |
+
notes = (ROOT / 'docs/deepmind_daily_benchmark.md').read_text()
|
| 157 |
+
self.assertIn('retained byte-for-byte', notes)
|
| 158 |
+
self.assertIn('handoff before import, not the current benchmark status', notes)
|
| 159 |
+
|
| 160 |
|
| 161 |
if __name__ == '__main__':
|
| 162 |
unittest.main()
|