euler314's picture
Publish completed WeatherNext Cyclones Mini daily comparison and audited three-model chart
5066a90 verified
|
Raw History Blame Contribute Delete
2.44 kB

Evaluation assets

The current completed three-model comparison scores Trackformer 1.1, Trackformer 1.2 and Google's WeatherNext Cyclones Mini <2024 (software v0.3.0) on 1,473 daily starts / 270 Western Pacific storms, through +120 h. Mini's RTX 3070 CUDA run is complete and verified; Trackformer 1.2 is a mean of 50 members and Mini is one member.

Scores average valid leads within each day, days within each storm, then storms equally. The historical 1980–1999 group overlaps Mini's fitting years; recent 2024+ results are reported separately. This is development evidence, not a certified untouched holdout or an equal-compute architecture ablation.

For geographically aligned routes, compare forecasts and observations at the same valid times on the same map. Centred shape similarity alone cannot establish alignment. Never translate, rotate or rescale a forecast to make it match the observed route.

Reproduce the current three-model announcement chart from saved results with python release_tools/plot_model_announcement.py. This renders audited scores without new inference or weather retrieval; NumPy and Matplotlib are required. The paper is in ../paper/, and saved pressure arrays and preview provenance are under release_data/.

Archived diagnostics and selected examples

Older figures, the TIP diagnostic and showcase selection are preserved for provenance, not presented as the current benchmark. See the showcase archive for selected pressure-map films and their limitations.