zhangrenchao commited on
Commit
4e054f1
·
verified ·
1 Parent(s): 186a48a

Add English model card

Browse files
Files changed (1) hide show
  1. README.md +145 -0
README.md ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ tags:
6
+ - OneScience
7
+ - Earth Science
8
+ - Hydrological Forecasting
9
+ - Streamflow
10
+ - LSTM
11
+ frameworks: PyTorch
12
+ ---
13
+
14
+ <p align="center"><strong><span style="font-size: 30px;">Streamflow-LSTM</span></strong></p>
15
+
16
+ # Model Introduction
17
+
18
+ Streamflow-LSTM is an engineering reproduction of the gauge-specific LSTM river-flow forecasting method proposed by Hunt et al. It uses the preceding seven days of six-hourly meteorological and hydrological sequences to forecast 40 six-hourly leads, or ten days, for ten western US gauges.
19
+
20
+ Paper: Using a long short-term memory (LSTM) neural network to boost river streamflow forecasts over the western United States
21
+ https://doi.org/10.5194/hess-26-5449-2022
22
+
23
+ # Model Description
24
+
25
+ The model was proposed by researchers from the University of Reading, the European Centre for Medium-Range Weather Forecasts, Loughborough University, and the UK Centre for Ecology and Hydrology. It was trained with streamflow observations from ten western US gauges together with catchment-averaged meteorological forecasts and hydrological histories. The model is suitable for improving gauge-specific river-flow forecasts over the following ten days and for comparison with persistence and GloFAS baselines. Its key feature is independent temporal learning for each gauge, combined with member selection and ensemble averaging to improve forecast robustness.
26
+
27
+ # Use Cases
28
+
29
+ | Use Case | Description |
30
+ | :---: | :--- |
31
+ | Gauge-specific forecasting | Generate 40 six-hourly streamflow leads for ten specified gauges. |
32
+ | Hydrological ensemble experiments | Select random-initialization members by validation NSE and average their forecasts. |
33
+ | Local engineering validation | Test the complete workflow with structured synthetic data retaining the `10x28x23x40` protocol. |
34
+ | ModelScope/OneCode execution | Validate data generation, training, inference, evaluation, and visualization entry points. |
35
+ | Multi-GPU training | Distribute gauge-member tasks with `torchrun` and consolidate one checkpoint. |
36
+
37
+ # Usage Instructions
38
+
39
+ ## 1.OneCode
40
+
41
+ [Try intelligent, one-click AI4S programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
42
+
43
+ ## 2. Download and Installation
44
+
45
+ ```bash
46
+ hf download OneScience-Group/Streamflow-LSTM --local-dir ./Streamflow-LSTM
47
+ cd Streamflow-LSTM
48
+ ```
49
+
50
+ ### Environment Dependencies
51
+
52
+ **Hardware Requirements**
53
+
54
+ - A GPU or DCU is recommended for the paper-scale ensemble.
55
+ - A CPU can be used for connectivity validation with the default small-sample configuration.
56
+ - DCU users must install DTK in advance. DTK 25.04.2 or later, or the OneScience-recommended version matching the current cluster, is recommended.
57
+
58
+ **DCU Environment**
59
+
60
+ ```bash
61
+ # Activate DTK and Conda first
62
+ conda create -n onescience311 python=3.11 -y
63
+ conda activate onescience311
64
+ pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
65
+ ```
66
+
67
+ **GPU Environment**
68
+
69
+ ```bash
70
+ # Activate Conda first
71
+ conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
72
+ conda activate onescience311
73
+ pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
74
+ ```
75
+
76
+ ### Training Data
77
+
78
+ The paper dataset combines sources including ERA5, IFS, GloFAS, and USGS for ten gauges, arranging the previous seven days as 28 six-hourly steps with 23 variables per step. The repository's deterministic synthetic data retain ten gauges, `[B,28,23]` inputs, and 40 six-hourly leads while reducing the numbers of training, validation, and forecast samples. These data validate engineering only and do not represent official distributions, paper-scale training, or paper performance; observed flow is frozen after issue time in forecast inputs to prevent future-observation leakage.
79
+
80
+ ```bash
81
+ python scripts/fake_data.py
82
+ ```
83
+
84
+ ### Training
85
+
86
+ For single-GPU training, use:
87
+
88
+ ```bash
89
+ python scripts/train.py
90
+ ```
91
+
92
+ For multi-GPU training, use:
93
+
94
+ ```bash
95
+ torchrun --nproc_per_node=8 --nnodes=1 --rdzv_id=1000 --rdzv_backend=c10d --max_restarts=0 --master_addr="localhost" --master_port=29500 scripts/train.py
96
+ ```
97
+
98
+ Training uses MSE, Adam at `0.001`, and `0.1` dropout for every gauge member, with member tasks divided among processes. Pass `--paper` to restore 50 hidden units, 100 members per gauge, and selection of the best five members. All gauges and members are consolidated into one checkpoint, and the artifacts are:
99
+
100
+ ```text
101
+ result/checkpoints/streamflow_lstm.pt
102
+ result/training/metrics.json
103
+ ```
104
+
105
+ ### Trained Weights
106
+
107
+ The paper does not provide directly loadable official weights, and `weight/` contains only a status note. Local training stores every ensemble member for all ten gauges, normalization statistics, and validation NSE values in one `result/checkpoints/streamflow_lstm.pt`; it must not be represented as an official pretrained checkpoint.
108
+
109
+ ### Inference
110
+
111
+ ```bash
112
+ python scripts/inference.py
113
+ ```
114
+
115
+ Inference restores every gauge member from the single checkpoint and validates its format version and ten-gauge order. It selects the best two members by validation NSE in default mode or the best five in paper mode, evaluates each `[C,40,28,23]` input, and averages the ensemble. Outputs are clipped to non-negative streamflow and retain 40 six-hourly leads in `m3 s-1`. Inference results are saved to:
116
+
117
+ ```text
118
+ result/output/predictions.npz
119
+ ```
120
+
121
+ ### Evaluation and Visualization
122
+
123
+ ```bash
124
+ python scripts/result.py
125
+ ```
126
+
127
+ Evaluation computes gauge-level KGE, NSE, and RMSE at two-, five-, and eight-day leads, including the correlation, variability, and bias components of KGE. Forecasts are compared with issue-time persistence and a synthetic GloFAS proxy. The figure presents mean RMSE over all leads and five-day KGE for each gauge; synthetic-data results validate engineering only. Evaluation artifacts are saved to:
128
+
129
+ ```text
130
+ result/evaluation/metrics.json
131
+ result/evaluation/comparison.png
132
+ ```
133
+
134
+ # Official OneScience Information
135
+
136
+ | Platform | OneScience Main Repository | Skills Repository |
137
+ | --- | --- | --- |
138
+ | Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
139
+ | GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
140
+
141
+ # Citation and License
142
+
143
+ This repository is an independent engineering reproduction of the public Streamflow-LSTM specifications, with code licensed under the Apache License 2.0.
144
+
145
+ The original paper is licensed under CC BY 4.0; the paper, model weights, and ERA5, IFS, GloFAS, and USGS data remain subject to the licenses and terms of their respective projects.