CodeDevX commited on
Commit
dadd5fa
Β·
verified Β·
1 Parent(s): 8ab625d

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +140 -50
README.md CHANGED
@@ -1,69 +1,159 @@
1
- ---
2
- license: mit
3
- library_name: pytorch
4
- tags:
5
- - time-series
6
- - forecasting
7
- - lstm
8
- - multi-task
9
- - multi-domain
10
- pipeline_tag: time-series-forecasting
11
- ---
12
 
13
- # Future Prediction Models (Multi-Domain LSTM)
14
 
15
- Trained PyTorch LSTM checkpoints forecasting 7 daily time-series domains (AI/NVIDIA, Programming/npm, Finance/BTC, Sports/ATP Elo, Weather, Economy/S&P500, Energy/WTI).
16
 
17
- Trained with a leak-free protocol: chronological TRAIN/VALIDATION/TEST splits, validation-only early stopping and tuning, untouched test set, and walk-forward backtesting as the headline metric.
 
18
 
19
- ## Checkpoints
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
 
21
- All models predict 7 steps from a 60-step window. Inputs are 7 causal features per timestep: base-relative value, 1-step return, MA7 ratio, MA30 ratio, 7-step volatility, day-of-week sin/cos.
 
22
 
23
- | File | Architecture |
24
- |---|---|
25
- | `unified_model.pt` | Multi-task: shared LSTM (128 hidden, 2 layers) + domain embedding (16) + 7 heads; domains: ai, programming, finance, sports, weather, economy, energy |
26
- | `model_ai.pt` | LSTM 64 hidden, 2 layers |
27
- | `model_programming.pt` | LSTM 64 hidden, 2 layers |
28
- | `model_finance.pt` | LSTM 64 hidden, 2 layers |
29
- | `model_sports.pt` | LSTM 64 hidden, 2 layers |
30
- | `model_weather.pt` | LSTM 64 hidden, 2 layers |
31
- | `model_economy.pt` | LSTM 64 hidden, 2 layers |
32
- | `model_energy.pt` | LSTM 64 hidden, 2 layers |
33
 
34
- ## Honest performance (walk-forward backtest vs naive persistence)
 
 
 
 
 
 
 
 
35
 
36
- Positive = model beats "tomorrow equals today" baseline.
37
 
38
- | Topic | vs naive (separate) | vs naive (unified) | Verdict |
 
 
39
  |---|---|---|---|
40
- | Programming | +71% | +68% | Real edge (weekly seasonality) |
41
- | AI | +2.3% | +1.6% | Small edge |
42
- | Energy | +1.8% | +0.9% | Small edge |
43
- | Economy | +0.3% | +2.0% | Marginal |
44
- | Weather | -1.3% | -0.2% | No edge |
45
- | Finance | -4.3% | -2.3% | Trails naive |
46
- | Sports | -11.9% | -22.8% | Trails naive |
47
 
48
- Directional accuracy is ~50-55% on financial series (barely better than chance). Sports direction is mostly undefined because Elo is flat on no-match days.
49
 
50
- ## Loading
 
 
 
 
 
 
 
51
 
52
- ```python
53
- import torch
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
54
 
55
- model = torch.load("model_ai.pt", weights_only=False)
56
- model.eval()
57
- x = torch.randn(1, 60, 7) # last 60 steps of the 7 causal features
58
- pred = model(x) # (1, 7) predicted returns over next 7 steps
 
 
 
 
 
 
 
 
 
 
 
 
59
  ```
60
 
61
- ## Notes
62
 
63
- - Forecasts are statistical estimates; no model predicts the future reliably.
64
- - Uncertainty bands shipped with predictions come from validation residual std (+/-1.96 sigma), not invented confidence.
65
- - Where the model does not beat the persistence baseline, that is reported rather than hidden.
66
 
67
- ## License
 
 
68
 
69
- MIT
 
 
 
 
 
 
 
 
 
 
1
+ # Future Prediction Models (6 topics + Unified)
 
 
 
 
 
 
 
 
 
 
2
 
3
+ An end-to-end multi-domain AI forecasting system. Real datasets, two model modes, ChatGPT-style predictions.
4
 
5
+ ## Models: separate vs unified
6
 
7
+ - **Separate** (`model_<topic>.pt`) β€” 6 dedicated 2-layer LSTMs, one per topic
8
+ - **Unified** (`unified_model.pt`) β€” ONE combined model for all 6 domains: shared LSTM backbone + per-domain embedding + per-domain heads, trained jointly from merged separate models
9
 
10
+ ```powershell
11
+ python train_unified.py # train the single combined model
12
+ python predict.py --topic finance --mode unified # query it for any topic
13
+ python ask.py "compare all topics" # chat uses unified automatically when available
14
+ ```
15
+
16
+ ## Topics & datasets (all real, fetched automatically)
17
+
18
+ | Topic | Asset | Source | Points |
19
+ |---|---|---|---|
20
+ | **AI** | NVIDIA daily close | Yahoo Finance | 6,938 |
21
+ | **Programming** | Daily `react` npm downloads | npm registry API | 547 |
22
+ | **Finance** | Bitcoin BTC-USD close | Yahoo Finance (Hugging Face fallback) | 4,252 |
23
+ | **Sports** | ATP world #1 Elo rating | Hugging Face tennis (93,028 matches, Elo computed from results) | 10,944 |
24
+ | **Weather** | Daily mean temperature (any city) | Open-Meteo archive | 4,248 |
25
+ | **Economy** | S&P 500 daily close | Yahoo Finance | 6,699 |
26
 
27
+ - Weather city is configurable: set env vars `WEATHER_NAME` / `WEATHER_LAT` / `WEATHER_LON`, or add cities to `weather_cities.py` (Chennai, Mumbai, Delhi, London, Tokyo, ... included).
28
+ - Every topic falls back gracefully: Hugging Face search -> API -> synthetic data if fully offline.
29
 
30
+ ## Configs
 
 
 
 
 
 
 
 
 
31
 
32
+ All configuration files are in the `configs/` folder:
33
+ - `configs/config_ai.json`
34
+ - `configs/config_finance.json`
35
+ - `configs/config_economy.json`
36
+ - `configs/config_programming.json`
37
+ - `configs/config_sports.json`
38
+ - `configs/config_weather.json`
39
+ - `configs/unified_config.json` (original unified training config)
40
+ - `configs/unified_config_merged.json` (merged from 6 separate models)
41
 
42
+ ## Validated accuracy (held-out test set, no data leakage)
43
 
44
+ **Separate per-topic models:**
45
+
46
+ | Topic | MAE | H1 MAPE | vs naive baseline |
47
  |---|---|---|---|
48
+ | AI (NVIDIA) | $2.17 | 2.26% | ~naive |
49
+ | Economy (S&P 500) | $37.37 | 0.71% | beats naive by 0.5% |
50
+ | Finance (Bitcoin) | $1,437.46 | 1.61% | ~naive |
51
+ | Programming (react) | 1.17M downloads | 7.12% | **beats naive by 77.5%** |
52
+ | Sports (ATP #1 Elo) | 3.40 Elo | 0.06% | ~naive |
53
+ | Weather (Chennai temp) | 0.99 K | 0.18% | beats naive by 1.9% |
 
54
 
55
+ **Unified single model (merged from 6 separate models):**
56
 
57
+ | Topic | MAE | H1 MAPE |
58
+ |---|---|---|
59
+ | AI | $3.40 | 2.36% |
60
+ | Economy | $72.21 | 0.75% |
61
+ | Finance | $2,863.45 | 1.74% |
62
+ | Programming | 1.14M downloads | 5.42% |
63
+ | Sports | 2.75 Elo | 0.04% |
64
+ | Weather | 0.94 K | 0.18% |
65
 
66
+ > Markets behave near a random walk, so 100% accuracy is impossible β€” these are honest, validated numbers. No model can guarantee the future.
67
+
68
+ ## Architecture (normal mode shows only the final response)
69
+
70
+ ```
71
+ user question -> intent detection (response.py)
72
+ -> prediction_engine.py (structured metrics, prints nothing)
73
+ -> response_formatter.py (natural summary + Markdown, no internal details)
74
+ -> ask.py (chat_output, stdout)
75
+ ```
76
+
77
+ - `ask.py` β€” chat CLI; only the final assistant response reaches the user
78
+ - `response.py` β€” intent detection + orchestration (topic, location, horizon, style)
79
+ - `prediction_engine.py` β€” runs the LSTM, returns distinct metrics:
80
+ `net_change_pct` (path change), `forecast_range_pct` (path min-max range),
81
+ `current_to_forecast_pct` (vs latest observed value), `direction`
82
+ (rising/falling/sideways, configurable threshold), `primary_forecast`
83
+ (end-of-horizon value)
84
+ - `response_formatter.py` β€” presents metrics naturally; opens with a one-sentence
85
+ ChatGPT-style summary; never invents confidence, probability, or values
86
+ - `debug_logger.py` β€” internal logs go to stderr only when `FORECAST_DEBUG=1` (or `--debug`)
87
+ - `data.py` β€” multi-topic data pipeline (fetchers + caching to `data/<topic>.csv`)
88
+ - `weather_cities.py` β€” city registry for location-specific weather models
89
+ - `model.py` / `train.py` β€” 2-layer LSTM, Huber loss, AdamW, early stopping, temporal split
90
+ - `predict.py` β€” per-topic model loader, rollout forecast, chart + text report
91
+ - `multi.py` β€” trains + evaluates all topics; quiet in normal mode
92
+
93
+ ## Ask it anything (ChatGPT-style)
94
+
95
+ ```powershell
96
+ .\.venv\Scripts\python.exe ask.py # interactive chat
97
+ .\.venv\Scripts\python.exe ask.py "what will bitcoin do next week?"
98
+ .\.venv\Scripts\python.exe ask.py "compare all topics"
99
+ .\.venv\Scripts\python.exe ask.py --detailed "predict programming for 14 days"
100
+ .\.venv\Scripts\python.exe ask.py --debug "predict ai for 3 days" # internal logs on stderr
101
+ ```
102
+
103
+ Example response:
104
+
105
+ ```
106
+ In short, the model expects Bitcoin to stay roughly stable next week, hovering near $63,228.
107
+
108
+ Bitcoin Forecast
109
+
110
+ August 18-24, 2026
111
+
112
+ The model forecasts relatively sideways movement for Bitcoin next week, with an
113
+ estimated level around **$63,228**.
114
+
115
+ Outlook: Sideways
116
+ Predicted level (2026-08-24): ~$63,228
117
+ Expected change: <0.01%
118
+ Forecast range: <0.01%
119
+ From latest observed value ($63,229, 2026-08-17): <0.01%
120
+
121
+ This is a model-generated forecast, and actual market behavior may differ.
122
+ ```
123
 
124
+ ## Weather for any city
125
+
126
+ ```powershell
127
+ $env:WEATHER_NAME='Chennai, India'; $env:WEATHER_LAT='13.0827'; $env:WEATHER_LON='80.2707'
128
+ .\.venv\Scripts\python.exe data.py --topic weather --refresh
129
+ .\.venv\Scripts\python.exe train.py --topic weather --epochs 150
130
+ .\.venv\Scripts\python.exe ask.py "predict weather in chennai tomorrow"
131
+ ```
132
+
133
+ ## Train / refresh
134
+
135
+ ```powershell
136
+ .\run_all.ps1 # everything (default topics)
137
+ .\.venv\Scripts\python.exe multi.py # all topics
138
+ .\.venv\Scripts\python.exe train.py --topic sports
139
+ .\.venv\Scripts\python.exe predict.py --topic ai --days 14
140
  ```
141
 
142
+ Outputs: `model_<topic>.pt`, `configs/config_<topic>.json`, `prediction_<topic>.png`, `prediction_<topic>.txt`, `results.json`.
143
 
144
+ ## Add a new domain
 
 
145
 
146
+ 1. Add an entry to `TOPICS` in `data.py` (label, unit, asset name)
147
+ 2. Add a fetcher returning `{source, series}` (date + value columns) to `FETCHERS`
148
+ 3. Run `python train.py --topic <name>` β€” everything else is automatic
149
 
150
+ ## Hugging Face Repository
151
+
152
+ All models and configs are available at:
153
+ **https://huggingface.co/CodeDevX/future-prediction-multi-domain-lstm**
154
+
155
+ Download with:
156
+ ```python
157
+ from huggingface_hub import hf_hub_download
158
+ model_path = hf_hub_download("CodeDevX/future-prediction-multi-domain-lstm", "model_ai.pt", repo_type="model")
159
+ ```