Are Multidimensional Models Worth Their Computational Cost in Demand Forecasting?
Cheng-Jui Fan · Nikolay Aristov · Elenna R. Dugundji
IEEE High Performance Extreme Computing Conference (HPEC), 2026 — Paper 372
Trained checkpoints and archived experiment results accompanying the paper. The benchmark evaluates whether explicitly modeling the commodity, state, and trade-flow axes improves demand-forecasting accuracy enough to justify the additional computation.
Code and paper overview · Dataset · Archived results · Reproduction guide · Citation metadata
Paper overview
We compare statistical baselines, XGBoost, recurrent networks (GRU and LSTM), Transformers, and state-space models (S4, S4ND, Mamba-2, Mamba-3, and Mamba-ND) on monthly U.S. Census merchandise-trade value and shipping weight. Experiments cover a fixed aggregate split, six rolling-origin aggregate folds, and forecasts for individual commodity–state–flow series.
The multidimensional comparisons include Axial Self-Attention (ASA), Convex Factorized Attention (CoFA), and the factorized-attention operator adapted from CaFA. Evaluation considers MSE, MAE, RMSE, sMAPE, MASE, model size, inference FLOPs, and statistical comparisons.
The paper reports that a 12-month moving average leads the fixed aggregate and full multidimensional benchmarks, while persistence leads the rolling-origin aggregate test. S4ND-3D achieves the best neural result on the multidimensional panel. ASA and CoFA help in some settings, but the benefit depends on the backbone and dimensionality; multidimensional structure does not consistently outperform strong statistical baselines. Mamba-3 exhibits substantial seed variability on the full multidimensional benchmark.
Dataset and protocol
| Property | Specification |
|---|---|
| Period | January 2010–December 2025; 192 months |
| Lattice | 1,343 HS6 commodities × 14 states × 2 flows |
| Active series | 30,087 of 37,604 possible combinations; 80.01% density |
| Numerical channels | Nine: five trade-value channels and four shipping-weight channels |
| Targets | Next-month aggregate value and aggregate weight for each series |
| Inputs | 36-month window; nine channels, 24 target-lag features, and month sine/cosine |
| Train / validation / test | 2010–2021 / 2022–2023 / 2024–2025 |
| Preprocessing | log1p, then per-series MinMax fitted on training months only |
| Main-matrix seeds | 947, 732, 619 |
The dataset repository
contains the processed lattice and its required JSON sidecar, plus raw Census
archives. Use repo_type="dataset" for data and repo_type="model" for this
checkpoint repository; both have the identifier Celsia/HPEC2026.
The exact training input pair is pinned to dataset revision
d3f61546516002a44f79962507422f47c4f80263.
Cohort filtering and commodity/state selection use the 144 training months only.
The earlier 28,292-series build used validation/test months in cohort selection
and is retained for historical inspection under
processed/legacy_precorrection/.
It is not the input pair for these checkpoints. The
release provenance
records exact dataset/model revisions and data checksums.
Checkpoints and experiment layout
The root manifest declares the 477-run main matrix (Tests 1, 1.1, 2, 3, 4, and 6). Learning-rate searches and the CoFA/FA-softmax diagnostics are additional experiments stored alongside that matrix.
| Directory or run-name pattern | Contents |
|---|---|
exp0/, exp0_fa/ |
Archived learning-rate searches |
*_aggregate_s* |
Test 1: fixed aggregate forecasts |
*_aggregate_roll_* |
Test 1.1: rolling-origin aggregate forecasts |
*_onehot_*, *_embeddings_* |
Flat and multidimensional panel runs |
*_asa_*, *_fa_*, *_grid_* |
Axial, factorized-attention, and native multidimensional variants |
*_id_* |
Test 4: identity-aware axial variants |
exp7_exp8_operator_study/ |
Separate CoFA and FA-softmax diagnostics |
results/final/ |
Archived tables, figures, workbooks, and analysis outputs |
Paper terminology maps to code as follows: asa_* is ASA; fa_local_* is
CoFA; fa_* is the CaFA authors' FA operator adapted to this categorical
lattice; and fa_sm_* is FA with its softmax switch. CoFA and FA differ in
multiple architectural choices, so their comparison is not a single-factor
ablation. FA has its own learning-rate selection; do not substitute the ASA
learning rate when reproducing it.
A typical saved run
contains best.pth, run_complete.json, logs/metrics.json, a training log,
and cost.json. The release checkpoint verifier checks manifest membership,
completion provenance, and checkpoint SHA-256 before loading tensors.
Results
- Results workbook: aggregate, rolling-origin, flat, and multidimensional panels, with per-seed metrics and analysis sheets.
- Tables and figures: archived LaTeX tables and PDF/PNG figures.
- Inference costs: the final inference-FLOP table; supersedes the earlier
results/model_cost.csvcounts. - Pairwise statistical results: archived comparisons derived from per-observation error dumps.
- Portable result tables and checksums: public code-release copies with original Hub paths in
SOURCES.json.
These are archived research outputs. File availability alone does not establish that every paper statistic has been independently reproduced. Consult the release verification notes and statistical audit data for the validation scope and qualifications. Inference FLOPs and training costs are distinct quantities.
Download and reproduce
Use the standalone code release, which
provides the model implementations, dependency files, and download/verification
commands. These research checkpoints are loaded with that code, rather than
with a generic Transformers from_pretrained call.
After cloning the release and installing its documented dependencies, run from the repository root:
# Download and verify the exact processed lattice and matching sidecar.
python scripts/release.py data
python scripts/release.py data-check
# Download saved result tables and verify an archived checkpoint.
python scripts/release.py results
python scripts/release.py checkpoint --run gru_embeddings_1d_s947
# Check the saved GRU on a small synthetic input.
python scripts/release.py smoke
The release pins model revision
51e064a2f8e1af1ff867794b5ce8aeeade4692cd
for reproducibility. The synthetic smoke check establishes compatibility, not
forecast accuracy. Full-roster training requires Linux, NVIDIA CUDA, and Triton;
see the training and evaluation guide.
Citation
@inproceedings{fan2026multidimensional,
title = {Are Multidimensional Models Worth Their Computational Cost in Demand Forecasting?},
author = {Fan, Cheng-Jui and Aristov, Nikolay and Dugundji, Elenna R.},
booktitle = {IEEE High Performance Extreme Computing Conference (HPEC)},
year = {2026}
}
License
This model repository is distributed under the MIT license. See the code-release license and the dataset card for the separate source-data licensing and provenance.