Are Multidimensional Models Worth Their Computational Cost in Demand Forecasting?

Cheng-Jui Fan · Nikolay Aristov · Elenna R. Dugundji
IEEE High Performance Extreme Computing Conference (HPEC), 2026 — Paper 372

Trained checkpoints and archived experiment results accompanying the paper. The benchmark evaluates whether explicitly modeling the commodity, state, and trade-flow axes improves demand-forecasting accuracy enough to justify the additional computation.

Code and paper overview · Dataset · Archived results · Reproduction guide · Citation metadata

Paper overview

We compare statistical baselines, XGBoost, recurrent networks (GRU and LSTM), Transformers, and state-space models (S4, S4ND, Mamba-2, Mamba-3, and Mamba-ND) on monthly U.S. Census merchandise-trade value and shipping weight. Experiments cover a fixed aggregate split, six rolling-origin aggregate folds, and forecasts for individual commodity–state–flow series.

The multidimensional comparisons include Axial Self-Attention (ASA), Convex Factorized Attention (CoFA), and the factorized-attention operator adapted from CaFA. Evaluation considers MSE, MAE, RMSE, sMAPE, MASE, model size, inference FLOPs, and statistical comparisons.

The paper reports that a 12-month moving average leads the fixed aggregate and full multidimensional benchmarks, while persistence leads the rolling-origin aggregate test. S4ND-3D achieves the best neural result on the multidimensional panel. ASA and CoFA help in some settings, but the benefit depends on the backbone and dimensionality; multidimensional structure does not consistently outperform strong statistical baselines. Mamba-3 exhibits substantial seed variability on the full multidimensional benchmark.

Dataset and protocol

Property Specification
Period January 2010–December 2025; 192 months
Lattice 1,343 HS6 commodities × 14 states × 2 flows
Active series 30,087 of 37,604 possible combinations; 80.01% density
Numerical channels Nine: five trade-value channels and four shipping-weight channels
Targets Next-month aggregate value and aggregate weight for each series
Inputs 36-month window; nine channels, 24 target-lag features, and month sine/cosine
Train / validation / test 2010–2021 / 2022–2023 / 2024–2025
Preprocessing log1p, then per-series MinMax fitted on training months only
Main-matrix seeds 947, 732, 619

The dataset repository contains the processed lattice and its required JSON sidecar, plus raw Census archives. Use repo_type="dataset" for data and repo_type="model" for this checkpoint repository; both have the identifier Celsia/HPEC2026.

The exact training input pair is pinned to dataset revision d3f61546516002a44f79962507422f47c4f80263. Cohort filtering and commodity/state selection use the 144 training months only. The earlier 28,292-series build used validation/test months in cohort selection and is retained for historical inspection under processed/legacy_precorrection/. It is not the input pair for these checkpoints. The release provenance records exact dataset/model revisions and data checksums.

Checkpoints and experiment layout

The root manifest declares the 477-run main matrix (Tests 1, 1.1, 2, 3, 4, and 6). Learning-rate searches and the CoFA/FA-softmax diagnostics are additional experiments stored alongside that matrix.

Directory or run-name pattern Contents
exp0/, exp0_fa/ Archived learning-rate searches
*_aggregate_s* Test 1: fixed aggregate forecasts
*_aggregate_roll_* Test 1.1: rolling-origin aggregate forecasts
*_onehot_*, *_embeddings_* Flat and multidimensional panel runs
*_asa_*, *_fa_*, *_grid_* Axial, factorized-attention, and native multidimensional variants
*_id_* Test 4: identity-aware axial variants
exp7_exp8_operator_study/ Separate CoFA and FA-softmax diagnostics
results/final/ Archived tables, figures, workbooks, and analysis outputs

Paper terminology maps to code as follows: asa_* is ASA; fa_local_* is CoFA; fa_* is the CaFA authors' FA operator adapted to this categorical lattice; and fa_sm_* is FA with its softmax switch. CoFA and FA differ in multiple architectural choices, so their comparison is not a single-factor ablation. FA has its own learning-rate selection; do not substitute the ASA learning rate when reproducing it.

A typical saved run contains best.pth, run_complete.json, logs/metrics.json, a training log, and cost.json. The release checkpoint verifier checks manifest membership, completion provenance, and checkpoint SHA-256 before loading tensors.

Results

These are archived research outputs. File availability alone does not establish that every paper statistic has been independently reproduced. Consult the release verification notes and statistical audit data for the validation scope and qualifications. Inference FLOPs and training costs are distinct quantities.

Download and reproduce

Use the standalone code release, which provides the model implementations, dependency files, and download/verification commands. These research checkpoints are loaded with that code, rather than with a generic Transformers from_pretrained call.

After cloning the release and installing its documented dependencies, run from the repository root:

# Download and verify the exact processed lattice and matching sidecar.
python scripts/release.py data
python scripts/release.py data-check

# Download saved result tables and verify an archived checkpoint.
python scripts/release.py results
python scripts/release.py checkpoint --run gru_embeddings_1d_s947

# Check the saved GRU on a small synthetic input.
python scripts/release.py smoke

The release pins model revision 51e064a2f8e1af1ff867794b5ce8aeeade4692cd for reproducibility. The synthetic smoke check establishes compatibility, not forecast accuracy. Full-roster training requires Linux, NVIDIA CUDA, and Triton; see the training and evaluation guide.

Citation

@inproceedings{fan2026multidimensional,
  title     = {Are Multidimensional Models Worth Their Computational Cost in Demand Forecasting?},
  author    = {Fan, Cheng-Jui and Aristov, Nikolay and Dugundji, Elenna R.},
  booktitle = {IEEE High Performance Extreme Computing Conference (HPEC)},
  year      = {2026}
}

License

This model repository is distributed under the MIT license. See the code-release license and the dataset card for the separate source-data licensing and provenance.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Celsia/HPEC2026