Papers
arxiv:2606.09257

BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

Published on Jun 8
Authors:
,
,
,

Abstract

High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by n ll m, where n = number of samples, and m = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in R^m ill-conditioned since n ll m. We propose BSTabDiff, a block-subunit generative framework that partitions the m observed features into M latent blocks (M ll m) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space R^M while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.09257
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.09257 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.09257 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.09257 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.