FTFM-1
A tabular foundation model that classifies in context: fit reads a
labelled support table, predict scores new rows. No gradient steps, no
per-dataset tuning.
pip install "git+https://github.com/machinelearningnuremberg/ftfm-official.git"
from ftfm import FTFMClassifier
clf = FTFMClassifier().fit(X_train, y_train) # weights fetched on first use
proba = clf.predict_proba(X_test)
FTFMClassifier() downloads these weights on the first fit, caches them
under ~/.cache/huggingface, and predicts with an 8-member view ensemble by
default. ftfm-download (or ftfm.download_checkpoint()) fetches them ahead
of time, e.g. for offline jobs.
Licence β read before use
| These weights | CC BY-NC 4.0 β NON-COMMERCIAL use only |
The ftfm source code |
Apache-2.0 |
Research use is granted. These weights are released for research, and the following need no further permission: academic and scientific research (including inside a company's research division, where the work itself is not directed at commercial advantage); publishing results, papers, theses and public leaderboard entries β including numbers unfavourable to FTFM; benchmarking and independent verification; teaching; and redistributing the weights, or derived weights, under these same terms with attribution.
Commercial use β deploying the weights or anything derived from them in a
product or service β is prohibited without prior written permission from
the rights holder, Prof. Dr. Josif Grabocka, University of Technology
Nuremberg. Commercial licensing is available on request. The same code/weights
split is used by TabFM and EXAONE. Full terms: LICENSE-WEIGHTS.md, shipped
beside these weights and in the source repository.
Derived weights β fine-tuned, retrained, merged, quantized, pruned or distilled β inherit CC BY-NC 4.0.
FTFM model weights, Prof. Dr. Josif Grabocka, University of Technology
Nuremberg. Licensed under CC BY-NC 4.0.
https://huggingface.co/josifgrabocka/ftfm
Model
| Checkpoint | stage C, update 130,000 (weight-EMA parameter set) |
| Parameters | 121,635,840 |
| Architecture | 16 factorized row/column blocks + readout, width 512, 8 heads |
| Precision | FP32 weights |
| Task | Classification; up to 10 classes natively, more through error-correcting output codes over 10-class members |
| Input | Numeric and categorical columns (DataFrames with string columns are accepted); missing values allowed |
Cell representations are built by factorizing one contextual latent per row
against one per column, rather than by cell-to-cell attention. The support side
is computed once at fit and cached; every query row is scored independently
of the others, so predict runs in arbitrary chunks without approximation.
Training
Three stages, each warm-started from the one before:
| Stage | Prior | Rows | Features | Updates |
|---|---|---|---|---|
| A | synthetic graph-SCM tables | 1,024 | 1β100 | 500k |
| B | synthetic + real-table episodes, 1:1 | 256β16,384 | 4β128 | 200k |
| C | synthetic + real-table episodes, 1:1 | 256β32,768 | 4β128 | 130k |
Stages B and C draw half their updates from self-supervised episodes over a
corpus of public tables, filtered against the evaluation benchmarks by dataset
name. A later content-level audit found that this name filter missed a small
number of benchmark tables, so results on those tables should be read with
that in mind. Provenance for this checkpoint is recorded in config.json
(trained_updates, tables_seen, weight_selection, and the full prior
envelope).
Limitations
- Classification only; regression weights are not released.
- The fitted context lives in device memory, so very large support tables are bounded by it. Query rows are not.
Citation
@software{grabocka_ftfm,
author = {Grabocka, Josif},
title = {FTFM: a factorized tabular foundation model},
url = {https://huggingface.co/josifgrabocka/ftfm}
}
- Downloads last month
- 35