VertiMosaic VFL Reference Models
Portable, deterministic CPU reference checkpoints for VertiMosaic vertical federated learning.
๐ Explore results in the Space ยท ๐ Open the benchmark dataset ยท ๐ป Source code
This repository contains two research checkpoints generated from the VertiMosaic source tree:
VFLLogisticRegressionโ first-principles vertical logistic regression.VFLHistGBDTโ vertical histogram gradient boosting with party-local routing state.
They are deliberately small, inspectable reference artifacts rather than opaque production binaries.
Dataset relationship: the linked Hugging Face benchmark dataset contains generated evaluation summaries and robustness-study tables. It is not the row-level training corpus used to fit these checkpoints.
At a glance
| Item | Reference bundle |
|---|---|
| Checkpoints | 2 |
| Benchmark population | 800 aligned entities |
| Seed | 42 |
| Entity split | 560 train / 120 validation / 120 test |
| Logistic release setting | max_iter=150 |
| GBDT release setting | n_estimators=8 |
| Runtime target | CPU-first |
| Round-trip validation | prediction-level for both formats |
| Source commit | ede09d38933a91aec31296b70865a5528c4ec958 |
Held-out reference evaluation
| Checkpoint | ROC-AUC | PR-AUC | F1 | Brier |
|---|---|---|---|---|
| VFL Logistic | 0.7872 | 0.7860 | 0.7361 | 0.1879 |
| VFL HistGBDT | 0.7041 | 0.6749 | 0.6622 | 0.2283 |
These are single deterministic synthetic benchmark measurements, not a leaderboard and not evidence that VFL generally outperforms centralized learning.
What is in the repository?
| Path | Purpose |
|---|---|
logistic/model.json |
coordinator-visible logistic configuration and per-party weights |
gbdt/model.json |
ensemble topology, leaf values and opaque party/feature/bin references |
gbdt/parties/*.json |
synthetic party-local routing thresholds kept as separate party shards |
evaluation.json |
held-out metrics, validation-selected thresholds and prediction digests |
metadata.json |
source commit, split sizes and exact generation configuration |
The JSON formats are intentionally inspectable so the published checkpoints can be independently validated.
Inspect locally
import json
from pathlib import Path
evaluation = json.loads(Path("evaluation.json").read_text())
metadata = json.loads(Path("metadata.json").read_text())
print(evaluation)
print(metadata["source_commit"])
To regenerate the complete bundle from source:
git clone https://github.com/sauravsingla/VertiMosaic.git
cd VertiMosaic
python -m pip install -e .
SOURCE_SHA=$(git rev-parse HEAD) python huggingface/build_model.py --output hf-model
The builder performs prediction-level round-trip checks for both checkpoint formats before the package is published.
Privacy and deployment boundary
The reference VertiMosaic protocols keep raw party feature matrices local, but that does not imply end-to-end cryptographic privacy. Gradient, Hessian, residual, logit, routing, timing and transport information can remain outside stronger privacy guarantees depending on the path used.
For the public synthetic GBDT checkpoint, train-derived routing thresholds are published as separate party shards for reproducibility. Real organizations should not centralize or publicly release analogous private party-local routing state merely because this synthetic research artifact does so.
These checkpoints are not production models, are not trained on real organizational data, and should not be used to make real-world decisions.
Follow the evidence
- Interactive Space: https://huggingface.co/spaces/sauravsingla08/VertiMosaic
- Benchmark dataset: https://huggingface.co/datasets/sauravsingla08/VertiMosaic-VFL-Benchmark
- Source repository: https://github.com/sauravsingla/VertiMosaic
- PyPI: https://pypi.org/project/vertimosaic/
Generated from source commit ede09d38933a91aec31296b70865a5528c4ec958.