PBISC-Diffusion
The first diffusion reference is paper-2k-e3-ema, the manuscript's fixed 2,000-output-gene epoch-3 EMA checkpoint (143,293 optimizer steps). It consumes the ordered 20,097-gene input axis. This is a manuscript reference, not a claim of superiority over later checkpoints or larger gene panels.
Code and installation: https://github.com/junidude/PBISC Training data: https://huggingface.co/datasets/junidude14/PBISC-data
variants/paper-2k-e3-ema/ contains safe inference weights, full input/output
gene axes, train-fitted conditioner/scaler, masks, sampling configuration and
provenance. The complete bound dataset manifest is included for identity checks;
the corpus itself is unnecessary for inference. Generated count-scale values
are continuous predictions and must not be presented as measured raw counts.
Install the PBISC package and load this repository with an immutable revision:
from pbisc import PBISCPipeline
pipe = PBISCPipeline.from_pretrained("junidude14/PBISC", revision="MODEL_COMMIT_SHA")
result = pipe.sample("conditions.npz", steps=50, seed=0)
pipe.save_sample("generated.npz", result)
Use the model SHA from the GitHub release lock. Inputs must supply all canonical gene IDs in exact order, bulk values, bulk measurement fractions, output assay availability and recipe K. Sampling uses CPU-drawn initial noise and DDIM eta=0. A repeated seed does not guarantee bitwise identity across hardware/software.
training/paper-2k-e3/resume.pt preserves the trusted original checkpoint with
raw/EMA/optimizer/RNG state. Follow the training documentation and verify its
SHA before loading. The lightweight inference bundle cannot resume training.
Prior CVAE/refiner checkpoints remain available at HF revision
b1b84d60ea084bcd12fb14747d3dc485fb61ec53. They are distinct model families.
No new blanket license is asserted for model weights or source datasets.
Release verification covers artifact identity and selected CPU parity tests,
not clinical utility or complete GPU training reproducibility.
0.1%-floor panel campaign (2026-09-23/24)
training/floor01_20260923/{2k,6k,p15k}/ holds the part-16 checkpoints of three panels selected at
a 0.1% detection floor instead of the published 1% one, which was relative and therefore excluded
every gene specific to a cell type rarer than 1% of PBMC (pDC, cDC1, HSPC, platelet, erythroid).
The panels differ only in how many genes the model emits: same 1x architecture, data, seed,
schedule, sampler and drawn read-out.
On the held-out childhood SLE cohort, scored on the 1,006 genes every panel contains, widening the
panel does not help: 2k 0.348 logFC r / 53 top-50 DEG / 0.131 composition L1, 6k 0.399 / 81 /
0.181, p15k 0.344 / 62 / 0.297. MANIFEST.json carries the per-file checksums and the full result;
CHECKPOINTS.sha256 the digests alone. Code and the campaign record are in
reproduce/floor01-panel-campaign.md.
These are training checkpoints, not inference bundles: they resume training and require the
repository's training code, not PBISCPipeline.