--- title: PhenoSeq emoji: ๐Ÿงซ colorFrom: green colorTo: pink sdk: gradio sdk_version: 6.26.0 app_file: app.py short_description: Cell Painting images to single-cell transcriptomes python_version: "3.12" startup_duration_timeout: 30m --- # PhenoSeq โ€” morphology โ†’ transcriptome Interactive demo of [**Sentinal4D/PhenoSeq**](https://huggingface.co/Sentinal4D/PhenoSeq), a conditional Gaussian diffusion model that generates **single-cell RNA-seq embeddings (scGPT, 512-d)** from **Cell Painting morphology features** (ViT-L/14 features of the five fluorescence channels โ€” DNA, RNA, AGP, Mito, ER โ€” 5 ร— 1024 = 5120-d per cell). Given the imaging features of 16 cells from a well, the model samples a population of synthetic transcriptomes for that well via 50-step DDIM. ## What the demo shows * the actual Cell Painting cells the model is conditioned on (4ร—4 false-colour montage), * a PCA scatter of the **generated** transcriptomes against the **real, held-out** scGPT profiles measured in the same well, * a per-dimension embedding profile (generated vs. real), * quantitative agreement โ€” mean-centred cosine to the correct well, with the cosine to the *other* held-out wells as a specificity baseline โ€” plus a **29-way perturbation retrieval** check against treatment centroids computed from the *training* wells only, * the raw generated `(n_cells, 512)` matrix as a downloadable `.npy`. You can also upload your own `(N, 5120)` ViT-L Cell Painting feature array (`.npy` / `.npz`) to condition the model on data of your own. All 15 wells offered in the dropdown are **held-out validation wells** โ€” they were reconstructed with the authors' own split (`RandomState(42)`, 85/15 over the 94 paired sample IDs) and the eval-mode cell selection (`np.linspace(0, n-1, 16)`), so nothing shown here was seen during training. ## What to expect scGPT embeddings share a large common mean, so raw cosine between any two of them is โ‰ˆ0.99 and tells you nothing. Everything here is therefore **mean-centred**. Sampling all 15 held-out wells (256 cells, 50 DDIM steps, seed 0): | | | |---|---| | median centred cosine to the **correct** well | **0.66** | | median centred cosine to the other 14 wells | 0.40 | | correct well ranked #1 / top-3 (of 15) | 5/15 ยท 9/15 (chance 1/15 ยท 3/15) | | 29-way treatment retrieval, median rank | 9 / 29 (chance 15), 3/15 at rank 1 | Clearly above chance, and far from solved: BAY 11-7082 and (R)-roscovitine are recovered exactly, while TPA in well `202502PA_12` is missed outright. ## Data attribution The bundled example assets (cell crops, imaging features, real RNA-seq embeddings, treatment centroids) are small excerpts derived from [**altoslabs/scGeneScope**](https://huggingface.co/datasets/altoslabs/scGeneScope), released by Altos Labs under **CC-BY-NC-4.0**. They are redistributed here for non-commercial research demonstration, with attribution, as the licence requires. Model weights: `Sentinal4D/PhenoSeq`, Apache-2.0.