File size: 3,001 Bytes
cf0c6ed
2e3eac1
 
 
 
cf0c6ed
 
 
2e3eac1
 
 
cf0c6ed
 
2e3eac1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c08a743
 
 
2e3eac1
 
 
 
 
 
 
 
 
 
c08a743
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2e3eac1
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
---
title: PhenoSeq
emoji: 🧫
colorFrom: green
colorTo: pink
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
short_description: Cell Painting images to single-cell transcriptomes
python_version: "3.12"
startup_duration_timeout: 30m
---

# PhenoSeq — morphology → transcriptome

Interactive demo of [**Sentinal4D/PhenoSeq**](https://huggingface.co/Sentinal4D/PhenoSeq),
a conditional Gaussian diffusion model that generates **single-cell RNA-seq embeddings
(scGPT, 512-d)** from **Cell Painting morphology features** (ViT-L/14 features of the five
fluorescence channels — DNA, RNA, AGP, Mito, ER — 5 × 1024 = 5120-d per cell).

Given the imaging features of 16 cells from a well, the model samples a population of
synthetic transcriptomes for that well via 50-step DDIM.

## What the demo shows

* the actual Cell Painting cells the model is conditioned on (4×4 false-colour montage),
* a PCA scatter of the **generated** transcriptomes against the **real, held-out**
  scGPT profiles measured in the same well,
* a per-dimension embedding profile (generated vs. real),
* quantitative agreement — mean-centred cosine to the correct well, with the cosine to the
  *other* held-out wells as a specificity baseline — plus a **29-way perturbation retrieval**
  check against treatment centroids computed from the *training* wells only,
* the raw generated `(n_cells, 512)` matrix as a downloadable `.npy`.

You can also upload your own `(N, 5120)` ViT-L Cell Painting feature array (`.npy` / `.npz`)
to condition the model on data of your own.

All 15 wells offered in the dropdown are **held-out validation wells** — they were
reconstructed with the authors' own split (`RandomState(42)`, 85/15 over the 94 paired
sample IDs) and the eval-mode cell selection (`np.linspace(0, n-1, 16)`), so nothing shown
here was seen during training.

## What to expect

scGPT embeddings share a large common mean, so raw cosine between any two of them is ≈0.99
and tells you nothing. Everything here is therefore **mean-centred**.

Sampling all 15 held-out wells (256 cells, 50 DDIM steps, seed 0):

| | |
|---|---|
| median centred cosine to the **correct** well | **0.66** |
| median centred cosine to the other 14 wells | 0.40 |
| correct well ranked #1 / top-3 (of 15) | 5/15 · 9/15 (chance 1/15 · 3/15) |
| 29-way treatment retrieval, median rank | 9 / 29 (chance 15), 3/15 at rank 1 |

Clearly above chance, and far from solved: BAY 11-7082 and (R)-roscovitine are recovered
exactly, while TPA in well `202502PA_12` is missed outright.

## Data attribution

The bundled example assets (cell crops, imaging features, real RNA-seq embeddings,
treatment centroids) are small excerpts derived from
[**altoslabs/scGeneScope**](https://huggingface.co/datasets/altoslabs/scGeneScope),
released by Altos Labs under **CC-BY-NC-4.0**. They are redistributed here for
non-commercial research demonstration, with attribution, as the licence requires.

Model weights: `Sentinal4D/PhenoSeq`, Apache-2.0.