CELL-FM / README.md
BoHuangLab's picture
Add the citation: paper link under the title, BibTeX in the app and the README
69a3abb verified
|
Raw History Blame Contribute Delete
9.14 kB
---
license: mit
title: CELL-FM
emoji: 🧬
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 5.9.1
app_file: app.py
pinned: false
short_description: Cell microscopy images from protein sequence
---
# CELL-FM
A virtual microscopy model that bridges microscopy images and protein sequences.
Each tab in the app is one **application** of the model. More will be added; the
sections below document the ones that exist.
---
## Condensate Titration
Predict how an Intrinsically Disordered Peptide (IDP) behaves as its
concentration rises, directly from sequence.
![CELL-FM condensate titration pipeline](images/cellfm_flow_diagram.png)
The only input is the sequence, and it must be **exactly 66 amino acids** — the
entire CondenSeq library the model was trained on is 66-mers (all 14,578 of them),
so other lengths are off-distribution. Sampling is fixed at 512 concentration
levels and 100 ODE steps, seed 6, matching the offline runs.
Enter a sequence and this section runs the three stages of the CELL-FM condensate
pipeline end to end:
1. **Generate.** CELL-FM CS samples one microscopy image per protein-intensity
level, over a log-spaced concentration ladder from 20 to 2048 a.u. Every image
is conditioned on the same fixed reference nucleus, so curves are comparable
across sequences.
2. **Classify.** A ViT calls each generated image **condensed** or **diffuse**.
3. **Summarise.** A centred moving average over those binary calls is the
**titration curve**; integrating it against log-intensity gives the two numbers
the demo reports.
### What AUC and AAC mean
**AUC** — the area under the titration curve. How much of the scanned concentration
range the sequence spends in the condensed state. Higher means it condenses more
readily, over a wider range.
**AAC** — the area *above* the curve. The curve is compared against the one it would
have followed without **reentrant dissolution**: everything past the curve's maximum
is held flat at that maximum, and AAC is the area between the two. It is zero when
the sequence condenses and stays condensed, and grows the more the sequence
re-dissolves at high concentration.
Both are normalised by the width of the log-intensity range, so both lie in [0, 1]
and are directly comparable between sequences.
Also reported: **c_sat**, the first intensity at which the smoothed curve reaches
probability 0.8, and the raw per-image calls as a downloadable CSV.
### Default sequence
The box is pre-filled with the **NUP98 IDP**, wild type. Its mutation panel — 92
sequences of 66 aa covering F→S, F→Y, charge and spacer variants — ships in
`assets/nup98_mutation_seqs.csv`; paste one in to compare its AUC against WT, which
is the quickest way to see the model respond to sequence grammar.
---
## Virtual OpenCell
Browse CELL-FM's virtual staining of 1,276 OpenCell proteins, laid out like an
[OpenCell](https://opencell.sf.czbiohub.org/) target page. Search for a protein:
- **Left** is OpenCell's annotation of the *real* cell line: localization grades
(bar height is the grade, 3 the most prominent), UniProt and Ensembl IDs, sequence
length, HEK293 abundance, and a link to the real images on OpenCell. None of it is
predicted.
- **Right** are CELL-FM's images, in OpenCell's colours: the nucleus blue and the
generated protein (the *target*) gray. OpenCell's channel switch and per-channel
intensity range and gamma adjust the view; the strip below holds every sample of
the protein, and **Download TIFF** saves the selected one as the raw 2-channel image.
Below the annotations, the **virtual staining map** places every protein's samples by
UMAP of their image embeddings (`opencell/vit.bin`), coloured by OpenCell localization in
the palette of the paper figures; the selected protein's samples are outlined and
labelled. Zoom, pan and hover work as in any Plotly chart, and clicking a legend entry
hides that category. The map is `opencell/vs_umap_embedding.npz` in the weights repo;
the backdrop draws 12 evenly spaced samples per protein to keep each update light.
Every image is generated from the protein sequence in **one shared cell**: the nucleus
channel is the same real nucleus (ATG7, crop 1) for every protein, so differences
between proteins come from the sequence, and a protein's samples are independent
draws for that cell.
The images are the dataset
[`BoHuangLab/CELL-FM`](https://huggingface.co/datasets/BoHuangLab/CELL-FM):
`OpenCell/virtual_staining/<GENE>/<NNNN>.tif`, 93,510 samples, each (2, 256, 256)
uint16 with the nucleus first. A protein's samples are fetched the first time it is
viewed (`VIRTUAL_OPENCELL_LOCAL_DIR` reads a local copy instead); while that repo is
private, the Space needs an `HF_TOKEN` secret that can read it. The protein table,
`assets/virtual_opencell_proteins.csv`, is built by
`notebooks/tools/build_virtual_opencell_assets.py`. This section runs on CPU.
---
## Hardware
**A GPU is required in practice.** The code itself runs on CPU — `flash-attn` is
deliberately left out of `requirements.txt`, so ESM-C uses its pure-torch rotary
fallback (verified bit-identical to the Triton path: same AUC, same AAC, same
256/256 per-image calls) — but a CPU run of the defaults would take hours and
time out. Use ZeroGPU or a GPU hardware tier.
The fixed settings — 512 images at 100 ODE steps — take **~85 s on an A40**
(measured: 11.5 s per batch of 64, plus ~9 s of one-off model loading). Runtime
scales linearly in both, so a section that needs to be faster should lower them in
`pipeline.py` (`NUM_IMAGES`, `NUM_STEPS`).
On ZeroGPU the per-call budget is set by `ZEROGPU_DURATION` (default 300 s), which
covers the defaults with room to spare.
---
## Running it without a Space
Creating a Gradio Space under an organization now needs a Team or Enterprise plan.
Until `HuangLab` has one, `run_local.sh` serves the identical app from a cluster GPU
over a temporary public URL:
```bash
sbatch run_local.sh # prints a *.gradio.live URL, valid 72 h
```
## Weights
The Space downloads ~2.1 GB of checkpoints at startup from
[`BoHuangLab/CELL-FM`](https://huggingface.co/BoHuangLab/CELL-FM), overridable with
`CELLFM_MODEL_REPO`. That repo keeps one subfolder per model family:
| File | Model |
|---|---|
| `condenseq/cellfm_seq2img.bin` | CELL-FM CS generator, ESM-C 600M encoder included |
| `condenseq/vae.bin` | 160×160 image VAE |
| `condenseq/vit_cls.bin` | Condensed/diffuse ViT classifier |
To run against a local copy instead, set `CELLFM_LOCAL_WEIGHTS` to a **flat**
directory holding the three `.bin` files — the prefix is stripped for local paths.
`upload_weights.py` pushes them from the cluster.
## Provenance
Ported from the CELL-FM repository, one module per stage, with behaviour preserved:
| Stage | Origin |
|---|---|
| Generation | `cell_fm/tasks_local/cell_fm_cs/evaluate_seq2img_dict.py` |
| Classification | `cell_fm/tasks_local/vit_cls_condenseq_img/evaluate_single_img.py` |
| Moving average | `cell_fm/tasks_local/vit_cls_condenseq_img/ma_plot.py` |
| AUC / AAC | `cell_diff/tasks/analysis/{utils,ana_all_mutation_log_scale}.py` |
Model hyperparameters mirror `scripts/cell_fm_cs/evaluate_seq2img.sh` and
`scripts/vit_cls_condenseq_img/pretrain.sh`. The metric port in `metrics.py` was
checked against the original on 40 stored curves and agrees to 0.0 in both log and
linear mode. Generated images make the same round-trip through uint16 that the
offline pipeline's TIFF write imposed, so demo numbers stay comparable to published
ones.
### Reference values, and how noisy they are
Measured at the fixed settings (512 images, 100 steps, **seed 6**) on an A40:
| | AUC | AAC | c_sat |
|---|---|---|---|
| NUP98 WT | 0.4912 | 0.0034 | 360 |
| 1F→S mutant | 0.4794 | 0.0035 | 412 |
These are single draws, and how much that matters depends on whether you pair.
**Taken alone, the numbers move a lot.** Over seeds 6–9 the WT AUC ranges
0.4660–0.5087 (sd 0.019) and c_sat ranges 289–531 (sd 102). Treat any individual
c_sat or AAC as indicative rather than precise.
**Compared at a matched seed, small differences hold up.** The 1F→S effect is
−0.0123 ± 0.0042 across four seeds and negative in every one, even though it is
smaller than the marginal spread above — driving both sequences from the same seed
cancels most of the sampling noise, so the comparison is far tighter than either
number is on its own. Compare sequences at the same seed and a single-residue
change is resolvable; the demo does this by construction, with the seed fixed at 6.
Measured on the cluster stack (torch 2.4.1); the Space runs torch 2.8.0, which is
not expected to change these but has not been compared directly.
---
## Citation
If you use CELL-FM, please cite:
```bibtex
@article{zheng2026virtual,
title = {Virtual experiments bridge sequence and microscopy with generative models},
author = {Zheng, Dihan and Hong, Kibeom and Huang, Bo},
journal = {bioRxiv},
year = {2026},
doi = {10.64898/2026.09.13.751243},
url = {https://www.biorxiv.org/content/10.64898/2026.09.13.751243v1}
}
```