CELL-FM / README.md
BoHuangLab's picture
Add the citation: paper link under the title, BibTeX in the app and the README
69a3abb verified
|
Raw History Blame Contribute Delete
9.14 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
license: mit
title: CELL-FM
emoji: 🧬
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 5.9.1
app_file: app.py
pinned: false
short_description: Cell microscopy images from protein sequence

CELL-FM

A virtual microscopy model that bridges microscopy images and protein sequences.

Each tab in the app is one application of the model. More will be added; the sections below document the ones that exist.


Condensate Titration

Predict how an Intrinsically Disordered Peptide (IDP) behaves as its concentration rises, directly from sequence.

CELL-FM condensate titration pipeline

The only input is the sequence, and it must be exactly 66 amino acids β€” the entire CondenSeq library the model was trained on is 66-mers (all 14,578 of them), so other lengths are off-distribution. Sampling is fixed at 512 concentration levels and 100 ODE steps, seed 6, matching the offline runs.

Enter a sequence and this section runs the three stages of the CELL-FM condensate pipeline end to end:

  1. Generate. CELL-FM CS samples one microscopy image per protein-intensity level, over a log-spaced concentration ladder from 20 to 2048 a.u. Every image is conditioned on the same fixed reference nucleus, so curves are comparable across sequences.
  2. Classify. A ViT calls each generated image condensed or diffuse.
  3. Summarise. A centred moving average over those binary calls is the titration curve; integrating it against log-intensity gives the two numbers the demo reports.

What AUC and AAC mean

AUC β€” the area under the titration curve. How much of the scanned concentration range the sequence spends in the condensed state. Higher means it condenses more readily, over a wider range.

AAC β€” the area above the curve. The curve is compared against the one it would have followed without reentrant dissolution: everything past the curve's maximum is held flat at that maximum, and AAC is the area between the two. It is zero when the sequence condenses and stays condensed, and grows the more the sequence re-dissolves at high concentration.

Both are normalised by the width of the log-intensity range, so both lie in [0, 1] and are directly comparable between sequences.

Also reported: c_sat, the first intensity at which the smoothed curve reaches probability 0.8, and the raw per-image calls as a downloadable CSV.

Default sequence

The box is pre-filled with the NUP98 IDP, wild type. Its mutation panel — 92 sequences of 66 aa covering F→S, F→Y, charge and spacer variants — ships in assets/nup98_mutation_seqs.csv; paste one in to compare its AUC against WT, which is the quickest way to see the model respond to sequence grammar.


Virtual OpenCell

Browse CELL-FM's virtual staining of 1,276 OpenCell proteins, laid out like an OpenCell target page. Search for a protein:

  • Left is OpenCell's annotation of the real cell line: localization grades (bar height is the grade, 3 the most prominent), UniProt and Ensembl IDs, sequence length, HEK293 abundance, and a link to the real images on OpenCell. None of it is predicted.
  • Right are CELL-FM's images, in OpenCell's colours: the nucleus blue and the generated protein (the target) gray. OpenCell's channel switch and per-channel intensity range and gamma adjust the view; the strip below holds every sample of the protein, and Download TIFF saves the selected one as the raw 2-channel image.

Below the annotations, the virtual staining map places every protein's samples by UMAP of their image embeddings (opencell/vit.bin), coloured by OpenCell localization in the palette of the paper figures; the selected protein's samples are outlined and labelled. Zoom, pan and hover work as in any Plotly chart, and clicking a legend entry hides that category. The map is opencell/vs_umap_embedding.npz in the weights repo; the backdrop draws 12 evenly spaced samples per protein to keep each update light.

Every image is generated from the protein sequence in one shared cell: the nucleus channel is the same real nucleus (ATG7, crop 1) for every protein, so differences between proteins come from the sequence, and a protein's samples are independent draws for that cell.

The images are the dataset BoHuangLab/CELL-FM: OpenCell/virtual_staining/<GENE>/<NNNN>.tif, 93,510 samples, each (2, 256, 256) uint16 with the nucleus first. A protein's samples are fetched the first time it is viewed (VIRTUAL_OPENCELL_LOCAL_DIR reads a local copy instead); while that repo is private, the Space needs an HF_TOKEN secret that can read it. The protein table, assets/virtual_opencell_proteins.csv, is built by notebooks/tools/build_virtual_opencell_assets.py. This section runs on CPU.


Hardware

A GPU is required in practice. The code itself runs on CPU β€” flash-attn is deliberately left out of requirements.txt, so ESM-C uses its pure-torch rotary fallback (verified bit-identical to the Triton path: same AUC, same AAC, same 256/256 per-image calls) β€” but a CPU run of the defaults would take hours and time out. Use ZeroGPU or a GPU hardware tier.

The fixed settings β€” 512 images at 100 ODE steps β€” take ~85 s on an A40 (measured: 11.5 s per batch of 64, plus ~9 s of one-off model loading). Runtime scales linearly in both, so a section that needs to be faster should lower them in pipeline.py (NUM_IMAGES, NUM_STEPS).

On ZeroGPU the per-call budget is set by ZEROGPU_DURATION (default 300 s), which covers the defaults with room to spare.


Running it without a Space

Creating a Gradio Space under an organization now needs a Team or Enterprise plan. Until HuangLab has one, run_local.sh serves the identical app from a cluster GPU over a temporary public URL:

sbatch run_local.sh      # prints a *.gradio.live URL, valid 72 h

Weights

The Space downloads ~2.1 GB of checkpoints at startup from BoHuangLab/CELL-FM, overridable with CELLFM_MODEL_REPO. That repo keeps one subfolder per model family:

File Model
condenseq/cellfm_seq2img.bin CELL-FM CS generator, ESM-C 600M encoder included
condenseq/vae.bin 160Γ—160 image VAE
condenseq/vit_cls.bin Condensed/diffuse ViT classifier

To run against a local copy instead, set CELLFM_LOCAL_WEIGHTS to a flat directory holding the three .bin files β€” the prefix is stripped for local paths. upload_weights.py pushes them from the cluster.

Provenance

Ported from the CELL-FM repository, one module per stage, with behaviour preserved:

Stage Origin
Generation cell_fm/tasks_local/cell_fm_cs/evaluate_seq2img_dict.py
Classification cell_fm/tasks_local/vit_cls_condenseq_img/evaluate_single_img.py
Moving average cell_fm/tasks_local/vit_cls_condenseq_img/ma_plot.py
AUC / AAC cell_diff/tasks/analysis/{utils,ana_all_mutation_log_scale}.py

Model hyperparameters mirror scripts/cell_fm_cs/evaluate_seq2img.sh and scripts/vit_cls_condenseq_img/pretrain.sh. The metric port in metrics.py was checked against the original on 40 stored curves and agrees to 0.0 in both log and linear mode. Generated images make the same round-trip through uint16 that the offline pipeline's TIFF write imposed, so demo numbers stay comparable to published ones.

Reference values, and how noisy they are

Measured at the fixed settings (512 images, 100 steps, seed 6) on an A40:

AUC AAC c_sat
NUP98 WT 0.4912 0.0034 360
1F→S mutant 0.4794 0.0035 412

These are single draws, and how much that matters depends on whether you pair.

Taken alone, the numbers move a lot. Over seeds 6–9 the WT AUC ranges 0.4660–0.5087 (sd 0.019) and c_sat ranges 289–531 (sd 102). Treat any individual c_sat or AAC as indicative rather than precise.

Compared at a matched seed, small differences hold up. The 1Fβ†’S effect is βˆ’0.0123 Β± 0.0042 across four seeds and negative in every one, even though it is smaller than the marginal spread above β€” driving both sequences from the same seed cancels most of the sampling noise, so the comparison is far tighter than either number is on its own. Compare sequences at the same seed and a single-residue change is resolvable; the demo does this by construction, with the seed fixed at 6.

Measured on the cluster stack (torch 2.4.1); the Space runs torch 2.8.0, which is not expected to change these but has not been compared directly.


Citation

If you use CELL-FM, please cite:

@article{zheng2026virtual,
  title   = {Virtual experiments bridge sequence and microscopy with generative models},
  author  = {Zheng, Dihan and Hong, Kibeom and Huang, Bo},
  journal = {bioRxiv},
  year    = {2026},
  doi     = {10.64898/2026.09.13.751243},
  url     = {https://www.biorxiv.org/content/10.64898/2026.09.13.751243v1}
}