--- license: mit title: CELL-FM emoji: 🧬 colorFrom: blue colorTo: indigo sdk: gradio sdk_version: 5.9.1 app_file: app.py pinned: false short_description: Cell microscopy images from protein sequence --- # CELL-FM A virtual microscopy model that bridges microscopy images and protein sequences. Each tab in the app is one **application** of the model. More will be added; the sections below document the ones that exist. --- ## Condensate Titration Predict how an Intrinsically Disordered Peptide (IDP) behaves as its concentration rises, directly from sequence. ![CELL-FM condensate titration pipeline](images/cellfm_flow_diagram.png) The only input is the sequence, and it must be **exactly 66 amino acids** — the entire CondenSeq library the model was trained on is 66-mers (all 14,578 of them), so other lengths are off-distribution. Sampling is fixed at 512 concentration levels and 100 ODE steps, seed 6, matching the offline runs. Enter a sequence and this section runs the three stages of the CELL-FM condensate pipeline end to end: 1. **Generate.** CELL-FM CS samples one microscopy image per protein-intensity level, over a log-spaced concentration ladder from 20 to 2048 a.u. Every image is conditioned on the same fixed reference nucleus, so curves are comparable across sequences. 2. **Classify.** A ViT calls each generated image **condensed** or **diffuse**. 3. **Summarise.** A centred moving average over those binary calls is the **titration curve**; integrating it against log-intensity gives the two numbers the demo reports. ### What AUC and AAC mean **AUC** — the area under the titration curve. How much of the scanned concentration range the sequence spends in the condensed state. Higher means it condenses more readily, over a wider range. **AAC** — the area *above* the curve. The curve is compared against the one it would have followed without **reentrant dissolution**: everything past the curve's maximum is held flat at that maximum, and AAC is the area between the two. It is zero when the sequence condenses and stays condensed, and grows the more the sequence re-dissolves at high concentration. Both are normalised by the width of the log-intensity range, so both lie in [0, 1] and are directly comparable between sequences. Also reported: **c_sat**, the first intensity at which the smoothed curve reaches probability 0.8, and the raw per-image calls as a downloadable CSV. ### Default sequence The box is pre-filled with the **NUP98 IDP**, wild type. Its mutation panel — 92 sequences of 66 aa covering F→S, F→Y, charge and spacer variants — ships in `assets/nup98_mutation_seqs.csv`; paste one in to compare its AUC against WT, which is the quickest way to see the model respond to sequence grammar. --- ## Virtual OpenCell Browse CELL-FM's virtual staining of 1,276 OpenCell proteins, laid out like an [OpenCell](https://opencell.sf.czbiohub.org/) target page. Search for a protein: - **Left** is OpenCell's annotation of the *real* cell line: localization grades (bar height is the grade, 3 the most prominent), UniProt and Ensembl IDs, sequence length, HEK293 abundance, and a link to the real images on OpenCell. None of it is predicted. - **Right** are CELL-FM's images, in OpenCell's colours: the nucleus blue and the generated protein (the *target*) gray. OpenCell's channel switch and per-channel intensity range and gamma adjust the view; the strip below holds every sample of the protein, and **Download TIFF** saves the selected one as the raw 2-channel image. Below the annotations, the **virtual staining map** places every protein's samples by UMAP of their image embeddings (`opencell/vit.bin`), coloured by OpenCell localization in the palette of the paper figures; the selected protein's samples are outlined and labelled. Zoom, pan and hover work as in any Plotly chart, and clicking a legend entry hides that category. The map is `opencell/vs_umap_embedding.npz` in the weights repo; the backdrop draws 12 evenly spaced samples per protein to keep each update light. Every image is generated from the protein sequence in **one shared cell**: the nucleus channel is the same real nucleus (ATG7, crop 1) for every protein, so differences between proteins come from the sequence, and a protein's samples are independent draws for that cell. The images are the dataset [`BoHuangLab/CELL-FM`](https://huggingface.co/datasets/BoHuangLab/CELL-FM): `OpenCell/virtual_staining//.tif`, 93,510 samples, each (2, 256, 256) uint16 with the nucleus first. A protein's samples are fetched the first time it is viewed (`VIRTUAL_OPENCELL_LOCAL_DIR` reads a local copy instead); while that repo is private, the Space needs an `HF_TOKEN` secret that can read it. The protein table, `assets/virtual_opencell_proteins.csv`, is built by `notebooks/tools/build_virtual_opencell_assets.py`. This section runs on CPU. --- ## Hardware **A GPU is required in practice.** The code itself runs on CPU — `flash-attn` is deliberately left out of `requirements.txt`, so ESM-C uses its pure-torch rotary fallback (verified bit-identical to the Triton path: same AUC, same AAC, same 256/256 per-image calls) — but a CPU run of the defaults would take hours and time out. Use ZeroGPU or a GPU hardware tier. The fixed settings — 512 images at 100 ODE steps — take **~85 s on an A40** (measured: 11.5 s per batch of 64, plus ~9 s of one-off model loading). Runtime scales linearly in both, so a section that needs to be faster should lower them in `pipeline.py` (`NUM_IMAGES`, `NUM_STEPS`). On ZeroGPU the per-call budget is set by `ZEROGPU_DURATION` (default 300 s), which covers the defaults with room to spare. --- ## Running it without a Space Creating a Gradio Space under an organization now needs a Team or Enterprise plan. Until `HuangLab` has one, `run_local.sh` serves the identical app from a cluster GPU over a temporary public URL: ```bash sbatch run_local.sh # prints a *.gradio.live URL, valid 72 h ``` ## Weights The Space downloads ~2.1 GB of checkpoints at startup from [`BoHuangLab/CELL-FM`](https://huggingface.co/BoHuangLab/CELL-FM), overridable with `CELLFM_MODEL_REPO`. That repo keeps one subfolder per model family: | File | Model | |---|---| | `condenseq/cellfm_seq2img.bin` | CELL-FM CS generator, ESM-C 600M encoder included | | `condenseq/vae.bin` | 160×160 image VAE | | `condenseq/vit_cls.bin` | Condensed/diffuse ViT classifier | To run against a local copy instead, set `CELLFM_LOCAL_WEIGHTS` to a **flat** directory holding the three `.bin` files — the prefix is stripped for local paths. `upload_weights.py` pushes them from the cluster. ## Provenance Ported from the CELL-FM repository, one module per stage, with behaviour preserved: | Stage | Origin | |---|---| | Generation | `cell_fm/tasks_local/cell_fm_cs/evaluate_seq2img_dict.py` | | Classification | `cell_fm/tasks_local/vit_cls_condenseq_img/evaluate_single_img.py` | | Moving average | `cell_fm/tasks_local/vit_cls_condenseq_img/ma_plot.py` | | AUC / AAC | `cell_diff/tasks/analysis/{utils,ana_all_mutation_log_scale}.py` | Model hyperparameters mirror `scripts/cell_fm_cs/evaluate_seq2img.sh` and `scripts/vit_cls_condenseq_img/pretrain.sh`. The metric port in `metrics.py` was checked against the original on 40 stored curves and agrees to 0.0 in both log and linear mode. Generated images make the same round-trip through uint16 that the offline pipeline's TIFF write imposed, so demo numbers stay comparable to published ones. ### Reference values, and how noisy they are Measured at the fixed settings (512 images, 100 steps, **seed 6**) on an A40: | | AUC | AAC | c_sat | |---|---|---|---| | NUP98 WT | 0.4912 | 0.0034 | 360 | | 1F→S mutant | 0.4794 | 0.0035 | 412 | These are single draws, and how much that matters depends on whether you pair. **Taken alone, the numbers move a lot.** Over seeds 6–9 the WT AUC ranges 0.4660–0.5087 (sd 0.019) and c_sat ranges 289–531 (sd 102). Treat any individual c_sat or AAC as indicative rather than precise. **Compared at a matched seed, small differences hold up.** The 1F→S effect is −0.0123 ± 0.0042 across four seeds and negative in every one, even though it is smaller than the marginal spread above — driving both sequences from the same seed cancels most of the sampling noise, so the comparison is far tighter than either number is on its own. Compare sequences at the same seed and a single-residue change is resolvable; the demo does this by construction, with the seed fixed at 6. Measured on the cluster stack (torch 2.4.1); the Space runs torch 2.8.0, which is not expected to change these but has not been compared directly. --- ## Citation If you use CELL-FM, please cite: ```bibtex @article{zheng2026virtual, title = {Virtual experiments bridge sequence and microscopy with generative models}, author = {Zheng, Dihan and Hong, Kibeom and Huang, Bo}, journal = {bioRxiv}, year = {2026}, doi = {10.64898/2026.09.13.751243}, url = {https://www.biorxiv.org/content/10.64898/2026.09.13.751243v1} } ```