Spaces:
Running on Zero
Running on Zero
|
Download README.md from BoHuangLab/CELL-FM: direct link, hf CLI and curl.
- Browser
- Download file 9.14 kB
-
https://huggingface.co/spaces/BoHuangLab/CELL-FM/resolve/main/README.md
- Command line
-
hf download hf://spaces/BoHuangLab/CELL-FM/README.md
-
curl -L -o README.md https://huggingface.co/spaces/BoHuangLab/CELL-FM/resolve/main/README.md
9.14 kB
| license: mit | |
| title: CELL-FM | |
| emoji: 🧬 | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 5.9.1 | |
| app_file: app.py | |
| pinned: false | |
| short_description: Cell microscopy images from protein sequence | |
| # CELL-FM | |
| A virtual microscopy model that bridges microscopy images and protein sequences. | |
| Each tab in the app is one **application** of the model. More will be added; the | |
| sections below document the ones that exist. | |
| --- | |
| ## Condensate Titration | |
| Predict how an Intrinsically Disordered Peptide (IDP) behaves as its | |
| concentration rises, directly from sequence. | |
|  | |
| The only input is the sequence, and it must be **exactly 66 amino acids** — the | |
| entire CondenSeq library the model was trained on is 66-mers (all 14,578 of them), | |
| so other lengths are off-distribution. Sampling is fixed at 512 concentration | |
| levels and 100 ODE steps, seed 6, matching the offline runs. | |
| Enter a sequence and this section runs the three stages of the CELL-FM condensate | |
| pipeline end to end: | |
| 1. **Generate.** CELL-FM CS samples one microscopy image per protein-intensity | |
| level, over a log-spaced concentration ladder from 20 to 2048 a.u. Every image | |
| is conditioned on the same fixed reference nucleus, so curves are comparable | |
| across sequences. | |
| 2. **Classify.** A ViT calls each generated image **condensed** or **diffuse**. | |
| 3. **Summarise.** A centred moving average over those binary calls is the | |
| **titration curve**; integrating it against log-intensity gives the two numbers | |
| the demo reports. | |
| ### What AUC and AAC mean | |
| **AUC** — the area under the titration curve. How much of the scanned concentration | |
| range the sequence spends in the condensed state. Higher means it condenses more | |
| readily, over a wider range. | |
| **AAC** — the area *above* the curve. The curve is compared against the one it would | |
| have followed without **reentrant dissolution**: everything past the curve's maximum | |
| is held flat at that maximum, and AAC is the area between the two. It is zero when | |
| the sequence condenses and stays condensed, and grows the more the sequence | |
| re-dissolves at high concentration. | |
| Both are normalised by the width of the log-intensity range, so both lie in [0, 1] | |
| and are directly comparable between sequences. | |
| Also reported: **c_sat**, the first intensity at which the smoothed curve reaches | |
| probability 0.8, and the raw per-image calls as a downloadable CSV. | |
| ### Default sequence | |
| The box is pre-filled with the **NUP98 IDP**, wild type. Its mutation panel — 92 | |
| sequences of 66 aa covering F→S, F→Y, charge and spacer variants — ships in | |
| `assets/nup98_mutation_seqs.csv`; paste one in to compare its AUC against WT, which | |
| is the quickest way to see the model respond to sequence grammar. | |
| --- | |
| ## Virtual OpenCell | |
| Browse CELL-FM's virtual staining of 1,276 OpenCell proteins, laid out like an | |
| [OpenCell](https://opencell.sf.czbiohub.org/) target page. Search for a protein: | |
| - **Left** is OpenCell's annotation of the *real* cell line: localization grades | |
| (bar height is the grade, 3 the most prominent), UniProt and Ensembl IDs, sequence | |
| length, HEK293 abundance, and a link to the real images on OpenCell. None of it is | |
| predicted. | |
| - **Right** are CELL-FM's images, in OpenCell's colours: the nucleus blue and the | |
| generated protein (the *target*) gray. OpenCell's channel switch and per-channel | |
| intensity range and gamma adjust the view; the strip below holds every sample of | |
| the protein, and **Download TIFF** saves the selected one as the raw 2-channel image. | |
| Below the annotations, the **virtual staining map** places every protein's samples by | |
| UMAP of their image embeddings (`opencell/vit.bin`), coloured by OpenCell localization in | |
| the palette of the paper figures; the selected protein's samples are outlined and | |
| labelled. Zoom, pan and hover work as in any Plotly chart, and clicking a legend entry | |
| hides that category. The map is `opencell/vs_umap_embedding.npz` in the weights repo; | |
| the backdrop draws 12 evenly spaced samples per protein to keep each update light. | |
| Every image is generated from the protein sequence in **one shared cell**: the nucleus | |
| channel is the same real nucleus (ATG7, crop 1) for every protein, so differences | |
| between proteins come from the sequence, and a protein's samples are independent | |
| draws for that cell. | |
| The images are the dataset | |
| [`BoHuangLab/CELL-FM`](https://huggingface.co/datasets/BoHuangLab/CELL-FM): | |
| `OpenCell/virtual_staining/<GENE>/<NNNN>.tif`, 93,510 samples, each (2, 256, 256) | |
| uint16 with the nucleus first. A protein's samples are fetched the first time it is | |
| viewed (`VIRTUAL_OPENCELL_LOCAL_DIR` reads a local copy instead); while that repo is | |
| private, the Space needs an `HF_TOKEN` secret that can read it. The protein table, | |
| `assets/virtual_opencell_proteins.csv`, is built by | |
| `notebooks/tools/build_virtual_opencell_assets.py`. This section runs on CPU. | |
| --- | |
| ## Hardware | |
| **A GPU is required in practice.** The code itself runs on CPU — `flash-attn` is | |
| deliberately left out of `requirements.txt`, so ESM-C uses its pure-torch rotary | |
| fallback (verified bit-identical to the Triton path: same AUC, same AAC, same | |
| 256/256 per-image calls) — but a CPU run of the defaults would take hours and | |
| time out. Use ZeroGPU or a GPU hardware tier. | |
| The fixed settings — 512 images at 100 ODE steps — take **~85 s on an A40** | |
| (measured: 11.5 s per batch of 64, plus ~9 s of one-off model loading). Runtime | |
| scales linearly in both, so a section that needs to be faster should lower them in | |
| `pipeline.py` (`NUM_IMAGES`, `NUM_STEPS`). | |
| On ZeroGPU the per-call budget is set by `ZEROGPU_DURATION` (default 300 s), which | |
| covers the defaults with room to spare. | |
| --- | |
| ## Running it without a Space | |
| Creating a Gradio Space under an organization now needs a Team or Enterprise plan. | |
| Until `HuangLab` has one, `run_local.sh` serves the identical app from a cluster GPU | |
| over a temporary public URL: | |
| ```bash | |
| sbatch run_local.sh # prints a *.gradio.live URL, valid 72 h | |
| ``` | |
| ## Weights | |
| The Space downloads ~2.1 GB of checkpoints at startup from | |
| [`BoHuangLab/CELL-FM`](https://huggingface.co/BoHuangLab/CELL-FM), overridable with | |
| `CELLFM_MODEL_REPO`. That repo keeps one subfolder per model family: | |
| | File | Model | | |
| |---|---| | |
| | `condenseq/cellfm_seq2img.bin` | CELL-FM CS generator, ESM-C 600M encoder included | | |
| | `condenseq/vae.bin` | 160×160 image VAE | | |
| | `condenseq/vit_cls.bin` | Condensed/diffuse ViT classifier | | |
| To run against a local copy instead, set `CELLFM_LOCAL_WEIGHTS` to a **flat** | |
| directory holding the three `.bin` files — the prefix is stripped for local paths. | |
| `upload_weights.py` pushes them from the cluster. | |
| ## Provenance | |
| Ported from the CELL-FM repository, one module per stage, with behaviour preserved: | |
| | Stage | Origin | | |
| |---|---| | |
| | Generation | `cell_fm/tasks_local/cell_fm_cs/evaluate_seq2img_dict.py` | | |
| | Classification | `cell_fm/tasks_local/vit_cls_condenseq_img/evaluate_single_img.py` | | |
| | Moving average | `cell_fm/tasks_local/vit_cls_condenseq_img/ma_plot.py` | | |
| | AUC / AAC | `cell_diff/tasks/analysis/{utils,ana_all_mutation_log_scale}.py` | | |
| Model hyperparameters mirror `scripts/cell_fm_cs/evaluate_seq2img.sh` and | |
| `scripts/vit_cls_condenseq_img/pretrain.sh`. The metric port in `metrics.py` was | |
| checked against the original on 40 stored curves and agrees to 0.0 in both log and | |
| linear mode. Generated images make the same round-trip through uint16 that the | |
| offline pipeline's TIFF write imposed, so demo numbers stay comparable to published | |
| ones. | |
| ### Reference values, and how noisy they are | |
| Measured at the fixed settings (512 images, 100 steps, **seed 6**) on an A40: | |
| | | AUC | AAC | c_sat | | |
| |---|---|---|---| | |
| | NUP98 WT | 0.4912 | 0.0034 | 360 | | |
| | 1F→S mutant | 0.4794 | 0.0035 | 412 | | |
| These are single draws, and how much that matters depends on whether you pair. | |
| **Taken alone, the numbers move a lot.** Over seeds 6–9 the WT AUC ranges | |
| 0.4660–0.5087 (sd 0.019) and c_sat ranges 289–531 (sd 102). Treat any individual | |
| c_sat or AAC as indicative rather than precise. | |
| **Compared at a matched seed, small differences hold up.** The 1F→S effect is | |
| −0.0123 ± 0.0042 across four seeds and negative in every one, even though it is | |
| smaller than the marginal spread above — driving both sequences from the same seed | |
| cancels most of the sampling noise, so the comparison is far tighter than either | |
| number is on its own. Compare sequences at the same seed and a single-residue | |
| change is resolvable; the demo does this by construction, with the seed fixed at 6. | |
| Measured on the cluster stack (torch 2.4.1); the Space runs torch 2.8.0, which is | |
| not expected to change these but has not been compared directly. | |
| --- | |
| ## Citation | |
| If you use CELL-FM, please cite: | |
| ```bibtex | |
| @article{zheng2026virtual, | |
| title = {Virtual experiments bridge sequence and microscopy with generative models}, | |
| author = {Zheng, Dihan and Hong, Kibeom and Huang, Bo}, | |
| journal = {bioRxiv}, | |
| year = {2026}, | |
| doi = {10.64898/2026.09.13.751243}, | |
| url = {https://www.biorxiv.org/content/10.64898/2026.09.13.751243v1} | |
| } | |
| ``` | |