BoHuangLab commited on
Commit
c07ba06
·
verified ·
1 Parent(s): fd76ce4

Document the hpa/ family on the card

Browse files
Files changed (1) hide show
  1. README.md +14 -27
README.md CHANGED
@@ -1,5 +1,4 @@
1
  ---
2
- license: mit
3
  library_name: cell-fm
4
  tags:
5
  - biology
@@ -9,34 +8,22 @@ tags:
9
  - flow-matching
10
  ---
11
 
12
- # CELL-FM
13
 
14
- Model checkpoints for **CELL-FM**, a flow-matching model that generates cell
15
- microscopy images from protein sequence.
16
-
17
- One subfolder per model family. Point a loader at the subfolder it needs; the
18
- demo code does this through `CELLFM_MODEL_REPO` + a path prefix.
19
-
20
- ## `condenseq/` — sequence-conditioned condensate imaging
21
-
22
- Trained on CondenSeq, 160x160 GFP images. Drives the condensate titration demo:
23
- sequence in, condensate-probability curve out, integrated into AUC and AAC.
24
 
25
  | File | Model | Source checkpoint |
26
  |---|---|---|
27
- | `condenseq/cellfm_seq2img.bin` | CELL-FM CS sequence-to-image generator, ESM-C 600M encoder included (745 M params) | `pretrain_condenseq/cellfm_seq2img/checkpoint-50000` |
28
  | `condenseq/vae.bin` | Image VAE, 160x160, 3 down blocks, 4 latent channels | `pretrain_condenseq/vae/checkpoint-50000` |
29
- | `condenseq/vit_cls.bin` | ViT condensed/diffuse classifier, 2-channel 160x160 input (26 M params) | `PT_CondenSeq_img_ViT_cls_R1/checkpoint-10000` |
30
-
31
- Hyperparameters mirror `scripts/cell_fm_cs/evaluate_seq2img.sh` and
32
- `scripts/vit_cls_condenseq_img/pretrain.sh` in the CELL-FM repository, and are set
33
- in `huggingface_space/pipeline.py` there.
34
-
35
- Reference values at 512 images / 100 ODE steps / seed 6: NUP98 WT gives AUC 0.4896
36
- and AAC 0.0001; its 1F->S mutant gives AUC 0.4561.
37
-
38
- ## Adding a model family
39
-
40
- Upload under a new prefix (`hpa/`, `opencell_3d/`, ...) and add a section here.
41
- `upload_weights.py` in the demo takes a `{path_in_repo: local_checkpoint}` map, so
42
- new families need only a new entry.
 
1
  ---
 
2
  library_name: cell-fm
3
  tags:
4
  - biology
 
8
  - flow-matching
9
  ---
10
 
11
+ # CELL-FM weights
12
 
13
+ Checkpoints behind the [CELL-FM CondenSeq demo](https://huggingface.co/spaces/BoHuangLab/CELL-FM).
 
 
 
 
 
 
 
 
 
14
 
15
  | File | Model | Source checkpoint |
16
  |---|---|---|
17
+ | `condenseq/cellfm_seq2img.bin` | CELL-FM CS sequence-to-image generator (includes the ESM-C 600M encoder) | `pretrain_condenseq/cellfm_seq2img/checkpoint-50000` |
18
  | `condenseq/vae.bin` | Image VAE, 160x160, 3 down blocks, 4 latent channels | `pretrain_condenseq/vae/checkpoint-50000` |
19
+ | `condenseq/vit_cls.bin` | ViT condensed/diffuse classifier, 2-channel 160x160 input | `PT_CondenSeq_img_ViT_cls_R1/checkpoint-10000` |
20
+ | `hpa/cellfm_seq2img.bin` | CELL-FM virtual-staining generator for HPA, 256x256, 3-channel conditioning (includes the ESM-C 600M encoder) | `pretrain_hpa/cellfm_seq2img/checkpoint-50000` |
21
+ | `hpa/vae.bin` | Image VAE, 256x256, 3 down blocks, 4 latent channels | `pretrain_hpa/vae/checkpoint-50000` |
22
+ | `hpa/anchor_cell.npy` | The fixed cell every NLS-screening image is conditioned on: `(3, 256, 256)` float32 in [-1, 1], channels nucleus, ER, microtubules. HPA gene H3C13, cell crop `1194_B2_2_4` | built |
23
+ | `hpa/anchor_masks.npz` | Two 256x256 boolean masks over that cell, `nucleus` and `cell`; cytoplasm is `cell & ~nucleus` | built |
24
+
25
+ Hyperparameters for the CondenSeq models are set in `pipeline.py` in the Space and mirror
26
+ `scripts/cell_fm_cs/evaluate_seq2img.sh` and
27
+ `scripts/vit_cls_condenseq_img/pretrain.sh` in the CELL-FM repository. The HPA
28
+ hyperparameters are spelled out in `notebooks/nls_screening.ipynb` and mirror
29
+ `scripts/cell_fm/evaluate_virtual_staining_hpa_dict.sh`.