BoHuangLab commited on
Commit
ba01cc2
Β·
verified Β·
1 Parent(s): ece3f79

Simplify card intro and drop the Source checkpoint column

Browse files
Files changed (1) hide show
  1. README.md +25 -25
README.md CHANGED
@@ -11,32 +11,32 @@ tags:
11
 
12
  # CELL-FM weights
13
 
14
- Checkpoints behind the [CELL-FM CondenSeq demo](https://huggingface.co/spaces/BoHuangLab/CELL-FM).
15
 
16
- | File | Model | Source checkpoint |
17
- |---|---|---|
18
- | `condenseq/cellfm_seq2img.bin` | CELL-FM CS sequence-to-image generator (includes the ESM-C 600M encoder) | `pretrain_condenseq/cellfm_seq2img/checkpoint-50000` |
19
- | `condenseq/vae.bin` | Image VAE, 160x160, 3 down blocks, 4 latent channels | `pretrain_condenseq/vae/checkpoint-50000` |
20
- | `condenseq/vit_cls.bin` | ViT condensed/diffuse classifier, 2-channel 160x160 input | `PT_CondenSeq_img_ViT_cls_R1/checkpoint-10000` |
21
- | `condenseq/reference_nucleus.npy` | The DAPI channel every CondenSeq generation is conditioned on: `(1, 160, 160)` float32 in [-1, 1]. CondenSeq protein index 12626, image 0 β€” the same one the offline seq2img runs used | built |
22
- | `hpa/cellfm_seq2img.bin` | CELL-FM virtual-staining generator for HPA, 256x256, 3-channel conditioning (includes the ESM-C 600M encoder) | `pretrain_hpa/cellfm_seq2img/checkpoint-50000` |
23
- | `hpa/vae.bin` | Image VAE at 256x256, the one `hpa/cellfm_seq2img.bin` was trained against | `pretrain_hpa/vae/checkpoint-50000` |
24
- | `hpa/vit.bin` | Image-embedding ViT for HPA: 12 layers, 512 hidden, 2048 MLP, 8 heads, patch 4, 256x256 input. Trained to identify the protein in an image over 13,908 classes; the embedding is the representation, not the prediction. Four input channels β€” protein, nucleus, ER, microtubules | `pretrain_hpa/vit/checkpoint-16000` |
25
- | `hpa/cellfm_img2seq.bin` | CELL-FM image-to-sequence model for HPA, 512x512, 3-channel conditioning (includes the ESM-C 600M encoder) | `pretrain_hpa/cellfm_img2seq/checkpoint-60000` |
26
- | `hpa/vae_512.bin` | Image VAE at 512x512, the one `hpa/cellfm_img2seq.bin` was trained against β€” a different model from `hpa/vae.bin`, not a rename | `pretrain_hpa/vae_512/checkpoint-50000` |
27
- | `opencell/cellfm_vs.bin` | CELL-FM virtual-staining generator for OpenCell, 256x256, single nucleus conditioning channel, fine-tuned (includes the ESM-C 600M encoder) | `finetune_opencell/cellfm_vs/checkpoint-100000` |
28
- | `opencell/vae.bin` | Image VAE at 256x256, the one `opencell/cellfm_vs.bin` was trained against β€” OpenCell-finetuned, so not interchangeable with `hpa/vae.bin` despite the matching shape | `finetune_opencell/vae/checkpoint-50000` |
29
- | `opencell/vit.bin` | The same ViT fine-tuned on OpenCell, 1,311 classes. Identical backbone to `hpa/vit.bin`, but its input stem takes **two** channels β€” protein and nucleus β€” where the HPA one takes four, so the two cannot be swapped: the count is fixed in `conv_proj` and a mismatch fails there | `finetune_opencell/vit/checkpoint-10000` |
30
- | `opencell/vs_anchor_nucleus.npy` | The nucleus every OpenCell generation is conditioned on: `(1, 256, 256)` float32 in [-1, 1]. Gene ATG7, crop `CID001813_FID00035838_proj_11` β€” bit-for-bit the conditioning channel of the published offline run | built |
31
- | `opencell/vs_genes.csv` | OpenCell's 1,311 genes: name, protein name, UniProt accession, Ensembl id, localization annotation and sequence. The metadata table minus its image paths | built |
32
- | `opencell/vs_reference_cells.npz` | Four proteins in two matched pairs (POLR1A/SNRPF nuclear, LSM14A/DDX6 both P-body), each with one real OpenCell crop as a `(nucleus, protein)` float16 pair β€” the image shown beside the generated one | built |
33
- | `hpa/anchor_cell.npy` | The fixed cell every NLS-screening image is conditioned on: `(3, 256, 256)` float32 in [-1, 1], channels nucleus, ER, microtubules. HPA gene H3C13, cell crop `1194_B2_2_4` | built |
34
- | `hpa/anchor_masks.npz` | Two 256x256 boolean masks over that cell, `nucleus` and `cell`; cytoplasm is `cell & ~nucleus` | built |
35
- | `hpa/pls_anchor_nls.npz` | The cell PLS generation conditions on for nuclear signals: `cell` `(3, 512, 512)` nucleus/ER/microtubules and `protein` `(1, 512, 512)`, float32 in [-1, 1]. HPA gene PPM1G (Nucleoplasm), crop `392_B9_1_11` | built |
36
- | `hpa/pls_anchor_nes.npz` | The same for export signals. HPA gene DIAPH1 (Cytosol, Plasma membrane), crop `1608_B3_1_1` | built |
37
- | `hpa/proteome_aa_counts.json` | Residue counts over the 12,894 HPA proteins (7,940,784 residues), the proteome baseline the frequency analysis compares against | built |
38
- | `hpa/pls_reference_nls.csv` | The 320 published NLS signals: 20 independent draws at each of 16 tail lengths, 10-25 aa | `output/hpa/pls_generation/nls` |
39
- | `hpa/pls_reference_nes.csv` | The same 320 for export signals | `output/hpa/pls_generation/nes` |
40
 
41
  Hyperparameters for the CondenSeq models are set in `pipeline.py` in the Space and mirror
42
  `scripts/cell_fm_cs/evaluate_seq2img.sh` and
 
11
 
12
  # CELL-FM weights
13
 
14
+ Checkpoints of CELL-FM.
15
 
16
+ | File | Model |
17
+ |---|---|
18
+ | `condenseq/cellfm_seq2img.bin` | CELL-FM CS sequence-to-image generator (includes the ESM-C 600M encoder) |
19
+ | `condenseq/vae.bin` | Image VAE, 160x160, 3 down blocks, 4 latent channels |
20
+ | `condenseq/vit_cls.bin` | ViT condensed/diffuse classifier, 2-channel 160x160 input |
21
+ | `condenseq/reference_nucleus.npy` | The DAPI channel every CondenSeq generation is conditioned on: `(1, 160, 160)` float32 in [-1, 1]. CondenSeq protein index 12626, image 0 β€” the same one the offline seq2img runs used |
22
+ | `hpa/cellfm_seq2img.bin` | CELL-FM virtual-staining generator for HPA, 256x256, 3-channel conditioning (includes the ESM-C 600M encoder) |
23
+ | `hpa/vae.bin` | Image VAE at 256x256, the one `hpa/cellfm_seq2img.bin` was trained against |
24
+ | `hpa/vit.bin` | Image-embedding ViT for HPA: 12 layers, 512 hidden, 2048 MLP, 8 heads, patch 4, 256x256 input. Trained to identify the protein in an image over 13,908 classes; the embedding is the representation, not the prediction. Four input channels β€” protein, nucleus, ER, microtubules |
25
+ | `hpa/cellfm_img2seq.bin` | CELL-FM image-to-sequence model for HPA, 512x512, 3-channel conditioning (includes the ESM-C 600M encoder) |
26
+ | `hpa/vae_512.bin` | Image VAE at 512x512, the one `hpa/cellfm_img2seq.bin` was trained against β€” a different model from `hpa/vae.bin`, not a rename |
27
+ | `opencell/cellfm_vs.bin` | CELL-FM virtual-staining generator for OpenCell, 256x256, single nucleus conditioning channel, fine-tuned (includes the ESM-C 600M encoder) |
28
+ | `opencell/vae.bin` | Image VAE at 256x256, the one `opencell/cellfm_vs.bin` was trained against β€” OpenCell-finetuned, so not interchangeable with `hpa/vae.bin` despite the matching shape |
29
+ | `opencell/vit.bin` | The same ViT fine-tuned on OpenCell, 1,311 classes. Identical backbone to `hpa/vit.bin`, but its input stem takes **two** channels β€” protein and nucleus β€” where the HPA one takes four, so the two cannot be swapped: the count is fixed in `conv_proj` and a mismatch fails there |
30
+ | `opencell/vs_anchor_nucleus.npy` | The nucleus every OpenCell generation is conditioned on: `(1, 256, 256)` float32 in [-1, 1]. Gene ATG7, crop `CID001813_FID00035838_proj_11` β€” bit-for-bit the conditioning channel of the published offline run |
31
+ | `opencell/vs_genes.csv` | OpenCell's 1,311 genes: name, protein name, UniProt accession, Ensembl id, localization annotation and sequence. The metadata table minus its image paths |
32
+ | `opencell/vs_reference_cells.npz` | Four proteins in two matched pairs (POLR1A/SNRPF nuclear, LSM14A/DDX6 both P-body), each with one real OpenCell crop as a `(nucleus, protein)` float16 pair β€” the image shown beside the generated one |
33
+ | `hpa/anchor_cell.npy` | The fixed cell every NLS-screening image is conditioned on: `(3, 256, 256)` float32 in [-1, 1], channels nucleus, ER, microtubules. HPA gene H3C13, cell crop `1194_B2_2_4` |
34
+ | `hpa/anchor_masks.npz` | Two 256x256 boolean masks over that cell, `nucleus` and `cell`; cytoplasm is `cell & ~nucleus` |
35
+ | `hpa/pls_anchor_nls.npz` | The cell PLS generation conditions on for nuclear signals: `cell` `(3, 512, 512)` nucleus/ER/microtubules and `protein` `(1, 512, 512)`, float32 in [-1, 1]. HPA gene PPM1G (Nucleoplasm), crop `392_B9_1_11` |
36
+ | `hpa/pls_anchor_nes.npz` | The same for export signals. HPA gene DIAPH1 (Cytosol, Plasma membrane), crop `1608_B3_1_1` |
37
+ | `hpa/proteome_aa_counts.json` | Residue counts over the 12,894 HPA proteins (7,940,784 residues), the proteome baseline the frequency analysis compares against |
38
+ | `hpa/pls_reference_nls.csv` | The 320 published NLS signals: 20 independent draws at each of 16 tail lengths, 10-25 aa |
39
+ | `hpa/pls_reference_nes.csv` | The same 320 for export signals |
40
 
41
  Hyperparameters for the CondenSeq models are set in `pipeline.py` in the Space and mirror
42
  `scripts/cell_fm_cs/evaluate_seq2img.sh` and