|
Download README.md from ChatterjeeLab/pCoMole: direct link, hf CLI and curl.
- Browser
- Download file 10.8 kB
-
https://huggingface.co/ChatterjeeLab/pCoMole/resolve/main/README.md
- Command line
-
hf download hf://ChatterjeeLab/pCoMole/README.md
-
curl -L -o README.md https://huggingface.co/ChatterjeeLab/pCoMole/resolve/main/README.md
10.8 kB
| # pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows | |
|  | |
| ## Repository layout | |
| ```text | |
| pCoMole/ | |
| train.py / evaluate.py # Edit Flow train / val-loss eval | |
| configs/ # GFP (protein) and SELFIES configs | |
| model/ logic/ flow_matching/ # Edit Flow implementation | |
| smiles_tokenizer/ # SMILES SPE + SELFIES vocab | |
| data/selfies/28k_mimetics/ # shipped SELFIES training split | |
| gfp/ # GFP generation + pCoMole | |
| pcomole.py # official GFP editor | |
| generate.py | |
| ckpt/last_2.ckpt | |
| classifier_ckpt/best.pt | |
| FPredX/ # excitation / brightness / emission models | |
| peptidomimetics/ # peptidomimetic generation + pCoMole | |
| pcomole.py | |
| generate.py | |
| objectives.py # PeptiVerse + Admetica/DeepDTAGen mix | |
| ckpt/SELFIES_EditFlows.ckpt | |
| ckpt/SMILES_BindEvaluator.ckpt | |
| ckpt/admetica/ # LD50, solubility, Caco-2, half-life | |
| ckpt/deepdtagen/ | |
| cas9/ # Cas9 generation + pCoMole | |
| pcomole.py # official Cas9 editor | |
| generate_cas9_csvs.py | |
| objectives.py # Cas9Classification, DeletionCount, PAMMatching | |
| constraints.py | |
| domain_detector.py # HNH / RuvC completeness (HMMER) | |
| pam_detector_hmm.py # PAM-interacting domain masking | |
| model/ logic/ # Cas9 Edit Flow variant (see note below) | |
| data/orthologs.fasta # 10 Cas9 ortholog seeds | |
| ckpt/last.ckpt | |
| classifier_ckpt/ # Cas9-likeness classifier | |
| hmm/ # HMM databases | |
| ``` | |
| > **Note.** `cas9/model/` and `cas9/logic/` are a *fork* of the repo-root | |
| > packages, not duplicates: the Cas9 stack keeps the reparameterized Edit Flow | |
| > models inside `base_models.py`, whereas the root splits them into | |
| > `reparam_models.py`. The two are not interchangeable, so the Cas9 task | |
| > imports its own copies (`from cas9.model...`). | |
| ## 1. Environment | |
| A CUDA GPU is strongly recommended. | |
| ```bash | |
| git clone <this-repo-url> pCoMole | |
| cd pCoMole | |
| python -m venv .venv | |
| source .venv/bin/activate | |
| pip install -r requirements.txt | |
| ``` | |
| HuggingFace model downloads happen on first use (`facebook/esm2_t33_650M_UR50D`, `aaronfeller/PeptideCLM-23M-all`, ChemBERTa). | |
| ### PeptiVerse (required for peptidomimetic pCoMole) | |
| Peptidomimetic property oracles use the latest [PeptiVerse](https://huggingface.co/ChatterjeeLab/PeptiVerse) SMILES predictors. Clone it **next to** this repo, or set `PEPTIVERSE_ROOT`: | |
| ```bash | |
| # recommended: sibling of pCoMole | |
| cd .. | |
| git clone https://huggingface.co/ChatterjeeLab/PeptiVerse | |
| cd PeptiVerse | |
| pip install -r requirements.txt | |
| # follow PeptiVerse README to download model weights | |
| ``` | |
| pCoMole looks for PeptiVerse in this order: | |
| 1. `$PEPTIVERSE_ROOT` | |
| 2. `../PeptiVerse` relative to this repository | |
| PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged. | |
| ### MAFFT (required for GFP pCoMole) | |
| GFP excitation, brightness, and emission oracles run through FPredX and need [MAFFT](https://mafft.cbrc.jp/alignment/software/) on your `PATH`. | |
| ```bash | |
| # conda | |
| conda install -c bioconda mafft | |
| # or point to a local binary | |
| export MAFFT_PATH=/path/to/mafft | |
| ``` | |
| Confirm with `mafft --version` before running GFP editing. | |
| ### HMMER and protein2pam (required for Cas9 pCoMole) | |
| The Cas9 domain-completeness and PAM-masking oracles shell out to `hmmscan`: | |
| ```bash | |
| conda install -c bioconda hmmer | |
| ``` | |
| The PAM objective needs the `protein2pam` model class (weights download from the | |
| Hub as `Profluent-Bio/protein2pam-cas9_full`): | |
| ```bash | |
| pip install protein2pam==0.2.0 | |
| ``` | |
| ## 2. Train, evaluate, and sample Edit Flows | |
| Run every command from the **repository root**. | |
| ### SELFIES peptidomimetic Edit Flow | |
| Training data is shipped at `data/selfies/28k_mimetics`. | |
| ```bash | |
| # train | |
| python train.py --config configs/config_selfies.yaml | |
| # optional: python train.py --config configs/config_selfies.yaml --wandb | |
| # evaluate a checkpoint on the validation split | |
| python evaluate.py \ | |
| --config configs/config_selfies.yaml \ | |
| --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt | |
| # unconditional / seeded generation | |
| python peptidomimetics/generate.py \ | |
| --config configs/config_selfies.yaml \ | |
| --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \ | |
| --input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \ | |
| --num_steps 30 \ | |
| --num_samples 2 \ | |
| --output_csv outputs/peptidomimetic_unconditional.csv | |
| ``` | |
| Or: `bash peptidomimetics/scripts/train.sh` and `bash peptidomimetics/scripts/generate.sh`. | |
| ### GFP protein Edit Flow | |
| A trained GFP checkpoint is shipped at `gfp/ckpt/last_2.ckpt`. The original GFP training arrows are **not** included. To retrain, put a HuggingFace `load_from_disk` dataset under `data/gfp/{train,validation}` (or edit `configs/config_gfp.yaml`). | |
| ```bash | |
| python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt | |
| python gfp/generate.py \ | |
| --config configs/config_gfp.yaml \ | |
| --ckpt gfp/ckpt/last_2.ckpt \ | |
| --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \ | |
| --num_steps 10 \ | |
| --num_samples 8 \ | |
| --output_csv outputs/gfp_unconditional.csv | |
| ``` | |
| ## 3. pCoMole multi-objective editing | |
| ### GFP | |
| Official editor: `gfp/pcomole.py` (length + excitation + brightness, with GFP-classifier and emission constraints). | |
| ```bash | |
| bash gfp/scripts/pcomole.sh | |
| ``` | |
| Equivalent command: | |
| ```bash | |
| python gfp/pcomole.py \ | |
| --root_dir gfp/FPredX \ | |
| --config configs/config_gfp.yaml \ | |
| --ckpt gfp/ckpt/last_2.ckpt \ | |
| --num_steps 10 \ | |
| --num_candidates 50 \ | |
| --num_rollouts 10 \ | |
| --objective_weights 3 1 1 \ | |
| --output_file outputs/gfp_length_excitation_brightness.csv \ | |
| --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' | |
| ``` | |
| If the output filename contains `length_excitation_brightness`, `length_brightness`, `length_excitation`, or `length`, the matching objective subset is used. Any other name uses all three objectives. | |
| ### Peptidomimetics | |
| Official editor: `peptidomimetics/pcomole.py`. | |
| Objectives (7 scores when `--specificity` is on): | |
| 1. non-toxicity | |
| 2. solubility | |
| 3. permeability | |
| 4. half-life | |
| 5. affinity | |
| 6. motif | |
| 7. specificity | |
| PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint. | |
| ```bash | |
| # after PeptiVerse is cloned and its weights are downloaded | |
| bash peptidomimetics/scripts/pcomole.sh | |
| ``` | |
| `--target` and `--motifs` are required for affinity and motif oracles. `--specificity` adds the seventh objective. | |
| ### Cas9 | |
| Official editor: `cas9/pcomole.py`. Three objectives — Cas9-likeness, deletion | |
| count, and PAM matching — under Cas9-score, PAM-matching and length constraints. | |
| Cas9 code is a package, so run it from the repository root with `-m`: | |
| ```bash | |
| ORTHOLOG=cjcas9 bash cas9/scripts/pcomole.sh | |
| ``` | |
| Equivalent command: | |
| ```bash | |
| python -m cas9.pcomole \ | |
| --pcomol_config configs/config_cas9.yaml \ | |
| --ckpt cas9/ckpt/last.ckpt \ | |
| --input "$(awk -v n='>cjcas9' '$1==n{f=1;next}/^>/{f=0}f{printf "%s",$0}' cas9/data/orthologs.fasta)" \ | |
| --cas9_objective --deletion_count_obj --pam_distr_objective \ | |
| --objective_weights 1 1 1 \ | |
| --target_pam matching \ | |
| --num_steps 5 --num_rollouts 5 --num_candidates 10 --num_sequences 50 \ | |
| --max_deletion_percentage 0.2 --deletion_rate_scale 2500 \ | |
| --pam_min_confidence 0.7 --cas9_score_threshold 0.5 \ | |
| --max_target_length 930 --min_target_length 890 \ | |
| --use_ce_loss --allow_equal_length --zero_lam_ins \ | |
| --output_file outputs/cas9_cjcas9_pcomole.csv | |
| ``` | |
| Ten ortholog seeds ship in `cas9/data/orthologs.fasta` (SpCas9, SaCas9, CjCas9, | |
| St1/St3Cas9, GeoCas9, NmCas9, iSpyMac, SpCas9++, SpRYc). Set `ORTHOLOG` to any | |
| of them. `--target_pam matching` optimizes toward the seed's own PAM; pass an | |
| explicit 10-nt `ACGTN` string to retarget. | |
| ## 4. Included checkpoints | |
| | Path | Role | | |
| |---|---| | |
| | `gfp/ckpt/last_2.ckpt` | GFP Edit Flow | | |
| | `gfp/classifier_ckpt/best.pt` | GFP hard constraint | | |
| | `gfp/FPredX/{ex,bright,em}_model` | FPredX property models | | |
| | `peptidomimetics/ckpt/SELFIES_EditFlows.ckpt` | peptidomimetic Edit Flow | | |
| | `peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt` | motif / specificity | | |
| | `peptidomimetics/ckpt/admetica/*.ckpt` | Admetica blending oracles | | |
| | `peptidomimetics/ckpt/deepdtagen/` | DeepDTAGen affinity | | |
| | `cas9/ckpt/last.ckpt` | Cas9 Edit Flow | | |
| | `cas9/classifier_ckpt/` | Cas9-likeness classifier (objective + hard constraint) | | |
| | `cas9/hmm/` | HNH / RuvC and PAM-interacting HMM databases | | |
| Together these are about 10 GB. HuggingFace uploads should use Git LFS for the `.ckpt` / `.pt` / `.pth` files. | |
| ## 5. Environment variables | |
| | Variable | Purpose | | |
| |---|---| | |
| | `PEPTIVERSE_ROOT` | PeptiVerse repo root (manifest + `training_classifiers/`) | | |
| | `MAFFT_PATH` | MAFFT binary, if it is not on `PATH` | | |
| | `GFP_CLASSIFIER_CKPT` | optional override for the GFP classifier | | |
| | `PROTEIN2PAM_ROOT` | protein2pam package root (PAM oracle model class) | | |
| | `CAS_PREDICTOR_ROOT` | cas_predictor package root (Cas9 classifier code) | | |
| | `CAS9_CLASSIFIER_CKPT` | override for the Cas9 classifier checkpoint | | |
| | `CAS9_CLASSIFIER_CONFIG` | optional Cas9 classifier config | | |
| | `CAS9_HMM_DIR` | directory holding the Cas9 HMM databases | | |
| ## 6. Notes | |
| - Run scripts from the repo root, or use the wrappers under `*/scripts/`. | |
| - GFP FPredX will fail immediately if `mafft` is missing; that is expected. | |
| - Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded. | |
| - Cas9 objectives fail at init if `hmmscan`, `protein2pam`, or `cas_predictor` are missing. | |
| - Training writes to `outputs/` by default. | |
| ## Citation | |
| If you use this code, please cite the pCoMole paper and [PeptiVerse](https://www.biorxiv.org/content/10.64898/2025.12.31.697180). | |