pCoMole / README.md
pranamanam's picture
Update README.md
3086e08 verified
|
Raw History Blame Contribute Delete
10.8 kB
# pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows
![pCoMole_overview_10s](https://cdn-uploads.huggingface.co/production/uploads/64cd5b3f0494187a9e8b7c69/_wLLPKFMOpANDzjc85xSZ.gif)
## Repository layout
```text
pCoMole/
train.py / evaluate.py # Edit Flow train / val-loss eval
configs/ # GFP (protein) and SELFIES configs
model/ logic/ flow_matching/ # Edit Flow implementation
smiles_tokenizer/ # SMILES SPE + SELFIES vocab
data/selfies/28k_mimetics/ # shipped SELFIES training split
gfp/ # GFP generation + pCoMole
pcomole.py # official GFP editor
generate.py
ckpt/last_2.ckpt
classifier_ckpt/best.pt
FPredX/ # excitation / brightness / emission models
peptidomimetics/ # peptidomimetic generation + pCoMole
pcomole.py
generate.py
objectives.py # PeptiVerse + Admetica/DeepDTAGen mix
ckpt/SELFIES_EditFlows.ckpt
ckpt/SMILES_BindEvaluator.ckpt
ckpt/admetica/ # LD50, solubility, Caco-2, half-life
ckpt/deepdtagen/
cas9/ # Cas9 generation + pCoMole
pcomole.py # official Cas9 editor
generate_cas9_csvs.py
objectives.py # Cas9Classification, DeletionCount, PAMMatching
constraints.py
domain_detector.py # HNH / RuvC completeness (HMMER)
pam_detector_hmm.py # PAM-interacting domain masking
model/ logic/ # Cas9 Edit Flow variant (see note below)
data/orthologs.fasta # 10 Cas9 ortholog seeds
ckpt/last.ckpt
classifier_ckpt/ # Cas9-likeness classifier
hmm/ # HMM databases
```
> **Note.** `cas9/model/` and `cas9/logic/` are a *fork* of the repo-root
> packages, not duplicates: the Cas9 stack keeps the reparameterized Edit Flow
> models inside `base_models.py`, whereas the root splits them into
> `reparam_models.py`. The two are not interchangeable, so the Cas9 task
> imports its own copies (`from cas9.model...`).
## 1. Environment
A CUDA GPU is strongly recommended.
```bash
git clone <this-repo-url> pCoMole
cd pCoMole
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```
HuggingFace model downloads happen on first use (`facebook/esm2_t33_650M_UR50D`, `aaronfeller/PeptideCLM-23M-all`, ChemBERTa).
### PeptiVerse (required for peptidomimetic pCoMole)
Peptidomimetic property oracles use the latest [PeptiVerse](https://huggingface.co/ChatterjeeLab/PeptiVerse) SMILES predictors. Clone it **next to** this repo, or set `PEPTIVERSE_ROOT`:
```bash
# recommended: sibling of pCoMole
cd ..
git clone https://huggingface.co/ChatterjeeLab/PeptiVerse
cd PeptiVerse
pip install -r requirements.txt
# follow PeptiVerse README to download model weights
```
pCoMole looks for PeptiVerse in this order:
1. `$PEPTIVERSE_ROOT`
2. `../PeptiVerse` relative to this repository
PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged.
### MAFFT (required for GFP pCoMole)
GFP excitation, brightness, and emission oracles run through FPredX and need [MAFFT](https://mafft.cbrc.jp/alignment/software/) on your `PATH`.
```bash
# conda
conda install -c bioconda mafft
# or point to a local binary
export MAFFT_PATH=/path/to/mafft
```
Confirm with `mafft --version` before running GFP editing.
### HMMER and protein2pam (required for Cas9 pCoMole)
The Cas9 domain-completeness and PAM-masking oracles shell out to `hmmscan`:
```bash
conda install -c bioconda hmmer
```
The PAM objective needs the `protein2pam` model class (weights download from the
Hub as `Profluent-Bio/protein2pam-cas9_full`):
```bash
pip install protein2pam==0.2.0
```
## 2. Train, evaluate, and sample Edit Flows
Run every command from the **repository root**.
### SELFIES peptidomimetic Edit Flow
Training data is shipped at `data/selfies/28k_mimetics`.
```bash
# train
python train.py --config configs/config_selfies.yaml
# optional: python train.py --config configs/config_selfies.yaml --wandb
# evaluate a checkpoint on the validation split
python evaluate.py \
--config configs/config_selfies.yaml \
--ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt
# unconditional / seeded generation
python peptidomimetics/generate.py \
--config configs/config_selfies.yaml \
--ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \
--input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \
--num_steps 30 \
--num_samples 2 \
--output_csv outputs/peptidomimetic_unconditional.csv
```
Or: `bash peptidomimetics/scripts/train.sh` and `bash peptidomimetics/scripts/generate.sh`.
### GFP protein Edit Flow
A trained GFP checkpoint is shipped at `gfp/ckpt/last_2.ckpt`. The original GFP training arrows are **not** included. To retrain, put a HuggingFace `load_from_disk` dataset under `data/gfp/{train,validation}` (or edit `configs/config_gfp.yaml`).
```bash
python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt
python gfp/generate.py \
--config configs/config_gfp.yaml \
--ckpt gfp/ckpt/last_2.ckpt \
--input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \
--num_steps 10 \
--num_samples 8 \
--output_csv outputs/gfp_unconditional.csv
```
## 3. pCoMole multi-objective editing
### GFP
Official editor: `gfp/pcomole.py` (length + excitation + brightness, with GFP-classifier and emission constraints).
```bash
bash gfp/scripts/pcomole.sh
```
Equivalent command:
```bash
python gfp/pcomole.py \
--root_dir gfp/FPredX \
--config configs/config_gfp.yaml \
--ckpt gfp/ckpt/last_2.ckpt \
--num_steps 10 \
--num_candidates 50 \
--num_rollouts 10 \
--objective_weights 3 1 1 \
--output_file outputs/gfp_length_excitation_brightness.csv \
--input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS'
```
If the output filename contains `length_excitation_brightness`, `length_brightness`, `length_excitation`, or `length`, the matching objective subset is used. Any other name uses all three objectives.
### Peptidomimetics
Official editor: `peptidomimetics/pcomole.py`.
Objectives (7 scores when `--specificity` is on):
1. non-toxicity
2. solubility
3. permeability
4. half-life
5. affinity
6. motif
7. specificity
PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint.
```bash
# after PeptiVerse is cloned and its weights are downloaded
bash peptidomimetics/scripts/pcomole.sh
```
`--target` and `--motifs` are required for affinity and motif oracles. `--specificity` adds the seventh objective.
### Cas9
Official editor: `cas9/pcomole.py`. Three objectives — Cas9-likeness, deletion
count, and PAM matching — under Cas9-score, PAM-matching and length constraints.
Cas9 code is a package, so run it from the repository root with `-m`:
```bash
ORTHOLOG=cjcas9 bash cas9/scripts/pcomole.sh
```
Equivalent command:
```bash
python -m cas9.pcomole \
--pcomol_config configs/config_cas9.yaml \
--ckpt cas9/ckpt/last.ckpt \
--input "$(awk -v n='>cjcas9' '$1==n{f=1;next}/^>/{f=0}f{printf "%s",$0}' cas9/data/orthologs.fasta)" \
--cas9_objective --deletion_count_obj --pam_distr_objective \
--objective_weights 1 1 1 \
--target_pam matching \
--num_steps 5 --num_rollouts 5 --num_candidates 10 --num_sequences 50 \
--max_deletion_percentage 0.2 --deletion_rate_scale 2500 \
--pam_min_confidence 0.7 --cas9_score_threshold 0.5 \
--max_target_length 930 --min_target_length 890 \
--use_ce_loss --allow_equal_length --zero_lam_ins \
--output_file outputs/cas9_cjcas9_pcomole.csv
```
Ten ortholog seeds ship in `cas9/data/orthologs.fasta` (SpCas9, SaCas9, CjCas9,
St1/St3Cas9, GeoCas9, NmCas9, iSpyMac, SpCas9++, SpRYc). Set `ORTHOLOG` to any
of them. `--target_pam matching` optimizes toward the seed's own PAM; pass an
explicit 10-nt `ACGTN` string to retarget.
## 4. Included checkpoints
| Path | Role |
|---|---|
| `gfp/ckpt/last_2.ckpt` | GFP Edit Flow |
| `gfp/classifier_ckpt/best.pt` | GFP hard constraint |
| `gfp/FPredX/{ex,bright,em}_model` | FPredX property models |
| `peptidomimetics/ckpt/SELFIES_EditFlows.ckpt` | peptidomimetic Edit Flow |
| `peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt` | motif / specificity |
| `peptidomimetics/ckpt/admetica/*.ckpt` | Admetica blending oracles |
| `peptidomimetics/ckpt/deepdtagen/` | DeepDTAGen affinity |
| `cas9/ckpt/last.ckpt` | Cas9 Edit Flow |
| `cas9/classifier_ckpt/` | Cas9-likeness classifier (objective + hard constraint) |
| `cas9/hmm/` | HNH / RuvC and PAM-interacting HMM databases |
Together these are about 10 GB. HuggingFace uploads should use Git LFS for the `.ckpt` / `.pt` / `.pth` files.
## 5. Environment variables
| Variable | Purpose |
|---|---|
| `PEPTIVERSE_ROOT` | PeptiVerse repo root (manifest + `training_classifiers/`) |
| `MAFFT_PATH` | MAFFT binary, if it is not on `PATH` |
| `GFP_CLASSIFIER_CKPT` | optional override for the GFP classifier |
| `PROTEIN2PAM_ROOT` | protein2pam package root (PAM oracle model class) |
| `CAS_PREDICTOR_ROOT` | cas_predictor package root (Cas9 classifier code) |
| `CAS9_CLASSIFIER_CKPT` | override for the Cas9 classifier checkpoint |
| `CAS9_CLASSIFIER_CONFIG` | optional Cas9 classifier config |
| `CAS9_HMM_DIR` | directory holding the Cas9 HMM databases |
## 6. Notes
- Run scripts from the repo root, or use the wrappers under `*/scripts/`.
- GFP FPredX will fail immediately if `mafft` is missing; that is expected.
- Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded.
- Cas9 objectives fail at init if `hmmscan`, `protein2pam`, or `cas_predictor` are missing.
- Training writes to `outputs/` by default.
## Citation
If you use this code, please cite the pCoMole paper and [PeptiVerse](https://www.biorxiv.org/content/10.64898/2025.12.31.697180).