# pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows ![pCoMole_overview_10s](https://cdn-uploads.huggingface.co/production/uploads/64cd5b3f0494187a9e8b7c69/_wLLPKFMOpANDzjc85xSZ.gif) ## Repository layout ```text pCoMole/ train.py / evaluate.py # Edit Flow train / val-loss eval configs/ # GFP (protein) and SELFIES configs model/ logic/ flow_matching/ # Edit Flow implementation smiles_tokenizer/ # SMILES SPE + SELFIES vocab data/selfies/28k_mimetics/ # shipped SELFIES training split gfp/ # GFP generation + pCoMole pcomole.py # official GFP editor generate.py ckpt/last_2.ckpt classifier_ckpt/best.pt FPredX/ # excitation / brightness / emission models peptidomimetics/ # peptidomimetic generation + pCoMole pcomole.py generate.py objectives.py # PeptiVerse + Admetica/DeepDTAGen mix ckpt/SELFIES_EditFlows.ckpt ckpt/SMILES_BindEvaluator.ckpt ckpt/admetica/ # LD50, solubility, Caco-2, half-life ckpt/deepdtagen/ cas9/ # Cas9 generation + pCoMole pcomole.py # official Cas9 editor generate_cas9_csvs.py objectives.py # Cas9Classification, DeletionCount, PAMMatching constraints.py domain_detector.py # HNH / RuvC completeness (HMMER) pam_detector_hmm.py # PAM-interacting domain masking model/ logic/ # Cas9 Edit Flow variant (see note below) data/orthologs.fasta # 10 Cas9 ortholog seeds ckpt/last.ckpt classifier_ckpt/ # Cas9-likeness classifier hmm/ # HMM databases ``` > **Note.** `cas9/model/` and `cas9/logic/` are a *fork* of the repo-root > packages, not duplicates: the Cas9 stack keeps the reparameterized Edit Flow > models inside `base_models.py`, whereas the root splits them into > `reparam_models.py`. The two are not interchangeable, so the Cas9 task > imports its own copies (`from cas9.model...`). ## 1. Environment A CUDA GPU is strongly recommended. ```bash git clone pCoMole cd pCoMole python -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` HuggingFace model downloads happen on first use (`facebook/esm2_t33_650M_UR50D`, `aaronfeller/PeptideCLM-23M-all`, ChemBERTa). ### PeptiVerse (required for peptidomimetic pCoMole) Peptidomimetic property oracles use the latest [PeptiVerse](https://huggingface.co/ChatterjeeLab/PeptiVerse) SMILES predictors. Clone it **next to** this repo, or set `PEPTIVERSE_ROOT`: ```bash # recommended: sibling of pCoMole cd .. git clone https://huggingface.co/ChatterjeeLab/PeptiVerse cd PeptiVerse pip install -r requirements.txt # follow PeptiVerse README to download model weights ``` pCoMole looks for PeptiVerse in this order: 1. `$PEPTIVERSE_ROOT` 2. `../PeptiVerse` relative to this repository PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged. ### MAFFT (required for GFP pCoMole) GFP excitation, brightness, and emission oracles run through FPredX and need [MAFFT](https://mafft.cbrc.jp/alignment/software/) on your `PATH`. ```bash # conda conda install -c bioconda mafft # or point to a local binary export MAFFT_PATH=/path/to/mafft ``` Confirm with `mafft --version` before running GFP editing. ### HMMER and protein2pam (required for Cas9 pCoMole) The Cas9 domain-completeness and PAM-masking oracles shell out to `hmmscan`: ```bash conda install -c bioconda hmmer ``` The PAM objective needs the `protein2pam` model class (weights download from the Hub as `Profluent-Bio/protein2pam-cas9_full`): ```bash pip install protein2pam==0.2.0 ``` ## 2. Train, evaluate, and sample Edit Flows Run every command from the **repository root**. ### SELFIES peptidomimetic Edit Flow Training data is shipped at `data/selfies/28k_mimetics`. ```bash # train python train.py --config configs/config_selfies.yaml # optional: python train.py --config configs/config_selfies.yaml --wandb # evaluate a checkpoint on the validation split python evaluate.py \ --config configs/config_selfies.yaml \ --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt # unconditional / seeded generation python peptidomimetics/generate.py \ --config configs/config_selfies.yaml \ --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \ --input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \ --num_steps 30 \ --num_samples 2 \ --output_csv outputs/peptidomimetic_unconditional.csv ``` Or: `bash peptidomimetics/scripts/train.sh` and `bash peptidomimetics/scripts/generate.sh`. ### GFP protein Edit Flow A trained GFP checkpoint is shipped at `gfp/ckpt/last_2.ckpt`. The original GFP training arrows are **not** included. To retrain, put a HuggingFace `load_from_disk` dataset under `data/gfp/{train,validation}` (or edit `configs/config_gfp.yaml`). ```bash python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt python gfp/generate.py \ --config configs/config_gfp.yaml \ --ckpt gfp/ckpt/last_2.ckpt \ --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \ --num_steps 10 \ --num_samples 8 \ --output_csv outputs/gfp_unconditional.csv ``` ## 3. pCoMole multi-objective editing ### GFP Official editor: `gfp/pcomole.py` (length + excitation + brightness, with GFP-classifier and emission constraints). ```bash bash gfp/scripts/pcomole.sh ``` Equivalent command: ```bash python gfp/pcomole.py \ --root_dir gfp/FPredX \ --config configs/config_gfp.yaml \ --ckpt gfp/ckpt/last_2.ckpt \ --num_steps 10 \ --num_candidates 50 \ --num_rollouts 10 \ --objective_weights 3 1 1 \ --output_file outputs/gfp_length_excitation_brightness.csv \ --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' ``` If the output filename contains `length_excitation_brightness`, `length_brightness`, `length_excitation`, or `length`, the matching objective subset is used. Any other name uses all three objectives. ### Peptidomimetics Official editor: `peptidomimetics/pcomole.py`. Objectives (7 scores when `--specificity` is on): 1. non-toxicity 2. solubility 3. permeability 4. half-life 5. affinity 6. motif 7. specificity PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint. ```bash # after PeptiVerse is cloned and its weights are downloaded bash peptidomimetics/scripts/pcomole.sh ``` `--target` and `--motifs` are required for affinity and motif oracles. `--specificity` adds the seventh objective. ### Cas9 Official editor: `cas9/pcomole.py`. Three objectives — Cas9-likeness, deletion count, and PAM matching — under Cas9-score, PAM-matching and length constraints. Cas9 code is a package, so run it from the repository root with `-m`: ```bash ORTHOLOG=cjcas9 bash cas9/scripts/pcomole.sh ``` Equivalent command: ```bash python -m cas9.pcomole \ --pcomol_config configs/config_cas9.yaml \ --ckpt cas9/ckpt/last.ckpt \ --input "$(awk -v n='>cjcas9' '$1==n{f=1;next}/^>/{f=0}f{printf "%s",$0}' cas9/data/orthologs.fasta)" \ --cas9_objective --deletion_count_obj --pam_distr_objective \ --objective_weights 1 1 1 \ --target_pam matching \ --num_steps 5 --num_rollouts 5 --num_candidates 10 --num_sequences 50 \ --max_deletion_percentage 0.2 --deletion_rate_scale 2500 \ --pam_min_confidence 0.7 --cas9_score_threshold 0.5 \ --max_target_length 930 --min_target_length 890 \ --use_ce_loss --allow_equal_length --zero_lam_ins \ --output_file outputs/cas9_cjcas9_pcomole.csv ``` Ten ortholog seeds ship in `cas9/data/orthologs.fasta` (SpCas9, SaCas9, CjCas9, St1/St3Cas9, GeoCas9, NmCas9, iSpyMac, SpCas9++, SpRYc). Set `ORTHOLOG` to any of them. `--target_pam matching` optimizes toward the seed's own PAM; pass an explicit 10-nt `ACGTN` string to retarget. ## 4. Included checkpoints | Path | Role | |---|---| | `gfp/ckpt/last_2.ckpt` | GFP Edit Flow | | `gfp/classifier_ckpt/best.pt` | GFP hard constraint | | `gfp/FPredX/{ex,bright,em}_model` | FPredX property models | | `peptidomimetics/ckpt/SELFIES_EditFlows.ckpt` | peptidomimetic Edit Flow | | `peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt` | motif / specificity | | `peptidomimetics/ckpt/admetica/*.ckpt` | Admetica blending oracles | | `peptidomimetics/ckpt/deepdtagen/` | DeepDTAGen affinity | | `cas9/ckpt/last.ckpt` | Cas9 Edit Flow | | `cas9/classifier_ckpt/` | Cas9-likeness classifier (objective + hard constraint) | | `cas9/hmm/` | HNH / RuvC and PAM-interacting HMM databases | Together these are about 10 GB. HuggingFace uploads should use Git LFS for the `.ckpt` / `.pt` / `.pth` files. ## 5. Environment variables | Variable | Purpose | |---|---| | `PEPTIVERSE_ROOT` | PeptiVerse repo root (manifest + `training_classifiers/`) | | `MAFFT_PATH` | MAFFT binary, if it is not on `PATH` | | `GFP_CLASSIFIER_CKPT` | optional override for the GFP classifier | | `PROTEIN2PAM_ROOT` | protein2pam package root (PAM oracle model class) | | `CAS_PREDICTOR_ROOT` | cas_predictor package root (Cas9 classifier code) | | `CAS9_CLASSIFIER_CKPT` | override for the Cas9 classifier checkpoint | | `CAS9_CLASSIFIER_CONFIG` | optional Cas9 classifier config | | `CAS9_HMM_DIR` | directory holding the Cas9 HMM databases | ## 6. Notes - Run scripts from the repo root, or use the wrappers under `*/scripts/`. - GFP FPredX will fail immediately if `mafft` is missing; that is expected. - Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded. - Cas9 objectives fail at init if `hmmscan`, `protein2pam`, or `cas_predictor` are missing. - Training writes to `outputs/` by default. ## Citation If you use this code, please cite the pCoMole paper and [PeptiVerse](https://www.biorxiv.org/content/10.64898/2025.12.31.697180).