File size: 10,761 Bytes
84ee494 7f316fe 3086e08 7f316fe 12fea4a 7f316fe 12fea4a 7f316fe 12fea4a 7f316fe 12fea4a 7f316fe 12fea4a 7f316fe 12fea4a 7f316fe 12fea4a 7f316fe | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 | # pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows

## Repository layout
```text
pCoMole/
train.py / evaluate.py # Edit Flow train / val-loss eval
configs/ # GFP (protein) and SELFIES configs
model/ logic/ flow_matching/ # Edit Flow implementation
smiles_tokenizer/ # SMILES SPE + SELFIES vocab
data/selfies/28k_mimetics/ # shipped SELFIES training split
gfp/ # GFP generation + pCoMole
pcomole.py # official GFP editor
generate.py
ckpt/last_2.ckpt
classifier_ckpt/best.pt
FPredX/ # excitation / brightness / emission models
peptidomimetics/ # peptidomimetic generation + pCoMole
pcomole.py
generate.py
objectives.py # PeptiVerse + Admetica/DeepDTAGen mix
ckpt/SELFIES_EditFlows.ckpt
ckpt/SMILES_BindEvaluator.ckpt
ckpt/admetica/ # LD50, solubility, Caco-2, half-life
ckpt/deepdtagen/
cas9/ # Cas9 generation + pCoMole
pcomole.py # official Cas9 editor
generate_cas9_csvs.py
objectives.py # Cas9Classification, DeletionCount, PAMMatching
constraints.py
domain_detector.py # HNH / RuvC completeness (HMMER)
pam_detector_hmm.py # PAM-interacting domain masking
model/ logic/ # Cas9 Edit Flow variant (see note below)
data/orthologs.fasta # 10 Cas9 ortholog seeds
ckpt/last.ckpt
classifier_ckpt/ # Cas9-likeness classifier
hmm/ # HMM databases
```
> **Note.** `cas9/model/` and `cas9/logic/` are a *fork* of the repo-root
> packages, not duplicates: the Cas9 stack keeps the reparameterized Edit Flow
> models inside `base_models.py`, whereas the root splits them into
> `reparam_models.py`. The two are not interchangeable, so the Cas9 task
> imports its own copies (`from cas9.model...`).
## 1. Environment
A CUDA GPU is strongly recommended.
```bash
git clone <this-repo-url> pCoMole
cd pCoMole
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```
HuggingFace model downloads happen on first use (`facebook/esm2_t33_650M_UR50D`, `aaronfeller/PeptideCLM-23M-all`, ChemBERTa).
### PeptiVerse (required for peptidomimetic pCoMole)
Peptidomimetic property oracles use the latest [PeptiVerse](https://huggingface.co/ChatterjeeLab/PeptiVerse) SMILES predictors. Clone it **next to** this repo, or set `PEPTIVERSE_ROOT`:
```bash
# recommended: sibling of pCoMole
cd ..
git clone https://huggingface.co/ChatterjeeLab/PeptiVerse
cd PeptiVerse
pip install -r requirements.txt
# follow PeptiVerse README to download model weights
```
pCoMole looks for PeptiVerse in this order:
1. `$PEPTIVERSE_ROOT`
2. `../PeptiVerse` relative to this repository
PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged.
### MAFFT (required for GFP pCoMole)
GFP excitation, brightness, and emission oracles run through FPredX and need [MAFFT](https://mafft.cbrc.jp/alignment/software/) on your `PATH`.
```bash
# conda
conda install -c bioconda mafft
# or point to a local binary
export MAFFT_PATH=/path/to/mafft
```
Confirm with `mafft --version` before running GFP editing.
### HMMER and protein2pam (required for Cas9 pCoMole)
The Cas9 domain-completeness and PAM-masking oracles shell out to `hmmscan`:
```bash
conda install -c bioconda hmmer
```
The PAM objective needs the `protein2pam` model class (weights download from the
Hub as `Profluent-Bio/protein2pam-cas9_full`):
```bash
pip install protein2pam==0.2.0
```
## 2. Train, evaluate, and sample Edit Flows
Run every command from the **repository root**.
### SELFIES peptidomimetic Edit Flow
Training data is shipped at `data/selfies/28k_mimetics`.
```bash
# train
python train.py --config configs/config_selfies.yaml
# optional: python train.py --config configs/config_selfies.yaml --wandb
# evaluate a checkpoint on the validation split
python evaluate.py \
--config configs/config_selfies.yaml \
--ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt
# unconditional / seeded generation
python peptidomimetics/generate.py \
--config configs/config_selfies.yaml \
--ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \
--input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \
--num_steps 30 \
--num_samples 2 \
--output_csv outputs/peptidomimetic_unconditional.csv
```
Or: `bash peptidomimetics/scripts/train.sh` and `bash peptidomimetics/scripts/generate.sh`.
### GFP protein Edit Flow
A trained GFP checkpoint is shipped at `gfp/ckpt/last_2.ckpt`. The original GFP training arrows are **not** included. To retrain, put a HuggingFace `load_from_disk` dataset under `data/gfp/{train,validation}` (or edit `configs/config_gfp.yaml`).
```bash
python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt
python gfp/generate.py \
--config configs/config_gfp.yaml \
--ckpt gfp/ckpt/last_2.ckpt \
--input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \
--num_steps 10 \
--num_samples 8 \
--output_csv outputs/gfp_unconditional.csv
```
## 3. pCoMole multi-objective editing
### GFP
Official editor: `gfp/pcomole.py` (length + excitation + brightness, with GFP-classifier and emission constraints).
```bash
bash gfp/scripts/pcomole.sh
```
Equivalent command:
```bash
python gfp/pcomole.py \
--root_dir gfp/FPredX \
--config configs/config_gfp.yaml \
--ckpt gfp/ckpt/last_2.ckpt \
--num_steps 10 \
--num_candidates 50 \
--num_rollouts 10 \
--objective_weights 3 1 1 \
--output_file outputs/gfp_length_excitation_brightness.csv \
--input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS'
```
If the output filename contains `length_excitation_brightness`, `length_brightness`, `length_excitation`, or `length`, the matching objective subset is used. Any other name uses all three objectives.
### Peptidomimetics
Official editor: `peptidomimetics/pcomole.py`.
Objectives (7 scores when `--specificity` is on):
1. non-toxicity
2. solubility
3. permeability
4. half-life
5. affinity
6. motif
7. specificity
PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint.
```bash
# after PeptiVerse is cloned and its weights are downloaded
bash peptidomimetics/scripts/pcomole.sh
```
`--target` and `--motifs` are required for affinity and motif oracles. `--specificity` adds the seventh objective.
### Cas9
Official editor: `cas9/pcomole.py`. Three objectives — Cas9-likeness, deletion
count, and PAM matching — under Cas9-score, PAM-matching and length constraints.
Cas9 code is a package, so run it from the repository root with `-m`:
```bash
ORTHOLOG=cjcas9 bash cas9/scripts/pcomole.sh
```
Equivalent command:
```bash
python -m cas9.pcomole \
--pcomol_config configs/config_cas9.yaml \
--ckpt cas9/ckpt/last.ckpt \
--input "$(awk -v n='>cjcas9' '$1==n{f=1;next}/^>/{f=0}f{printf "%s",$0}' cas9/data/orthologs.fasta)" \
--cas9_objective --deletion_count_obj --pam_distr_objective \
--objective_weights 1 1 1 \
--target_pam matching \
--num_steps 5 --num_rollouts 5 --num_candidates 10 --num_sequences 50 \
--max_deletion_percentage 0.2 --deletion_rate_scale 2500 \
--pam_min_confidence 0.7 --cas9_score_threshold 0.5 \
--max_target_length 930 --min_target_length 890 \
--use_ce_loss --allow_equal_length --zero_lam_ins \
--output_file outputs/cas9_cjcas9_pcomole.csv
```
Ten ortholog seeds ship in `cas9/data/orthologs.fasta` (SpCas9, SaCas9, CjCas9,
St1/St3Cas9, GeoCas9, NmCas9, iSpyMac, SpCas9++, SpRYc). Set `ORTHOLOG` to any
of them. `--target_pam matching` optimizes toward the seed's own PAM; pass an
explicit 10-nt `ACGTN` string to retarget.
## 4. Included checkpoints
| Path | Role |
|---|---|
| `gfp/ckpt/last_2.ckpt` | GFP Edit Flow |
| `gfp/classifier_ckpt/best.pt` | GFP hard constraint |
| `gfp/FPredX/{ex,bright,em}_model` | FPredX property models |
| `peptidomimetics/ckpt/SELFIES_EditFlows.ckpt` | peptidomimetic Edit Flow |
| `peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt` | motif / specificity |
| `peptidomimetics/ckpt/admetica/*.ckpt` | Admetica blending oracles |
| `peptidomimetics/ckpt/deepdtagen/` | DeepDTAGen affinity |
| `cas9/ckpt/last.ckpt` | Cas9 Edit Flow |
| `cas9/classifier_ckpt/` | Cas9-likeness classifier (objective + hard constraint) |
| `cas9/hmm/` | HNH / RuvC and PAM-interacting HMM databases |
Together these are about 10 GB. HuggingFace uploads should use Git LFS for the `.ckpt` / `.pt` / `.pth` files.
## 5. Environment variables
| Variable | Purpose |
|---|---|
| `PEPTIVERSE_ROOT` | PeptiVerse repo root (manifest + `training_classifiers/`) |
| `MAFFT_PATH` | MAFFT binary, if it is not on `PATH` |
| `GFP_CLASSIFIER_CKPT` | optional override for the GFP classifier |
| `PROTEIN2PAM_ROOT` | protein2pam package root (PAM oracle model class) |
| `CAS_PREDICTOR_ROOT` | cas_predictor package root (Cas9 classifier code) |
| `CAS9_CLASSIFIER_CKPT` | override for the Cas9 classifier checkpoint |
| `CAS9_CLASSIFIER_CONFIG` | optional Cas9 classifier config |
| `CAS9_HMM_DIR` | directory holding the Cas9 HMM databases |
## 6. Notes
- Run scripts from the repo root, or use the wrappers under `*/scripts/`.
- GFP FPredX will fail immediately if `mafft` is missing; that is expected.
- Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded.
- Cas9 objectives fail at init if `hmmscan`, `protein2pam`, or `cas_predictor` are missing.
- Training writes to `outputs/` by default.
## Citation
If you use this code, please cite the pCoMole paper and [PeptiVerse](https://www.biorxiv.org/content/10.64898/2025.12.31.697180).
|