File size: 10,761 Bytes
84ee494
7f316fe
 
3086e08
7f316fe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12fea4a
 
 
 
 
 
 
 
 
 
 
 
7f316fe
 
12fea4a
 
 
 
 
 
7f316fe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12fea4a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7f316fe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12fea4a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7f316fe
 
 
 
 
 
 
 
 
 
 
12fea4a
 
 
7f316fe
 
 
 
 
 
 
 
 
 
12fea4a
 
 
 
 
7f316fe
 
 
 
 
 
12fea4a
7f316fe
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
# pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows


![pCoMole_overview_10s](https://cdn-uploads.huggingface.co/production/uploads/64cd5b3f0494187a9e8b7c69/_wLLPKFMOpANDzjc85xSZ.gif)

## Repository layout

```text
pCoMole/
  train.py / evaluate.py          # Edit Flow train / val-loss eval
  configs/                        # GFP (protein) and SELFIES configs
  model/ logic/ flow_matching/    # Edit Flow implementation
  smiles_tokenizer/               # SMILES SPE + SELFIES vocab
  data/selfies/28k_mimetics/      # shipped SELFIES training split
  gfp/                            # GFP generation + pCoMole
    pcomole.py                    # official GFP editor
    generate.py
    ckpt/last_2.ckpt
    classifier_ckpt/best.pt
    FPredX/                       # excitation / brightness / emission models
  peptidomimetics/                # peptidomimetic generation + pCoMole
    pcomole.py
    generate.py
    objectives.py                 # PeptiVerse + Admetica/DeepDTAGen mix
    ckpt/SELFIES_EditFlows.ckpt
    ckpt/SMILES_BindEvaluator.ckpt
    ckpt/admetica/                # LD50, solubility, Caco-2, half-life
    ckpt/deepdtagen/
  cas9/                           # Cas9 generation + pCoMole
    pcomole.py                    # official Cas9 editor
    generate_cas9_csvs.py
    objectives.py                 # Cas9Classification, DeletionCount, PAMMatching
    constraints.py
    domain_detector.py            # HNH / RuvC completeness (HMMER)
    pam_detector_hmm.py           # PAM-interacting domain masking
    model/ logic/                 # Cas9 Edit Flow variant (see note below)
    data/orthologs.fasta          # 10 Cas9 ortholog seeds
    ckpt/last.ckpt
    classifier_ckpt/              # Cas9-likeness classifier
    hmm/                          # HMM databases
```

> **Note.** `cas9/model/` and `cas9/logic/` are a *fork* of the repo-root
> packages, not duplicates: the Cas9 stack keeps the reparameterized Edit Flow
> models inside `base_models.py`, whereas the root splits them into
> `reparam_models.py`. The two are not interchangeable, so the Cas9 task
> imports its own copies (`from cas9.model...`).

## 1. Environment

A CUDA GPU is strongly recommended.

```bash
git clone <this-repo-url> pCoMole
cd pCoMole
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

HuggingFace model downloads happen on first use (`facebook/esm2_t33_650M_UR50D`, `aaronfeller/PeptideCLM-23M-all`, ChemBERTa).

### PeptiVerse (required for peptidomimetic pCoMole)

Peptidomimetic property oracles use the latest [PeptiVerse](https://huggingface.co/ChatterjeeLab/PeptiVerse) SMILES predictors. Clone it **next to** this repo, or set `PEPTIVERSE_ROOT`:

```bash
# recommended: sibling of pCoMole
cd ..
git clone https://huggingface.co/ChatterjeeLab/PeptiVerse
cd PeptiVerse
pip install -r requirements.txt
# follow PeptiVerse README to download model weights
```

pCoMole looks for PeptiVerse in this order:

1. `$PEPTIVERSE_ROOT`
2. `../PeptiVerse` relative to this repository

PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged.

### MAFFT (required for GFP pCoMole)

GFP excitation, brightness, and emission oracles run through FPredX and need [MAFFT](https://mafft.cbrc.jp/alignment/software/) on your `PATH`.

```bash
# conda
conda install -c bioconda mafft

# or point to a local binary
export MAFFT_PATH=/path/to/mafft
```

Confirm with `mafft --version` before running GFP editing.

### HMMER and protein2pam (required for Cas9 pCoMole)

The Cas9 domain-completeness and PAM-masking oracles shell out to `hmmscan`:

```bash
conda install -c bioconda hmmer
```

The PAM objective needs the `protein2pam` model class (weights download from the
Hub as `Profluent-Bio/protein2pam-cas9_full`):

```bash
pip install protein2pam==0.2.0
```

## 2. Train, evaluate, and sample Edit Flows

Run every command from the **repository root**.

### SELFIES peptidomimetic Edit Flow

Training data is shipped at `data/selfies/28k_mimetics`.

```bash
# train
python train.py --config configs/config_selfies.yaml
# optional: python train.py --config configs/config_selfies.yaml --wandb

# evaluate a checkpoint on the validation split
python evaluate.py \
  --config configs/config_selfies.yaml \
  --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt

# unconditional / seeded generation
python peptidomimetics/generate.py \
  --config configs/config_selfies.yaml \
  --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \
  --input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \
  --num_steps 30 \
  --num_samples 2 \
  --output_csv outputs/peptidomimetic_unconditional.csv
```

Or: `bash peptidomimetics/scripts/train.sh` and `bash peptidomimetics/scripts/generate.sh`.

### GFP protein Edit Flow

A trained GFP checkpoint is shipped at `gfp/ckpt/last_2.ckpt`. The original GFP training arrows are **not** included. To retrain, put a HuggingFace `load_from_disk` dataset under `data/gfp/{train,validation}` (or edit `configs/config_gfp.yaml`).

```bash
python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt

python gfp/generate.py \
  --config configs/config_gfp.yaml \
  --ckpt gfp/ckpt/last_2.ckpt \
  --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \
  --num_steps 10 \
  --num_samples 8 \
  --output_csv outputs/gfp_unconditional.csv
```

## 3. pCoMole multi-objective editing

### GFP

Official editor: `gfp/pcomole.py` (length + excitation + brightness, with GFP-classifier and emission constraints).

```bash
bash gfp/scripts/pcomole.sh
```

Equivalent command:

```bash
python gfp/pcomole.py \
  --root_dir gfp/FPredX \
  --config configs/config_gfp.yaml \
  --ckpt gfp/ckpt/last_2.ckpt \
  --num_steps 10 \
  --num_candidates 50 \
  --num_rollouts 10 \
  --objective_weights 3 1 1 \
  --output_file outputs/gfp_length_excitation_brightness.csv \
  --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS'
```

If the output filename contains `length_excitation_brightness`, `length_brightness`, `length_excitation`, or `length`, the matching objective subset is used. Any other name uses all three objectives.

### Peptidomimetics

Official editor: `peptidomimetics/pcomole.py`.

Objectives (7 scores when `--specificity` is on):

1. non-toxicity
2. solubility
3. permeability
4. half-life
5. affinity
6. motif
7. specificity

PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint.

```bash
# after PeptiVerse is cloned and its weights are downloaded
bash peptidomimetics/scripts/pcomole.sh
```

`--target` and `--motifs` are required for affinity and motif oracles. `--specificity` adds the seventh objective.

### Cas9

Official editor: `cas9/pcomole.py`. Three objectives — Cas9-likeness, deletion
count, and PAM matching — under Cas9-score, PAM-matching and length constraints.

Cas9 code is a package, so run it from the repository root with `-m`:

```bash
ORTHOLOG=cjcas9 bash cas9/scripts/pcomole.sh
```

Equivalent command:

```bash
python -m cas9.pcomole \
  --pcomol_config configs/config_cas9.yaml \
  --ckpt cas9/ckpt/last.ckpt \
  --input "$(awk -v n='>cjcas9' '$1==n{f=1;next}/^>/{f=0}f{printf "%s",$0}' cas9/data/orthologs.fasta)" \
  --cas9_objective --deletion_count_obj --pam_distr_objective \
  --objective_weights 1 1 1 \
  --target_pam matching \
  --num_steps 5 --num_rollouts 5 --num_candidates 10 --num_sequences 50 \
  --max_deletion_percentage 0.2 --deletion_rate_scale 2500 \
  --pam_min_confidence 0.7 --cas9_score_threshold 0.5 \
  --max_target_length 930 --min_target_length 890 \
  --use_ce_loss --allow_equal_length --zero_lam_ins \
  --output_file outputs/cas9_cjcas9_pcomole.csv
```

Ten ortholog seeds ship in `cas9/data/orthologs.fasta` (SpCas9, SaCas9, CjCas9,
St1/St3Cas9, GeoCas9, NmCas9, iSpyMac, SpCas9++, SpRYc). Set `ORTHOLOG` to any
of them. `--target_pam matching` optimizes toward the seed's own PAM; pass an
explicit 10-nt `ACGTN` string to retarget.

## 4. Included checkpoints

| Path | Role |
|---|---|
| `gfp/ckpt/last_2.ckpt` | GFP Edit Flow |
| `gfp/classifier_ckpt/best.pt` | GFP hard constraint |
| `gfp/FPredX/{ex,bright,em}_model` | FPredX property models |
| `peptidomimetics/ckpt/SELFIES_EditFlows.ckpt` | peptidomimetic Edit Flow |
| `peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt` | motif / specificity |
| `peptidomimetics/ckpt/admetica/*.ckpt` | Admetica blending oracles |
| `peptidomimetics/ckpt/deepdtagen/` | DeepDTAGen affinity |
| `cas9/ckpt/last.ckpt` | Cas9 Edit Flow |
| `cas9/classifier_ckpt/` | Cas9-likeness classifier (objective + hard constraint) |
| `cas9/hmm/` | HNH / RuvC and PAM-interacting HMM databases |

Together these are about 10 GB. HuggingFace uploads should use Git LFS for the `.ckpt` / `.pt` / `.pth` files.

## 5. Environment variables

| Variable | Purpose |
|---|---|
| `PEPTIVERSE_ROOT` | PeptiVerse repo root (manifest + `training_classifiers/`) |
| `MAFFT_PATH` | MAFFT binary, if it is not on `PATH` |
| `GFP_CLASSIFIER_CKPT` | optional override for the GFP classifier |
| `PROTEIN2PAM_ROOT` | protein2pam package root (PAM oracle model class) |
| `CAS_PREDICTOR_ROOT` | cas_predictor package root (Cas9 classifier code) |
| `CAS9_CLASSIFIER_CKPT` | override for the Cas9 classifier checkpoint |
| `CAS9_CLASSIFIER_CONFIG` | optional Cas9 classifier config |
| `CAS9_HMM_DIR` | directory holding the Cas9 HMM databases |

## 6. Notes

- Run scripts from the repo root, or use the wrappers under `*/scripts/`.
- GFP FPredX will fail immediately if `mafft` is missing; that is expected.
- Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded.
- Cas9 objectives fail at init if `hmmscan`, `protein2pam`, or `cas_predictor` are missing.
- Training writes to `outputs/` by default.

## Citation

If you use this code, please cite the pCoMole paper and [PeptiVerse](https://www.biorxiv.org/content/10.64898/2025.12.31.697180).