ESM-C 300M masked-flow generator for antimicrobial peptides (stage C)
Weights of the generator behind the AMP Challenge 2027 submission, code at github.com/eamag/esmc-flow-amp.
ESM-C 300M fine-tuned as a masked-flow (discrete unmasking) generator of 8-50 residue peptides. It is conditioned on three
control tokens (net charge, length, mean hydrophobicity), each a bucket id 0-6 or NULL (7). Input layout:
[BOS][charge][length][hydro][residues ...][EOS]. The model only generates sequences. It has no activity head, so it
does not score or rank peptides.
What it is for
- Draw a candidate pool of AMP-like peptides. This is what the submission does: sample tens of thousands of sequences, then score and select them with separate predictors (AMPredictor, TabPFN). The pool is the raw material, not the answer.
- Steer a library by net charge. Ask for a charge bucket and the samples follow it. Across 100 peptides of 18 residues each, mean charge at pH 7 was 2.04, 5.80 and 8.47 for buckets 1, 3 and 5. Cationic charge is the main lever for antimicrobial activity, so this is the control that matters most.
- Steer by length and hydrophobicity. Both are trained controls. The sampler also fixes the output length you ask for. The CLI draws lengths from the DBAASP distribution and leaves hydrophobicity NULL, so hydrophobicity is reachable only through the Python API. I have not measured how tightly either control steers.
- Unconditional or partly conditioned sampling. Set any control to NULL (7) and the model samples freely along that axis.
- Start from these weights for your own data. The repo's
esmc_flow_amp.trainresumes fromcheckpoint.ptand continues with your own peptide table. Resuming for 2 steps was tested. Full retraining was not repeated for this card. - Reproduce or compare against the challenge submission. The weights plus the repo regenerate the sampling pool. See
REPRODUCE.mdin the code repository.
Not for: predicting MIC or toxicity, sequences outside 8-50 residues, or non-canonical residues.
Usage
The checkpoint uses the state-dict layout of esm==3.2.1.post1. Newer esm releases (3.4.x) rename the keys and will not
load it. The repo lock file pins the right version, so run everything through uv from the repo root. Both weight files
load strictly into esmc_flow_amp.flow_model.ESMCFlowBackbone and give finite logits.
Quick start: 200 peptides
git clone https://github.com/eamag/esmc-flow-amp && cd esmc-flow-amp
uv run --extra pipeline --locked hf download eamag/esmc-flow-amp checkpoint.pt --local-dir weights
uv run --extra pipeline --locked python -m esmc_flow_amp.sample \
--checkpoint weights/checkpoint.pt --spec "b1:100,b5:100" --parts 1 --out pool.fasta
--spec is a list of b<charge bucket>:<number of draws>. Lengths are drawn from the DBAASP length distribution and
hydrophobicity is NULL. This run wrote 200 unique sequences in 13 s on an Apple M5 Pro (MPS). The sampler picks CUDA, then
MPS, then CPU.
The submission pool used the default spec, --parts 1 of which is 30,000 draws:
--spec "b0:900,b1:3600,b2:11100,b3:10800,b4:3000,b5:600"
Python API: choose the controls yourself
import torch
from esmc_flow_amp.controls import NULL_BUCKET
from esmc_flow_amp.flow_model import ESMCFlowBackbone, adapt_state_dict
from esmc_flow_amp.sample import residue_map, sample_sequences
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
ckpt = torch.load("weights/checkpoint.pt", map_location="cpu", weights_only=False)
backbone = ESMCFlowBackbone(conds=tuple(ckpt["conds"]), device="cpu")
backbone.model.load_state_dict(adapt_state_dict(ckpt["model"], backbone.base_vocab, backbone.n_controls))
backbone.model.to(device).eval()
# 100 peptides of 18 residues: charge bucket 5 (+8 to +11), length bucket 2 (17-20), hydrophobicity free
seqs = sample_sequences(
backbone, residue_map(backbone), n=100,
lengths=[18], ctrls_list=[(5, 2, NULL_BUCKET)], # (charge, length, hydro)
seed=0,
)
Sampling defaults: 12 confident-unmasking steps, temperature 1.0, top-p 0.95, guidance weight 1.0 (no extra forward pass).
Changing cfg_weight turns on classifier-free guidance, logits_null + w * (logits_cond - logits_null), at twice the cost.
Control buckets
A bucket id is the number of edges the value exceeds, so a value equal to an edge falls in the lower bucket. The length
control is the exception at sampling time. sampling_length_bucket puts an exact edge length one bucket higher, as it did
when the submission pool was drawn.
| Bucket | Net charge, pH 7 (Bjellqvist) | Length (residues) | Mean hydrophobicity (Eisenberg) |
|---|---|---|---|
| 0 | ≤ 0 | ≤ 12 | ≤ -0.45 |
| 1 | 0 to 2 | 13-16 | -0.45 to -0.25 |
| 2 | 2 to 4 | 17-20 | -0.25 to -0.10 |
| 3 | 4 to 6 | 21-26 | -0.10 to 0.05 |
| 4 | 6 to 8 | 27-34 | 0.05 to 0.25 |
| 5 | 8 to 11 | 35-44 | 0.25 to 0.50 |
| 6 | > 11 | > 44 | > 0.50 |
| 7 | NULL, no constraint | NULL | NULL |
Files
| File | What it is |
|---|---|
model.safetensors |
fp32 state dict, 309 tensors, 333,043,288 parameters |
checkpoint.pt |
the same tensors as {"model", "step", "val", "conds"}, the format esmc_flow_amp.sample --checkpoint loads |
config.json |
controls, base model, step and validation loss |
Optimizer state is not included. The two weight files hold identical tensors.
Training
Three stages, each resumed from the previous one. From the run's own report.json files: batch 64, AdamW with learning
rate 1e-5 (control embeddings 1e-3), weight decay 0.01, control dropout 0.15, uniform mask rate, fp16 mixed precision on a
Tesla T4.
| Stage | Steps | Table |
|---|---|---|
| A, prior | 0-3,999 | broad peptide prior, 5,634,791 rows |
| B, AMP-like | 4,000-9,999 | AMPSphere plus curated peptide databases, 889,293 rows |
| C, potent | 10,000-11,999 | 4,460 measured peptides with broad-spectrum MIC ≤ 16 µM |
Stages A and B are published as eamag/amp-plm. Stage C is rebuilt from
data/oracle/ by the code repository.
This checkpoint is global step 11,999, validation loss 1.909. Stage C was resumed from a step-10,199 snapshot after a spot-instance preemption.
Limits
- Activity is predicted (AMPredictor, TabPFN), not measured. There is no wet-lab data for anything this model generated.
- The weights are derived from ESM C 300M. The upstream licence applies to them; read it before redistributing.
- Training tables include sources with academic-use terms or no licence file. The list is in
METHOD.mdsection 7 anddata/plm/README.md. - Retraining is not bit-identical (GPU nondeterminism, spot-instance restarts).
- Downloads last month
- 1
Model tree for eamag/esmc-flow-amp
Base model
biohub/esmc-300m-2024-12