ESM

Model Introduction

ESM (Evolutionary Scale Modeling) is a family of protein language models released by Meta AI / FAIR. It can be used for protein representation extraction, structure prediction, variant effect scoring, and fixed-backbone sequence design.

Paper: Evolutionary-scale prediction of atomic-level protein structure with a language model
https://www.science.org/doi/10.1126/science.ade2574

Model Description

This model package provides PyTorch inference support for ESM-1, ESM-2, MSA Transformer, ESMFold, ESM-1v, and ESM-IF1, together with adaptations for running on DCUs. Sample data is distributed with the Hugging Face model repository OneScience-Sugon/ESM.

Use Cases

Scenario Description
Protein representation extraction Takes a FASTA file as input and outputs per-token, mean-pooled, BOS, or contact representations
Protein structure prediction Takes one or more amino acid sequences as input and outputs corresponding PDB structure files
Variant effect scoring Takes a wild-type sequence and a DMS mutation table as input and outputs mutation effect scores
Fixed-backbone sequence design Takes a PDB / CIF structure and chain ID as input and samples candidate sequences that satisfy the backbone constraints
Structure-conditioned sequence scoring Takes a structure and candidate sequences as input and computes their conditional log-likelihoods
Hugging Face / OneCode execution After downloading the model project, quickly verifies that the scripts run correctly in a life-sciences runtime environment

Usage Guide

1. OneCode Usage

Try one-click AI4S development in the OneCode online environment:

Try one-click AI4S development

2. Manual Installation and Usage

Hardware Requirements

  • GPU or DCU is recommended.
  • A CPU can be used for import checks and lightweight configuration tests; full training and inference will be slow.
  • DCU users must install DTK in advance. DTK 25.04.2 or later is recommended, or a OneScience-recommended version matching the current cluster.

Environment Check

  • NVIDIA GPU:
nvidia-smi
  • Hygon DCU:
hy-smi

Download the Model Package

hf download --model OneScience-Sugon/ESM --local-dir ./ESM
cd ESM

This model package includes a small set of sample data that can be used directly to validate the default workflow.

Install the Runtime Environment

DCU Environment

# Activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation supported
pip install onescience[bio-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai
# If required libraries cannot be found, activate the CUDA compatibility environment as follows:
source ${ROCM_PATH}/cuda/env.sh
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"

After installation, return to the model package directory:

cd ./ESM

Training and Inference Data Overview

The FASTA, PDB / CIF, and DMS files used in the ESM examples are distributed with the Hugging Face model repository OneScience-Sugon/ESM. After downloading the complete model package, the files are available under data/. This model package does not include a training entry point; the data is intended for example inference and workflow validation. You can also download only the data directory:

hf download --model OneScience-Sugon/ESM ESM/data --local-dir ./data

Model Weights

The repository includes multiple ESM model checkpoints under weight/; select the appropriate checkpoint for inference.

Preparing Weights

Place the required ESM weights in the following directory:

weight/
  checkpoints/
    esm2_t6_8M_UR50D.pt
    esmfold_3B_v1.pt
    esm1v_t33_650M_UR90S_1.pt
    esm_if1_gvp4_t16_142M_UR50.pt
    ...

When using a shared runtime environment, you can specify the weights location through an environment variable:

export ESM_WEIGHT_DIR=/path/to/esm/weight

The default example uses:

  • weight/checkpoints/esm2_t6_8M_UR50D.pt

The ESMFold, ESM-1v, and ESM-IF1 examples require their respective weights to be available.

Default Example

bash scripts/infer.sh

The default example reads data/fasta/few_proteins.fasta, extracts protein representations using esm2_t6_8M_UR50D.pt, and saves the results to outputs/embeddings/.

Sequence Representation Extraction

python scripts/extract.py \
  weight/checkpoints/esm2_t6_8M_UR50D.pt \
  data/fasta/few_proteins.fasta \
  outputs/embeddings \
  --include mean per_tok \
  --repr_layers 6

ESMFold Structure Prediction

python scripts/fold.py \
  -i data/fasta/few_proteins.fasta \
  -o outputs/pdb \
  --model-dir weight \
  --cpu-only

The output directory will contain one or more .pdb files. For production GPU / DCU inference, remove --cpu-only and configure --chunk-size or --max-tokens-per-batch according to the available accelerator memory.

You can also explicitly enable ESMFold via the default script:

RUN_ESMFOLD=1 bash scripts/infer.sh

Inverse Folding โ€” Sequence Sampling

python scripts/inverse_folding/sample_sequences.py \
  data/inverse_folding/5YH2.pdb \
  --chain A \
  --outpath outputs/sampled_seqs.fasta \
  --num-samples 1 \
  --nogpu

Inverse Folding โ€” Sequence Scoring

python scripts/inverse_folding/score_log_likelihoods.py \
  data/inverse_folding/5YH2.pdb \
  data/inverse_folding/5YH2_mutated_seqs.fasta \
  --chain A \
  --outpath outputs/sequence_scores.csv \
  --nogpu

Variant Effect Prediction

Variant effect prediction requires a wild-type sequence that is consistent with the mutation annotations in the DMS table:

python scripts/variant_prediction/predict.py \
  --model-location esm1v_t33_650M_UR90S_1 \
  --sequence "${ESM_VARIANT_SEQUENCE}" \
  --dms-input data/variant_prediction/BLAT_ECOLX_Ranganathan2015.csv \
  --mutation-col mutant \
  --dms-output outputs/variant_prediction.csv \
  --offset-idx 24 \
  --scoring-strategy wt-marginals

Data Format

Sample data is stored under data/ by default:

data/
  fasta/
    few_proteins.fasta
    some_proteins.fasta
  inverse_folding/
    5YH2.pdb
    5YH2.cif
    5YH2_mutated_seqs.fasta
    example.json
  variant_prediction/
    BLAT_ECOLX_Ranganathan2015.csv
    rho_pp.csv
    aggregated_rho.csv
    aggregated_rho_round3.csv

In this structure:

  • FASTA files are used for sequence representation extraction and structure prediction.
  • PDB / CIF files are used for inverse folding sampling and structure-conditioned sequence scoring.
  • Variant effect prediction CSV files must include a mutation column. The default column name is mutant, and mutations use notation such as A123B.
  • For custom DMS data, the wild-type amino acid at each mutated position in the sequence provided via --sequence must match the corresponding mutation annotation.

Verification

Static import check:

python scripts/check_import_boundaries.py

Syntax check:

python -B -c "import ast, pathlib; [ast.parse(p.read_text(encoding='utf-8'), filename=str(p)) for root in ['model', 'scripts', 'tests'] for p in pathlib.Path(root).rglob('*.py')]"

OneScience Official Information

Citation & License

  • This repository is adapted from the open-source ESM model to support DCUs.
  • The ESM source code is licensed under the MIT License; see LICENSE. For the usage terms governing model weights and data, refer to the documentation provided by the respective publishers.
  • For scientific use, please cite the corresponding original ESM paper for each submodel used. For ESM-2 / ESMFold, cite: Evolutionary-scale prediction of atomic-level protein structure with a language model.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support