GenScore

Model Overview

GenScore is a protein-ligand scoring framework that can predict protein-small molecule binding affinity and evaluate docking pose quality. It demonstrates balanced scoring, ranking, docking, and virtual screening capabilities across multiple datasets.

Model Description

GenScore is based on graph neural networks and extends RTMScore.

Weights and datasets are not provided at the moment. They will be uploaded to ModelScope soon, and command-line downloads will be supported later.

Use Cases

Use case Description
Protein-ligand scoring Predicts binding scores for a given protein, or an extracted binding pocket, and ligand conformation.
Automatic pocket generation Automatically extracts the binding pocket from a full protein structure based on the reference ligand position.
Contribution analysis Outputs atom-level and residue-level contributions to the final score for interpretability analysis.
Model training Trains the GenScore scoring network using preprocessed PDBbind protein-ligand graph data.
CASF-2016 benchmark evaluation Supports three standard tests: scoring/ranking, docking, and virtual screening.

Usage

1. Using OneCode

You can try intelligent one-click AI4S programming through the OneCode online environment:

Try intelligent one-click AI4S programming

2. Manual Installation and Usage

Hardware Requirements

  • Running on a GPU or DCU is recommended.
  • CPU can be used for connectivity checks, but it is relatively slow.
  • DCU users need to install DTK in advance. DTK 25.04.2 or later is recommended, or the OneScience-recommended version that matches the current cluster.

Software Requirements

DCU users who want to learn more about adaptation details can contact liubiao@sugon.com.

Environment Checks

  • NVIDIA GPU:
nvidia-smi
  • Hygon DCU:
hy-smi

Quick Start

1. Install the Runtime Environment

conda create -n onescience311 python=3.11 -y
conda activate onescience311
pip install onescience[bio] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai

If the following code cannot find required libraries at runtime, activate CUDA as shown below.

source ${ROCM_PATH}/cuda/env.sh
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"

2. Download the Model Package

# By default, the package is downloaded to the GenScore folder under the current path. To change this, adjust the path after local_dir.
modelscope download --model OneScience/GenScore --local_dir ./model

cd model

Training Weights and Datasets

Weights and datasets are not provided at the moment. They will be uploaded to ModelScope soon, and command-line downloads will be supported later.

3. Usage Modes

1. Protein-Ligand Scoring Inference (run_genscore.sh)

cd ./scripts
bash run_genscore.sh

This script runs four inference examples in sequence, covering common use cases:

  1. GT model + automatic pocket generation: inputs a full protein, a reference ligand, and a decoy ligand, then automatically generates the binding pocket and scores it.
  2. GatedGCN model + pre-extracted pocket: inputs a pre-extracted pocket PDB and a decoy ligand for scoring.
  3. Atom contribution analysis: uses the GatedGCN model to compute each atom's contribution to the score.
  4. Residue contribution analysis: uses the GatedGCN model to compute each residue's contribution to the score.

2. Model Training

2.1 Small-Scale Smoke Test (train_genscore_smoke.sh)
cd ./scripts
bash train_genscore_smoke.sh

Default configuration:

Parameter Default value
Number of training epochs 100
Batch size 16
Number of validation samples 1500
Encoder gatedgcn
Output model genscore_smoke_bs16.pth
2.2 Full Training (train_genscore_full.sh)
cd ./scripts
bash train_genscore_full.sh

Key parameters can be overridden through environment variables:

Environment variable Description Default value
GENSCORE_DATA_DIR Training data directory ${ONESCIENCE_DATASETS_DIR}/GenScore/genscore_data/rtmscore_s
GENSCORE_DATA_PREFIX Data file prefix v2020_train
GENSCORE_ENCODER Graph encoder type gatedgcn
GENSCORE_MODEL_PATH Output model path examples/biosciences/genscore/genscore_${GENSCORE_ENCODER}_full_3000.pth
GENSCORE_NUM_EPOCHS Number of training epochs 3000
GENSCORE_BATCH_SIZE Batch size 64
GENSCORE_NUM_WORKERS Number of data loading workers 8
GENSCORE_VALNUM Number of validation samples 1500
GENSCORE_PATIENCE Early-stopping patience 70

3. CASF-2016 Benchmark Evaluation (run_genscore_benchmarks.sh)

cd ./scripts
bash run_genscore_benchmarks.sh all

You can also run a single evaluation task:

bash run_genscore_benchmarks.sh scoring
bash run_genscore_benchmarks.sh docking
bash run_genscore_benchmarks.sh screening

Evaluation task descriptions:

Task Description
scoring Scoring capability evaluation (scoring/ranking)
docking Docking capability evaluation (docking power)
screening Virtual screening capability evaluation (screening power)

Data Preprocessing

Training requires preprocessed PDBbind graph data, including:

<data_prefix>_ids.npy
<data_prefix>_lig.pt
<data_prefix>_prot.pt

You can preprocess the raw PDBbind data with the following command:

cd ./scripts
export PYTHONPATH=../../../src:$PYTHONPATH
python preprocess_pdbbind.py \
  --dir /path/to/pdbbind \
  --ref /path/to/pdbbind_2020_general.csv \
  --cutoff 10.0 \
  --outprefix /path/to/preprocessed/pdbbind/v2020_train

Main parameters:

Parameter Description
--dir Raw PDBbind data directory
--ref PDBbind index CSV file
--cutoff Protein-ligand distance cutoff, default 10.0 Å
--outprefix Output file prefix

Notes

  • Make sure the ONESCIENCE_DATASETS_DIR environment variable is correctly set before running the scripts.
  • The scripts automatically set ROCm/DCU-related LD_LIBRARY_PATH values and can run directly on Hygon DCU platforms.
  • Automatic pocket generation depends on OpenBabel and ProDy. If OpenBabel requires explicit data paths, set BABEL_LIBDIR and BABEL_DATADIR in advance.
  • Training and inference use a single device by default (HIP_VISIBLE_DEVICES=0 or CUDA_VISIBLE_DEVICES=0). For multi-device training, adjust the parallel strategy in the training scripts.
  • The training scripts check whether <data_prefix>_ids.npy, <data_prefix>_lig.pt, and <data_prefix>_prot.pt exist. If any are missing, the scripts will report an error and exit.
  • The docking and screening tasks in CASF-2016 evaluation require complete CASF-2016 and PDBbind v2020 data.

Official OneScience Information

Citation and License

If you use GenScore in your research, please cite the original work:

@article{genScore,
  title={GenScore: a generalized protein-ligand scoring framework},
  author={Shen, Chao and Hu, Yafeng and Wang, Zhe and Zhang, Xujun and Li, Jianxin and Wang, Guisheng and Wang, Tingjun and Chen, Yen-Wei and Pan, Peichen and Hou, Tingjun},
  journal={Journal of Chemical Information and Modeling},
  year={2023}
}

For more information, see the official GenScore repository: https://github.com/sc8668/GenScore

For license information, refer to the description in the official GenScore repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support