GenScore
Model Overview
GenScore is a protein-ligand scoring framework that can predict protein-small molecule binding affinity and evaluate docking pose quality. It demonstrates balanced scoring, ranking, docking, and virtual screening capabilities across multiple datasets.
Model Description
GenScore is based on graph neural networks and extends RTMScore.
Weights and datasets are not provided at the moment. They will be uploaded to ModelScope soon, and command-line downloads will be supported later.
Use Cases
| Use case | Description |
|---|---|
| Protein-ligand scoring | Predicts binding scores for a given protein, or an extracted binding pocket, and ligand conformation. |
| Automatic pocket generation | Automatically extracts the binding pocket from a full protein structure based on the reference ligand position. |
| Contribution analysis | Outputs atom-level and residue-level contributions to the final score for interpretability analysis. |
| Model training | Trains the GenScore scoring network using preprocessed PDBbind protein-ligand graph data. |
| CASF-2016 benchmark evaluation | Supports three standard tests: scoring/ranking, docking, and virtual screening. |
Usage
1. Using OneCode
You can try intelligent one-click AI4S programming through the OneCode online environment:
Try intelligent one-click AI4S programming
2. Manual Installation and Usage
Hardware Requirements
- Running on a GPU or DCU is recommended.
- CPU can be used for connectivity checks, but it is relatively slow.
- DCU users need to install DTK in advance. DTK 25.04.2 or later is recommended, or the OneScience-recommended version that matches the current cluster.
Software Requirements
DCU users who want to learn more about adaptation details can contact liubiao@sugon.com.
Environment Checks
- NVIDIA GPU:
nvidia-smi
- Hygon DCU:
hy-smi
Quick Start
1. Install the Runtime Environment
conda create -n onescience311 python=3.11 -y
conda activate onescience311
pip install onescience[bio] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
If the following code cannot find required libraries at runtime, activate CUDA as shown below.
source ${ROCM_PATH}/cuda/env.sh
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"
2. Download the Model Package
# By default, the package is downloaded to the GenScore folder under the current path. To change this, adjust the path after local_dir.
modelscope download --model OneScience/GenScore --local_dir ./model
cd model
Training Weights and Datasets
Weights and datasets are not provided at the moment. They will be uploaded to ModelScope soon, and command-line downloads will be supported later.
3. Usage Modes
1. Protein-Ligand Scoring Inference (run_genscore.sh)
cd ./scripts
bash run_genscore.sh
This script runs four inference examples in sequence, covering common use cases:
- GT model + automatic pocket generation: inputs a full protein, a reference ligand, and a decoy ligand, then automatically generates the binding pocket and scores it.
- GatedGCN model + pre-extracted pocket: inputs a pre-extracted pocket PDB and a decoy ligand for scoring.
- Atom contribution analysis: uses the GatedGCN model to compute each atom's contribution to the score.
- Residue contribution analysis: uses the GatedGCN model to compute each residue's contribution to the score.
2. Model Training
2.1 Small-Scale Smoke Test (train_genscore_smoke.sh)
cd ./scripts
bash train_genscore_smoke.sh
Default configuration:
| Parameter | Default value |
|---|---|
| Number of training epochs | 100 |
| Batch size | 16 |
| Number of validation samples | 1500 |
| Encoder | gatedgcn |
| Output model | genscore_smoke_bs16.pth |
2.2 Full Training (train_genscore_full.sh)
cd ./scripts
bash train_genscore_full.sh
Key parameters can be overridden through environment variables:
| Environment variable | Description | Default value |
|---|---|---|
GENSCORE_DATA_DIR |
Training data directory | ${ONESCIENCE_DATASETS_DIR}/GenScore/genscore_data/rtmscore_s |
GENSCORE_DATA_PREFIX |
Data file prefix | v2020_train |
GENSCORE_ENCODER |
Graph encoder type | gatedgcn |
GENSCORE_MODEL_PATH |
Output model path | examples/biosciences/genscore/genscore_${GENSCORE_ENCODER}_full_3000.pth |
GENSCORE_NUM_EPOCHS |
Number of training epochs | 3000 |
GENSCORE_BATCH_SIZE |
Batch size | 64 |
GENSCORE_NUM_WORKERS |
Number of data loading workers | 8 |
GENSCORE_VALNUM |
Number of validation samples | 1500 |
GENSCORE_PATIENCE |
Early-stopping patience | 70 |
3. CASF-2016 Benchmark Evaluation (run_genscore_benchmarks.sh)
cd ./scripts
bash run_genscore_benchmarks.sh all
You can also run a single evaluation task:
bash run_genscore_benchmarks.sh scoring
bash run_genscore_benchmarks.sh docking
bash run_genscore_benchmarks.sh screening
Evaluation task descriptions:
| Task | Description |
|---|---|
scoring |
Scoring capability evaluation (scoring/ranking) |
docking |
Docking capability evaluation (docking power) |
screening |
Virtual screening capability evaluation (screening power) |
Data Preprocessing
Training requires preprocessed PDBbind graph data, including:
<data_prefix>_ids.npy
<data_prefix>_lig.pt
<data_prefix>_prot.pt
You can preprocess the raw PDBbind data with the following command:
cd ./scripts
export PYTHONPATH=../../../src:$PYTHONPATH
python preprocess_pdbbind.py \
--dir /path/to/pdbbind \
--ref /path/to/pdbbind_2020_general.csv \
--cutoff 10.0 \
--outprefix /path/to/preprocessed/pdbbind/v2020_train
Main parameters:
| Parameter | Description |
|---|---|
--dir |
Raw PDBbind data directory |
--ref |
PDBbind index CSV file |
--cutoff |
Protein-ligand distance cutoff, default 10.0 Å |
--outprefix |
Output file prefix |
Notes
- Make sure the
ONESCIENCE_DATASETS_DIRenvironment variable is correctly set before running the scripts. - The scripts automatically set ROCm/DCU-related
LD_LIBRARY_PATHvalues and can run directly on Hygon DCU platforms. - Automatic pocket generation depends on OpenBabel and ProDy. If OpenBabel requires explicit data paths, set
BABEL_LIBDIRandBABEL_DATADIRin advance. - Training and inference use a single device by default (
HIP_VISIBLE_DEVICES=0orCUDA_VISIBLE_DEVICES=0). For multi-device training, adjust the parallel strategy in the training scripts. - The training scripts check whether
<data_prefix>_ids.npy,<data_prefix>_lig.pt, and<data_prefix>_prot.ptexist. If any are missing, the scripts will report an error and exit. - The
dockingandscreeningtasks in CASF-2016 evaluation require complete CASF-2016 and PDBbind v2020 data.
Official OneScience Information
| Platform | OneScience main repository | Skills repository |
|---|---|---|
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
Citation and License
If you use GenScore in your research, please cite the original work:
@article{genScore,
title={GenScore: a generalized protein-ligand scoring framework},
author={Shen, Chao and Hu, Yafeng and Wang, Zhe and Zhang, Xujun and Li, Jianxin and Wang, Guisheng and Wang, Tingjun and Chen, Yen-Wei and Pan, Peichen and Hou, Tingjun},
journal={Journal of Chemical Information and Modeling},
year={2023}
}
For more information, see the official GenScore repository: https://github.com/sc8668/GenScore
For license information, refer to the description in the official GenScore repository.