|
Download README.md from OneScience-Group/GenScore: direct link, hf CLI and curl.
- Browser
- Download file 8.39 kB
-
https://huggingface.co/OneScience-Group/GenScore/resolve/main/README.md
- Command line
-
hf download hf://OneScience-Group/GenScore/README.md
-
curl -L -o README.md https://huggingface.co/OneScience-Group/GenScore/resolve/main/README.md
8.39 kB
| frameworks: | |
| - "" | |
| language: | |
| - en | |
| license: mit | |
| tags: | |
| - OneScience | |
| - protein-ligand scoring | |
| - graph neural network | |
| <p align="center"> | |
| <strong> | |
| <span style="font-size: 30px;">GenScore</span> | |
| </strong> | |
| </p> | |
| # Model Overview | |
| GenScore is a **protein-ligand scoring framework** that can predict protein-small molecule binding affinity and evaluate docking pose quality. It demonstrates balanced scoring, ranking, docking, and virtual screening capabilities across multiple datasets. | |
| # Model Description | |
| GenScore is based on graph neural networks and extends RTMScore. | |
| Weights and datasets are not provided at the moment. They will be uploaded to Hugging Face soon, and command-line downloads will be supported later. | |
| # Use Cases | |
| | Use case | Description | | |
| | :---: | :---: | | |
| | Protein-ligand scoring | Predicts binding scores for a given protein, or an extracted binding pocket, and ligand conformation. | | |
| | Automatic pocket generation | Automatically extracts the binding pocket from a full protein structure based on the reference ligand position. | | |
| | Contribution analysis | Outputs atom-level and residue-level contributions to the final score for interpretability analysis. | | |
| | Model training | Trains the GenScore scoring network using preprocessed PDBbind protein-ligand graph data. | | |
| | CASF-2016 benchmark evaluation | Supports three standard tests: scoring/ranking, docking, and virtual screening. | | |
| # Usage | |
| ## 1. Using OneCode | |
| You can try intelligent one-click AI4S programming through the OneCode online environment: | |
| [Try intelligent one-click AI4S programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home) | |
| ## 2. Manual Installation and Usage | |
| **Hardware Requirements** | |
| - Running on a GPU or DCU is recommended. | |
| - CPU can be used for connectivity checks, but it is relatively slow. | |
| - DCU users need to install DTK in advance. DTK 25.04.2 or later is recommended, or the OneScience-recommended version that matches the current cluster. | |
| **Software Requirements** | |
| DCU users who want to learn more about adaptation details can contact liubiao@sugon.com. | |
| **Environment Checks** | |
| - NVIDIA GPU: | |
| ```bash | |
| nvidia-smi | |
| ``` | |
| - Hygon DCU: | |
| ```bash | |
| hy-smi | |
| ``` | |
| ## Quick Start | |
| ### 1. Install the Runtime Environment | |
| ```bash | |
| conda create -n onescience311 python=3.11 -y | |
| conda activate onescience311 | |
| pip install onescience[bio] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai | |
| ``` | |
| If the following code cannot find required libraries at runtime, activate CUDA as shown below. | |
| ```bash | |
| source ${ROCM_PATH}/cuda/env.sh | |
| export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH" | |
| export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH" | |
| ``` | |
| ### 2. Download the Model Package | |
| ```bash | |
| # By default, the package is downloaded to the GenScore folder under the current path. To change this, adjust the path after local_dir. | |
| hf download OneScience-Group/GenScore --local-dir ./model | |
| cd model | |
| ``` | |
| ### Training Weights and Datasets | |
| Weights and datasets are not provided at the moment. They will be uploaded to Hugging Face soon, and command-line downloads will be supported later. | |
| ### 3. Usage Modes | |
| #### 1. Protein-Ligand Scoring Inference (`run_genscore.sh`) | |
| ```bash | |
| cd ./scripts | |
| bash run_genscore.sh | |
| ``` | |
| This script runs four inference examples in sequence, covering common use cases: | |
| 1. **GT model + automatic pocket generation**: inputs a full protein, a reference ligand, and a decoy ligand, then automatically generates the binding pocket and scores it. | |
| 2. **GatedGCN model + pre-extracted pocket**: inputs a pre-extracted pocket PDB and a decoy ligand for scoring. | |
| 3. **Atom contribution analysis**: uses the GatedGCN model to compute each atom's contribution to the score. | |
| 4. **Residue contribution analysis**: uses the GatedGCN model to compute each residue's contribution to the score. | |
| #### 2. Model Training | |
| ##### 2.1 Small-Scale Smoke Test (`train_genscore_smoke.sh`) | |
| ```bash | |
| cd ./scripts | |
| bash train_genscore_smoke.sh | |
| ``` | |
| Default configuration: | |
| | Parameter | Default value | | |
| |------|--------| | |
| | Number of training epochs | 100 | | |
| | Batch size | 16 | | |
| | Number of validation samples | 1500 | | |
| | Encoder | `gatedgcn` | | |
| | Output model | `genscore_smoke_bs16.pth` | | |
| ##### 2.2 Full Training (`train_genscore_full.sh`) | |
| ```bash | |
| cd ./scripts | |
| bash train_genscore_full.sh | |
| ``` | |
| Key parameters can be overridden through environment variables: | |
| | Environment variable | Description | Default value | | |
| |----------|------|--------| | |
| | `GENSCORE_DATA_DIR` | Training data directory | `${ONESCIENCE_DATASETS_DIR}/GenScore/genscore_data/rtmscore_s` | | |
| | `GENSCORE_DATA_PREFIX` | Data file prefix | `v2020_train` | | |
| | `GENSCORE_ENCODER` | Graph encoder type | `gatedgcn` | | |
| | `GENSCORE_MODEL_PATH` | Output model path | `examples/biosciences/genscore/genscore_${GENSCORE_ENCODER}_full_3000.pth` | | |
| | `GENSCORE_NUM_EPOCHS` | Number of training epochs | `3000` | | |
| | `GENSCORE_BATCH_SIZE` | Batch size | `64` | | |
| | `GENSCORE_NUM_WORKERS` | Number of data loading workers | `8` | | |
| | `GENSCORE_VALNUM` | Number of validation samples | `1500` | | |
| | `GENSCORE_PATIENCE` | Early-stopping patience | `70` | | |
| --- | |
| #### 3. CASF-2016 Benchmark Evaluation (`run_genscore_benchmarks.sh`) | |
| ```bash | |
| cd ./scripts | |
| bash run_genscore_benchmarks.sh all | |
| ``` | |
| You can also run a single evaluation task: | |
| ```bash | |
| bash run_genscore_benchmarks.sh scoring | |
| bash run_genscore_benchmarks.sh docking | |
| bash run_genscore_benchmarks.sh screening | |
| ``` | |
| Evaluation task descriptions: | |
| | Task | Description | | |
| |------|------| | |
| | `scoring` | Scoring capability evaluation (scoring/ranking) | | |
| | `docking` | Docking capability evaluation (docking power) | | |
| | `screening` | Virtual screening capability evaluation (screening power) | | |
| ### Data Preprocessing | |
| Training requires preprocessed PDBbind graph data, including: | |
| ```text | |
| <data_prefix>_ids.npy | |
| <data_prefix>_lig.pt | |
| <data_prefix>_prot.pt | |
| ``` | |
| You can preprocess the raw PDBbind data with the following command: | |
| ```bash | |
| cd ./scripts | |
| export PYTHONPATH=../../../src:$PYTHONPATH | |
| python preprocess_pdbbind.py \ | |
| --dir /path/to/pdbbind \ | |
| --ref /path/to/pdbbind_2020_general.csv \ | |
| --cutoff 10.0 \ | |
| --outprefix /path/to/preprocessed/pdbbind/v2020_train | |
| ``` | |
| Main parameters: | |
| | Parameter | Description | | |
| |------|------| | |
| | `--dir` | Raw PDBbind data directory | | |
| | `--ref` | PDBbind index CSV file | | |
| | `--cutoff` | Protein-ligand distance cutoff, default `10.0` Å | | |
| | `--outprefix` | Output file prefix | | |
| ### Notes | |
| - Make sure the `ONESCIENCE_DATASETS_DIR` environment variable is correctly set before running the scripts. | |
| - The scripts automatically set ROCm/DCU-related `LD_LIBRARY_PATH` values and can run directly on Hygon DCU platforms. | |
| - Automatic pocket generation depends on OpenBabel and ProDy. If OpenBabel requires explicit data paths, set `BABEL_LIBDIR` and `BABEL_DATADIR` in advance. | |
| - Training and inference use a single device by default (`HIP_VISIBLE_DEVICES=0` or `CUDA_VISIBLE_DEVICES=0`). For multi-device training, adjust the parallel strategy in the training scripts. | |
| - The training scripts check whether `<data_prefix>_ids.npy`, `<data_prefix>_lig.pt`, and `<data_prefix>_prot.pt` exist. If any are missing, the scripts will report an error and exit. | |
| - The `docking` and `screening` tasks in CASF-2016 evaluation require complete CASF-2016 and PDBbind v2020 data. | |
| # Official OneScience Information | |
| | Platform | OneScience main repository | Skills repository | | |
| | --- | --- | --- | | |
| | Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills | | |
| | GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills | | |
| # Citation and License | |
| If you use GenScore in your research, please cite the original work: | |
| ```bibtex | |
| @article{genScore, | |
| title={GenScore: a generalized protein-ligand scoring framework}, | |
| author={Shen, Chao and Hu, Yafeng and Wang, Zhe and Zhang, Xujun and Li, Jianxin and Wang, Guisheng and Wang, Tingjun and Chen, Yen-Wei and Pan, Peichen and Hou, Tingjun}, | |
| journal={Journal of Chemical Information and Modeling}, | |
| year={2023} | |
| } | |
| ``` | |
| For more information, see the official GenScore repository: https://github.com/sc8668/GenScore | |
| For license information, refer to the description in the [official GenScore repository](https://github.com/sc8668/GenScore). | |