Official repository for the ProteinGym benchmarks
See the codeProteinGym is an extensive set of Deep Mutational Scanning (DMS) assays and annotated human clinical variants curated to enable thorough comparisons of various mutation effect predictors in different regimes. Both the DMS assays and clinical variants are divided into 1) a substitution benchmark which currently consists of the experimental characterisation of ~2.7M missense variants across 217 DMS assays and 2,525 clinical proteins, and 2) an indel benchmark that includes ∼300k mutants across 74 DMS assays and 1,555 clinical proteins.
Each processed file in each benchmark corresponds to a single DMS assay or clinical protein, and contains the following variables:
Additionally, we provide two reference files for each benchmark that give further details on each assay and contain in particular:
To download the benchmarks, please see DMS benchmark - Substitutions and DMS benchmark - Indels in the "Resources" section below.
The benchmarks folder provides detailed performance files for all baselines on the DMS and clinical benchmarks.
We report the following metrics:
Metrics are aggregated as follows:
These files are named e.g. DMS_substitutions_Spearman_DMS_level.csv, DMS_substitutions_Spearman_Uniprot_level and DMS_substitutions_Spearman_Uniprot_Selection_Type_level respectively for these different steps.
For other deep dives (performance split by taxa, MSA depth, mutational depth and more), these are all contained in the benchmarks/DMS_zero_shot/substitutions/Spearman/Summary_performance_DMS_substitutions_Spearman.csv folder (resp. DMS_indels/clinical_substitutions/clinical_indels & their supervised counterparts). These files are also what are hosted on the website.
We also include, as on the website, a bootstrapped standard error of these aggregated metrics to reflect the variance in the final numbers with respect to the individual assays.
To calculate the DMS substitution benchmark metrics:
./scripts/scoring_DMS_zero_shot/performance_substitutions.shAnd for indels, follow step #1 and run ./scripts/scoring_DMS_zero_shot/performance_substitutions_indels.sh.
The full ProteinGym benchmarks performance files are also accessible via our dedicated website: https://www.proteingym.org/. It includes leaderboards for the substitution and indel benchmarks, as well as detailed DMS-level performance files for all baselines. The current version of the substitution benchmark includes the following baselines:
For clinical baselines, we used dbNSFP 4.4a as detailed in the manuscript appendix (and in proteingym/clinical_benchmark_notebooks/clinical_subs_processing.ipynb).
To download and unzip the data, use the following template, replacing {VERSION} with the desired version number (e.g., "v1.3") and {FILENAME} with the specific file you want to download, as listed in the table below. The latest version is v1.3. For example, you can download & unzip the zero-shot predictions for all baselines for all DMS substitution assays as follows:
VERSION="v1.3"
FILENAME="DMS_ProteinGym_substitutions.zip"
curl -o ${FILENAME} https://marks.hms.harvard.edu/proteingym/ProteinGym_${VERSION}/${FILENAME}
unzip ${FILENAME} && rm ${FILENAME}
| Data | Size (unzipped) | Filename |
|---|---|---|
| DMS benchmark - Substitutions | 1.0GB | DMS_ProteinGym_substitutions.zip |
| DMS benchmark - Indels | 200MB | DMS_ProteinGym_indels.zip |
| Zero-shot DMS Model scores - Substitutions | 4.4GB | zero_shot_substitutions_scores.zip |
| Zero-shot DMS Model scores - Indels | 313MB | zero_shot_indels_scores.zip |
| Supervised DMS Model scores - Substitutions | 3.3GB | DMS_supervised_substitutions_scores.zip |
| Supervised DMS Model scores - Indels | 215MB | DMS_supervised_indels_scores.zip |
| Multiple Sequence Alignments (MSAs) for DMS assays | 5.2GB | DMS_msa_files.zip |
| Redundancy-based sequence weights for DMS assays | 200MB | DMS_msa_weights.zip |
| Predicted 3D structures from inverse-folding models | 84MB | ProteinGym_AF2_structures.zip |
| Clinical benchmark - Substitutions | 123MB | clinical_ProteinGym_substitutions.zip |
| Clinical benchmark - Indels | 2.8MB | clinical_ProteinGym_indels.zip |
| Clinical MSAs | 17.8GB | clinical_msa_files.zip |
| Clinical MSA weights | 250MB | clinical_msa_weights.zip |
| Clinical Model scores - Substitutions | 0.9GB | zero_shot_clinical_substitutions_scores.zip |
| Clinical Model scores - Indels | 0.7GB | zero_shot_clinical_indels_scores.zip |
| CV folds - Substitutions - Singles | 50M | cv_folds_singles_substitutions.zip |
| CV folds - Substitutions - Multiples | 81M | cv_folds_multiples_substitutions.zip |
| CV folds - Indels | 19MB | cv_folds_indels.zip |
Then we also host the raw DMS assays (before preprocessing)
| Data | Size (unzipped) | Link |
|---|---|---|
| DMS benchmark: Substitutions (raw) | 500MB | substitutions_raw_DMS.zip |
| DMS benchmark: Indels (raw) | 450MB | indels_raw_DMS.zip |
| Clinical benchmark: Substitutions (raw) | 58MB | substitutions_raw_clinical.zip |
| Clinical benchmark: Indels (raw) | 12.4MB | indels_raw_clinical.zip |
If you would like to suggest new assays to be part of ProteinGym, please raise an issue on this repository with a `new_assay' label. The criteria we typically consider for inclusion are as follows:
If you would like new baselines to be included in ProteinGym (ie., website, performance files, detailed scoring files), please follow the following steps:
At this point we are only considering new baselines satisfying the following conditions:
At this stage, we are only considering requests for which all model scores for all mutants in a given benchmark (substitution or indel) are provided by the requester; but we are planning on regularly scoring new baselines ourselves for methods with wide adoption by the community and/or suggestions with many upvotes.
12 December 2023: The code for training and evaluating supervised models is currently shared in https://github.com/OATML-Markslab/ProteinNPT. We are in the process of integrating the code into this repo.
If you would like to compute all performance metrics for the various benchmarks, please follow the following steps:
Our codebase leveraged code from the following repositories to compute baselines:
We would like to thank the GEMME team for providing model scores on an earlier version of the benchmark (ProteinGym v0.1), and the ProtSSN, SaProt, PoET, MULAN, VespaG, ProSST, ESCOTT, VenusREM, RSALOR, SiteRM and AIDO teams for integrating their model in the ProteinGym repo.
Special thanks the teams of experimentalists who developed and performed the assays that ProteinGym is built on. If you are using ProteinGym in your work, please consider citing the corresponding papers. To facilitate this, we have prepared a file (assays.bib) containing the bibtex entries for all these papers.
This project is available under the MIT license found in the LICENSE file in this GitHub repository.
If you use ProteinGym in your work, please cite the following paper:
@inproceedings{NEURIPS2023_cac723e5,
author = {Notin, Pascal and Kollasch, Aaron and Ritter, Daniel and van Niekerk, Lood and Paul, Steffanie and Spinner, Han and Rollins, Nathan and Shaw, Ada and Orenbuch, Rose and Weitzman, Ruben and Frazer, Jonathan and Dias, Mafalda and Franceschi, Dinko and Gal, Yarin and Marks, Debora},
booktitle = {Advances in Neural Information Processing Systems},
editor = {A. Oh and T. Neumann and A. Globerson and K. Saenko and M. Hardt and S. Levine},
pages = {64331--64379},
publisher = {Curran Associates, Inc.},
title = {ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design},
url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/cac723e5ff29f65e3fcbb0739ae91bee-Paper-Datasets_and_Benchmarks.pdf},
volume = {36},
year = {2023}
}
76 followers · starred Mar 2024
57 followers · starred Jun 2023
HTML
50.7%
Python
43.5%
TeX
4.2%
Shell
1.4%
Official repository for the ProteinGym benchmarks
See the codeProteinGym is an extensive set of Deep Mutational Scanning (DMS) assays and annotated human clinical variants curated to enable thorough comparisons of various mutation effect predictors in different regimes. Both the DMS assays and clinical variants are divided into 1) a substitution benchmark which currently consists of the experimental characterisation of ~2.7M missense variants across 217 DMS assays and 2,525 clinical proteins, and 2) an indel benchmark that includes ∼300k mutants across 74 DMS assays and 1,555 clinical proteins.
Each processed file in each benchmark corresponds to a single DMS assay or clinical protein, and contains the following variables:
Additionally, we provide two reference files for each benchmark that give further details on each assay and contain in particular:
To download the benchmarks, please see DMS benchmark - Substitutions and DMS benchmark - Indels in the "Resources" section below.
The benchmarks folder provides detailed performance files for all baselines on the DMS and clinical benchmarks.
We report the following metrics:
Metrics are aggregated as follows:
These files are named e.g. DMS_substitutions_Spearman_DMS_level.csv, DMS_substitutions_Spearman_Uniprot_level and DMS_substitutions_Spearman_Uniprot_Selection_Type_level respectively for these different steps.
For other deep dives (performance split by taxa, MSA depth, mutational depth and more), these are all contained in the benchmarks/DMS_zero_shot/substitutions/Spearman/Summary_performance_DMS_substitutions_Spearman.csv folder (resp. DMS_indels/clinical_substitutions/clinical_indels & their supervised counterparts). These files are also what are hosted on the website.
We also include, as on the website, a bootstrapped standard error of these aggregated metrics to reflect the variance in the final numbers with respect to the individual assays.
To calculate the DMS substitution benchmark metrics:
./scripts/scoring_DMS_zero_shot/performance_substitutions.shAnd for indels, follow step #1 and run ./scripts/scoring_DMS_zero_shot/performance_substitutions_indels.sh.
The full ProteinGym benchmarks performance files are also accessible via our dedicated website: https://www.proteingym.org/. It includes leaderboards for the substitution and indel benchmarks, as well as detailed DMS-level performance files for all baselines. The current version of the substitution benchmark includes the following baselines:
For clinical baselines, we used dbNSFP 4.4a as detailed in the manuscript appendix (and in proteingym/clinical_benchmark_notebooks/clinical_subs_processing.ipynb).
To download and unzip the data, use the following template, replacing {VERSION} with the desired version number (e.g., "v1.3") and {FILENAME} with the specific file you want to download, as listed in the table below. The latest version is v1.3. For example, you can download & unzip the zero-shot predictions for all baselines for all DMS substitution assays as follows:
VERSION="v1.3"
FILENAME="DMS_ProteinGym_substitutions.zip"
curl -o ${FILENAME} https://marks.hms.harvard.edu/proteingym/ProteinGym_${VERSION}/${FILENAME}
unzip ${FILENAME} && rm ${FILENAME}
| Data | Size (unzipped) | Filename |
|---|---|---|
| DMS benchmark - Substitutions | 1.0GB | DMS_ProteinGym_substitutions.zip |
| DMS benchmark - Indels | 200MB | DMS_ProteinGym_indels.zip |
| Zero-shot DMS Model scores - Substitutions | 4.4GB | zero_shot_substitutions_scores.zip |
| Zero-shot DMS Model scores - Indels | 313MB | zero_shot_indels_scores.zip |
| Supervised DMS Model scores - Substitutions | 3.3GB | DMS_supervised_substitutions_scores.zip |
| Supervised DMS Model scores - Indels | 215MB | DMS_supervised_indels_scores.zip |
| Multiple Sequence Alignments (MSAs) for DMS assays | 5.2GB | DMS_msa_files.zip |
| Redundancy-based sequence weights for DMS assays | 200MB | DMS_msa_weights.zip |
| Predicted 3D structures from inverse-folding models | 84MB | ProteinGym_AF2_structures.zip |
| Clinical benchmark - Substitutions | 123MB | clinical_ProteinGym_substitutions.zip |
| Clinical benchmark - Indels | 2.8MB | clinical_ProteinGym_indels.zip |
| Clinical MSAs | 17.8GB | clinical_msa_files.zip |
| Clinical MSA weights | 250MB | clinical_msa_weights.zip |
| Clinical Model scores - Substitutions | 0.9GB | zero_shot_clinical_substitutions_scores.zip |
| Clinical Model scores - Indels | 0.7GB | zero_shot_clinical_indels_scores.zip |
| CV folds - Substitutions - Singles | 50M | cv_folds_singles_substitutions.zip |
| CV folds - Substitutions - Multiples | 81M | cv_folds_multiples_substitutions.zip |
| CV folds - Indels | 19MB | cv_folds_indels.zip |
Then we also host the raw DMS assays (before preprocessing)
| Data | Size (unzipped) | Link |
|---|---|---|
| DMS benchmark: Substitutions (raw) | 500MB | substitutions_raw_DMS.zip |
| DMS benchmark: Indels (raw) | 450MB | indels_raw_DMS.zip |
| Clinical benchmark: Substitutions (raw) | 58MB | substitutions_raw_clinical.zip |
| Clinical benchmark: Indels (raw) | 12.4MB | indels_raw_clinical.zip |
If you would like to suggest new assays to be part of ProteinGym, please raise an issue on this repository with a `new_assay' label. The criteria we typically consider for inclusion are as follows:
If you would like new baselines to be included in ProteinGym (ie., website, performance files, detailed scoring files), please follow the following steps:
At this point we are only considering new baselines satisfying the following conditions:
At this stage, we are only considering requests for which all model scores for all mutants in a given benchmark (substitution or indel) are provided by the requester; but we are planning on regularly scoring new baselines ourselves for methods with wide adoption by the community and/or suggestions with many upvotes.
12 December 2023: The code for training and evaluating supervised models is currently shared in https://github.com/OATML-Markslab/ProteinNPT. We are in the process of integrating the code into this repo.
If you would like to compute all performance metrics for the various benchmarks, please follow the following steps:
Our codebase leveraged code from the following repositories to compute baselines:
We would like to thank the GEMME team for providing model scores on an earlier version of the benchmark (ProteinGym v0.1), and the ProtSSN, SaProt, PoET, MULAN, VespaG, ProSST, ESCOTT, VenusREM, RSALOR, SiteRM and AIDO teams for integrating their model in the ProteinGym repo.
Special thanks the teams of experimentalists who developed and performed the assays that ProteinGym is built on. If you are using ProteinGym in your work, please consider citing the corresponding papers. To facilitate this, we have prepared a file (assays.bib) containing the bibtex entries for all these papers.
This project is available under the MIT license found in the LICENSE file in this GitHub repository.
If you use ProteinGym in your work, please cite the following paper:
@inproceedings{NEURIPS2023_cac723e5,
author = {Notin, Pascal and Kollasch, Aaron and Ritter, Daniel and van Niekerk, Lood and Paul, Steffanie and Spinner, Han and Rollins, Nathan and Shaw, Ada and Orenbuch, Rose and Weitzman, Ruben and Frazer, Jonathan and Dias, Mafalda and Franceschi, Dinko and Gal, Yarin and Marks, Debora},
booktitle = {Advances in Neural Information Processing Systems},
editor = {A. Oh and T. Neumann and A. Globerson and K. Saenko and M. Hardt and S. Levine},
pages = {64331--64379},
publisher = {Curran Associates, Inc.},
title = {ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design},
url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/cac723e5ff29f65e3fcbb0739ae91bee-Paper-Datasets_and_Benchmarks.pdf},
volume = {36},
year = {2023}
}
76 followers · starred Mar 2024
57 followers · starred Jun 2023
HTML
50.7%
Python
43.5%
TeX
4.2%
Shell
1.4%