Align-to-Innovate/data-saturation-and-scaling

All inference and training scripts and relevant data for Spinner, A., DeBenedictis, E., & Hudson, C. M. (2025). Scaling and Data Saturation in Protein Language Models. Proceedings of the ICML GenBio Workshop 2025.

6

stars

30

commits

Jupyter Notebook

primary language

Aug 15, 2025

updated

README

Scaling and Data Saturation in Protein Language Models

This repository contains all inference and training scripts, along with relevant data processing utilities, used in:

Spinner, A., DeBenedictis, E., & Hudson, C. M. (2025).
Scaling and Data Saturation in Protein Language Models.
Proceedings of the ICML GenBio Workshop 2025.


Repository Structure

  • Analysis/ – notebooks and scripts for analysis and visualization
  • Data/ – input datasets used in the study
  • Results/ – model outputs, figures, and metrics

Requirements

Install dependencies with:

pip install -r requirements.txt

Data

We used the Substitution DMS dataset from the ProteinGym benchmark for all supervised experiments.

Specifically, the following files in Data/ were downloaded from the ProteinGym repository:

For full dataset access, download instructions are available on the ProteinGym Resources page.

Contributors

avivspinner

26 commits

aviv-beep

4 commits

Align-to-Innovate/data-saturation-and-scaling

All inference and training scripts and relevant data for Spinner, A., DeBenedictis, E., & Hudson, C. M. (2025). Scaling and Data Saturation in Protein Language Models. Proceedings of the ICML GenBio Workshop 2025.

6

stars

30

commits

Jupyter Notebook

primary language

Aug 15, 2025

updated

README

Scaling and Data Saturation in Protein Language Models

This repository contains all inference and training scripts, along with relevant data processing utilities, used in:

Spinner, A., DeBenedictis, E., & Hudson, C. M. (2025).
Scaling and Data Saturation in Protein Language Models.
Proceedings of the ICML GenBio Workshop 2025.


Repository Structure

  • Analysis/ – notebooks and scripts for analysis and visualization
  • Data/ – input datasets used in the study
  • Results/ – model outputs, figures, and metrics

Requirements

Install dependencies with:

pip install -r requirements.txt

Data

We used the Substitution DMS dataset from the ProteinGym benchmark for all supervised experiments.

Specifically, the following files in Data/ were downloaded from the ProteinGym repository:

For full dataset access, download instructions are available on the ProteinGym Resources page.

Contributors

avivspinner

26 commits

aviv-beep

4 commits

Languages

Jupyter Notebook

99.4%