BAAI/OPI

Dataset

Github:

10

72 commits

2 linked in READMEs

updated Mar 12, 2025

See the code

README

image.png

Github:

https://github.com/baaihealth/opi

Paper:

OPI: An Open Instruction Dataset for Adapting Large Language Models to Protein-Related Tasks has been accepted by NeurIPS 2024 Workshop: Foundation Models for Science: Progress, Opportunities, and Challenges.

Dataset Overview

Dataset size:
- Thera are 1.64M samples, including training (1,615,661) and testing (26,607) sets, in OPI dataset, covering 9 protein-related tasks.

We are excited to announce the release of the Open Protein Instructions (OPI) dataset, a curated collection of instructions covering 9 tasks for adapting LLMs to protein biology. The dataset is designed to advance LLM-driven research in the field of protein biology. We welcome contributions and enhancements to this dataset from the community.

OPI is the initial part of Open Biology Instructions(OBI) project, together with the subsequent Open Molecule Instructions(OMI), Open DNA Instructions(ODI), Open RNA Instructions(ORI) and Open Single-cell Instructions (OSCI). OBI is a project which aims to fully leverage the potential ability of Large Language Models(LLMs), especially the scientific LLMs like Galactica, to facilitate research in AI for Life Science community. While OBI is still in an early stage, we hope to provide a starting point for the community to bridge LLMs and biological domain knowledge.

Dataset Update

The previous version of OPI dataset is based on the release 2022_01 of UniProtKB/Swiss-Prot protein knowledgebase. At current, OPI is updated to contain the latest release 2023_05, which can be accessed via the dataset file OPI_updated_160k.json.

Reference:

OPI Dataset Construction Pipeline

The OPI dataset is curated on our own by extracting key information from Swiss-Prot database. The following figure shows the general construction process. image.png

OPI Dataset Folder Structure

The OPI dataset is organized into the three subfoldersβ€”AP, KM, and SUβ€”by in the OPI_DATA directory within this repository, where you can find a subset for each specific task as well as the full dataset file: OPI_full_1.61M_train.json.

./OPI_DATA/
└── SU
β”‚   β”œβ”€β”€ EC_number
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   β”œβ”€β”€ CLEAN_EC_number_new_test.jsonl
β”‚   β”‚   β”‚   └── CLEAN_EC_number_price_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       β”œβ”€β”€ CLEAN_EC_number_train.json
β”‚   β”œβ”€β”€ Fold_type
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   └── fold_type_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       └── fold_type_train.json
β”‚   └── Subcellular_localization
β”‚       β”œβ”€β”€ test
β”‚       β”‚   β”œβ”€β”€ subcell_loc_test.jsonl
β”‚       └── train
            └── subcell_loc_train.json
β”œβ”€β”€ AP
β”‚   └── Keywords
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   β”œβ”€β”€ CASPSimilarSeq_keywords_test.jsonl
β”‚   β”‚   β”‚   β”œβ”€β”€ IDFilterSeq_keywords_test.jsonl
β”‚   β”‚   β”‚   └── UniProtSeq_keywords_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       β”œβ”€β”€ keywords_train.json
β”‚   β”œβ”€β”€ GO
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   β”œβ”€β”€ CASPSimilarSeq_go_terms_test.jsonl
β”‚   β”‚   β”‚   β”œβ”€β”€ IDFilterSeq_go_terms_test.jsonl
β”‚   β”‚   β”‚   └── UniProtSeq_go_terms_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       β”œβ”€β”€ go_terms_train.json
β”‚   β”œβ”€β”€ Function
β”‚       β”œβ”€β”€ test
β”‚       β”‚   β”œβ”€β”€ CASPSimilarSeq_function_test.jsonl
β”‚       β”‚   β”œβ”€β”€ IDFilterSeq_function_test.jsonl
β”‚       β”‚   └── UniProtSeq_function_test.jsonl
β”‚       └── train
β”‚           β”œβ”€β”€ function_train.json
β”œβ”€β”€ KM
    └── gSymbol2Tissue
    β”‚   β”œβ”€β”€ test
    β”‚   β”‚   └── gene_symbol_to_tissue_test.jsonl
    β”‚   └── train
    β”‚       └── gene_symbol_to_tissue_train.json
    β”œβ”€β”€ gSymbol2Cancer
    β”‚   β”œβ”€β”€ test
    β”‚   β”‚   └── gene_symbol_to_cancer_test.jsonl
    β”‚   └── train
    β”‚       └── gene_symbol_to_cancer_train.json
    β”œβ”€β”€ gName2Cancer
        β”œβ”€β”€ test
        β”‚   └── gene_name_to_cancer_test.jsonl
        └── train
            └── gene_name_to_cancer_train.json

Dataset Examples

An example of OPI training data:

instruction: 
    What is the EC classification of the input protein sequence based on its biological function?
input:                         
    MGLVSSKKPDKEKPIKEKDKGQWSPLKVSAQDKDAPPLPPLVVFNHLTPPPPDEHLDEDKHFVVALYDYTAMNDRDLQMLKGEKLQVLKGTGDWWLARS
    LVTGREGYVPSNFVARVESLEMERWFFRSQGRKEAERQLLAPINKAGSFLIRESETNKGAFSLSVKDVTTQGELIKHYKIRCLDEGGYYISPRITFPSL
    QALVQHYSKKGDGLCQRLTLPCVRPAPQNPWAQDEWEIPRQSLRLVRKLGSGQFGEVWMGYYKNNMKVAIKTLKEGTMSPEAFLGEANVMKALQHERLV
    RLYAVVTKEPIYIVTEYMARGCLLDFLKTDEGSRLSLPRLIDMSAQIAEGMAYIERMNSIHRDLRAANILVSEALCCKIADFGLARIIDSEYTAQEGAK
    FPIKWTAPEAIHFGVFTIKADVWSFGVLLMEVVTYGRVPYPGMSNPEVIRNLERGYRMPRPDTCPPELYRGVIAECWRSRPEERPTFEFLQSVLEDFYT
    ATERQYELQP
output: 
    2.7.10.2

An example of OPI testing data:

{"id": "seed_task_0", "name": "EC number of price dataset from CLEAN", "instruction":
"Return the EC number of the protein sequence.", "instances": [{"input":
"MAIPPYPDFRSAAFLRQHLRATMAFYDPVATDASGGQFHFFLDDGTVYNTHTRHLVSATRFVVTHAMLYRTTGEARYQVGMRHALEFLRTAFLDPATGGY
AWLIDWQDGRATVQDTTRHCYGMAFVMLAYARAYEAGVPEARVWLAEAFDTAEQHFWQPAAGLYADEASPDWQLTSYRGQNANMHACEAMISAFRATGERR
YIERAEQLAQGICQRQAALSDRTHAPAAEGWVWEHFHADWSVDWDYNRHDRSNIFRPWGYQVGHQTEWAKLLLQLDALLPADWHLPCAQRLFDTAVERGWD
AEHGGLYYGMAPDGSICDDGKYHWVQAESMAAAAVLAVRTGDARYWQWYDRIWAYCWAHFVDHEHGAWFRILHRDNRNTTREKSNAGKVDYHNMGACYDVL
LWALDAPGFSKESRSAALGRP", "output": "5.3.1.7"}], "is_classification": false}

OPEval: Nine evaluation tasks using the OPI dataset

To assess the effectiveness of instruction tuning with the OPI dataset, we developed OPEval, which comprises three categories of evaluation tasks. Each category includes three specific tasks. The table below outlines the task types, names, and the corresponding sizes of the training and testing sets.

Task TypeType Abbr.Task NameTask Abbr.Training set sizeTesting set size
Sequence UnderstandingSUEC Number PredictionEC_number227,362392 (NEW-392), 149 (Price-149)
Fold Type PredictionFold_type12,312718 (Fold), 1254 (Superfamily), 1272 (Family)
Subcellular Localization PredictionSubcellular_localization11,2302,772
Annotation PredictionAPFunction Keywords PredictionKeywords451,618184 (CASPSimilarSeq), 1,112 (IDFilterSeq), 4562 (UniprotSeq)
Gene Ontology(GO) Terms PredictionGO451,618184 (CASPSimilarSeq), 1,112 (IDFilterSeq), 4562 (UniprotSeq)
Function Description PredictionFunction451,618184 (CASPSimilarSeq), 1,112 (IDFilterSeq), 4562 (UniprotSeq)
Knowledge MiningKMTissue Location Prediction from Gene SymbolgSymbol2Tissue8,7232,181
Cancer Prediction from Gene SymbolgSymbol2Cancer590148
Cancer Prediction from Gene NamegName2Cancer590148

License

The dataset is licensed under a Creative Commons Attribution Non Commercial 4.0 License. The use of this dataset should also abide by the original License & Disclaimer and Privacy Notice of UniProt.

AI4Science
biology
instruction tuning
Life Science
protein

BAAI/OPI

Dataset

Github:

10

72 commits

2 linked in READMEs

updated Mar 12, 2025

See the code

README

image.png

Github:

https://github.com/baaihealth/opi

Paper:

OPI: An Open Instruction Dataset for Adapting Large Language Models to Protein-Related Tasks has been accepted by NeurIPS 2024 Workshop: Foundation Models for Science: Progress, Opportunities, and Challenges.

Dataset Overview

Dataset size:
- Thera are 1.64M samples, including training (1,615,661) and testing (26,607) sets, in OPI dataset, covering 9 protein-related tasks.

We are excited to announce the release of the Open Protein Instructions (OPI) dataset, a curated collection of instructions covering 9 tasks for adapting LLMs to protein biology. The dataset is designed to advance LLM-driven research in the field of protein biology. We welcome contributions and enhancements to this dataset from the community.

OPI is the initial part of Open Biology Instructions(OBI) project, together with the subsequent Open Molecule Instructions(OMI), Open DNA Instructions(ODI), Open RNA Instructions(ORI) and Open Single-cell Instructions (OSCI). OBI is a project which aims to fully leverage the potential ability of Large Language Models(LLMs), especially the scientific LLMs like Galactica, to facilitate research in AI for Life Science community. While OBI is still in an early stage, we hope to provide a starting point for the community to bridge LLMs and biological domain knowledge.

Dataset Update

The previous version of OPI dataset is based on the release 2022_01 of UniProtKB/Swiss-Prot protein knowledgebase. At current, OPI is updated to contain the latest release 2023_05, which can be accessed via the dataset file OPI_updated_160k.json.

Reference:

OPI Dataset Construction Pipeline

The OPI dataset is curated on our own by extracting key information from Swiss-Prot database. The following figure shows the general construction process. image.png

OPI Dataset Folder Structure

The OPI dataset is organized into the three subfoldersβ€”AP, KM, and SUβ€”by in the OPI_DATA directory within this repository, where you can find a subset for each specific task as well as the full dataset file: OPI_full_1.61M_train.json.

./OPI_DATA/
└── SU
β”‚   β”œβ”€β”€ EC_number
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   β”œβ”€β”€ CLEAN_EC_number_new_test.jsonl
β”‚   β”‚   β”‚   └── CLEAN_EC_number_price_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       β”œβ”€β”€ CLEAN_EC_number_train.json
β”‚   β”œβ”€β”€ Fold_type
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   └── fold_type_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       └── fold_type_train.json
β”‚   └── Subcellular_localization
β”‚       β”œβ”€β”€ test
β”‚       β”‚   β”œβ”€β”€ subcell_loc_test.jsonl
β”‚       └── train
            └── subcell_loc_train.json
β”œβ”€β”€ AP
β”‚   └── Keywords
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   β”œβ”€β”€ CASPSimilarSeq_keywords_test.jsonl
β”‚   β”‚   β”‚   β”œβ”€β”€ IDFilterSeq_keywords_test.jsonl
β”‚   β”‚   β”‚   └── UniProtSeq_keywords_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       β”œβ”€β”€ keywords_train.json
β”‚   β”œβ”€β”€ GO
β”‚   β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”‚   β”œβ”€β”€ CASPSimilarSeq_go_terms_test.jsonl
β”‚   β”‚   β”‚   β”œβ”€β”€ IDFilterSeq_go_terms_test.jsonl
β”‚   β”‚   β”‚   └── UniProtSeq_go_terms_test.jsonl
β”‚   β”‚   └── train
β”‚   β”‚       β”œβ”€β”€ go_terms_train.json
β”‚   β”œβ”€β”€ Function
β”‚       β”œβ”€β”€ test
β”‚       β”‚   β”œβ”€β”€ CASPSimilarSeq_function_test.jsonl
β”‚       β”‚   β”œβ”€β”€ IDFilterSeq_function_test.jsonl
β”‚       β”‚   └── UniProtSeq_function_test.jsonl
β”‚       └── train
β”‚           β”œβ”€β”€ function_train.json
β”œβ”€β”€ KM
    └── gSymbol2Tissue
    β”‚   β”œβ”€β”€ test
    β”‚   β”‚   └── gene_symbol_to_tissue_test.jsonl
    β”‚   └── train
    β”‚       └── gene_symbol_to_tissue_train.json
    β”œβ”€β”€ gSymbol2Cancer
    β”‚   β”œβ”€β”€ test
    β”‚   β”‚   └── gene_symbol_to_cancer_test.jsonl
    β”‚   └── train
    β”‚       └── gene_symbol_to_cancer_train.json
    β”œβ”€β”€ gName2Cancer
        β”œβ”€β”€ test
        β”‚   └── gene_name_to_cancer_test.jsonl
        └── train
            └── gene_name_to_cancer_train.json

Dataset Examples

An example of OPI training data:

instruction: 
    What is the EC classification of the input protein sequence based on its biological function?
input:                         
    MGLVSSKKPDKEKPIKEKDKGQWSPLKVSAQDKDAPPLPPLVVFNHLTPPPPDEHLDEDKHFVVALYDYTAMNDRDLQMLKGEKLQVLKGTGDWWLARS
    LVTGREGYVPSNFVARVESLEMERWFFRSQGRKEAERQLLAPINKAGSFLIRESETNKGAFSLSVKDVTTQGELIKHYKIRCLDEGGYYISPRITFPSL
    QALVQHYSKKGDGLCQRLTLPCVRPAPQNPWAQDEWEIPRQSLRLVRKLGSGQFGEVWMGYYKNNMKVAIKTLKEGTMSPEAFLGEANVMKALQHERLV
    RLYAVVTKEPIYIVTEYMARGCLLDFLKTDEGSRLSLPRLIDMSAQIAEGMAYIERMNSIHRDLRAANILVSEALCCKIADFGLARIIDSEYTAQEGAK
    FPIKWTAPEAIHFGVFTIKADVWSFGVLLMEVVTYGRVPYPGMSNPEVIRNLERGYRMPRPDTCPPELYRGVIAECWRSRPEERPTFEFLQSVLEDFYT
    ATERQYELQP
output: 
    2.7.10.2

An example of OPI testing data:

{"id": "seed_task_0", "name": "EC number of price dataset from CLEAN", "instruction":
"Return the EC number of the protein sequence.", "instances": [{"input":
"MAIPPYPDFRSAAFLRQHLRATMAFYDPVATDASGGQFHFFLDDGTVYNTHTRHLVSATRFVVTHAMLYRTTGEARYQVGMRHALEFLRTAFLDPATGGY
AWLIDWQDGRATVQDTTRHCYGMAFVMLAYARAYEAGVPEARVWLAEAFDTAEQHFWQPAAGLYADEASPDWQLTSYRGQNANMHACEAMISAFRATGERR
YIERAEQLAQGICQRQAALSDRTHAPAAEGWVWEHFHADWSVDWDYNRHDRSNIFRPWGYQVGHQTEWAKLLLQLDALLPADWHLPCAQRLFDTAVERGWD
AEHGGLYYGMAPDGSICDDGKYHWVQAESMAAAAVLAVRTGDARYWQWYDRIWAYCWAHFVDHEHGAWFRILHRDNRNTTREKSNAGKVDYHNMGACYDVL
LWALDAPGFSKESRSAALGRP", "output": "5.3.1.7"}], "is_classification": false}

OPEval: Nine evaluation tasks using the OPI dataset

To assess the effectiveness of instruction tuning with the OPI dataset, we developed OPEval, which comprises three categories of evaluation tasks. Each category includes three specific tasks. The table below outlines the task types, names, and the corresponding sizes of the training and testing sets.

Task TypeType Abbr.Task NameTask Abbr.Training set sizeTesting set size
Sequence UnderstandingSUEC Number PredictionEC_number227,362392 (NEW-392), 149 (Price-149)
Fold Type PredictionFold_type12,312718 (Fold), 1254 (Superfamily), 1272 (Family)
Subcellular Localization PredictionSubcellular_localization11,2302,772
Annotation PredictionAPFunction Keywords PredictionKeywords451,618184 (CASPSimilarSeq), 1,112 (IDFilterSeq), 4562 (UniprotSeq)
Gene Ontology(GO) Terms PredictionGO451,618184 (CASPSimilarSeq), 1,112 (IDFilterSeq), 4562 (UniprotSeq)
Function Description PredictionFunction451,618184 (CASPSimilarSeq), 1,112 (IDFilterSeq), 4562 (UniprotSeq)
Knowledge MiningKMTissue Location Prediction from Gene SymbolgSymbol2Tissue8,7232,181
Cancer Prediction from Gene SymbolgSymbol2Cancer590148
Cancer Prediction from Gene NamegName2Cancer590148

License

The dataset is licensed under a Creative Commons Attribution Non Commercial 4.0 License. The use of this dataset should also abide by the original License & Disclaimer and Privacy Notice of UniProt.

AI4Science
biology
instruction tuning
Life Science
protein