wenyuan-wu/chemprot-drugprot_testing_ground

1

stars

74

commits

Jupyter Notebook

primary language

Sep 28, 2021

updated

README

ChemProt/DrugProt Testing Ground

Testing ground for task 5 from BioCreative VI

Task 5: Text mining chemical-protein interactions (CHEMPROT)

The aim of the CHEMPROT task of BioCreative VI is to promote the development and evaluation of systems that are able to automatically detect in running text (PubMed abstracts) relations between chemical compounds/drug and genes/proteins. We will therefore release a manually annotated corpus, the CHEMPROT corpus, where domain experts have exhaustively labeled: (a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (CHEMPROT relation classes).

Working Environment

  • Ubuntu 20.04
  • CUDA 11.2
  • PyTorch 1.8.1+cu111

Data Structure

data/
├── chemprot_development
│   ├── chemprot_development_abstracts.tsv
│   ├── chemprot_development_entities.tsv
│   ├── chemprot_development_gold_standard.tsv
│   ├── chemprot_development_relations.tsv
│   └── Readme.pdf
├── chemprot_sample
│   ├── chemprot_sample_abstracts.tsv
│   ├── chemprot_sample_entities.tsv
│   ├── chemprot_sample_gold_standard.tsv
│   ├── chemprot_sample_predictions_eval.txt
│   ├── chemprot_sample_predictions.tsv
│   ├── chemprot_sample_relations.tsv
│   ├── guidelines
│   │   ├── CEM_guidelines.pdf
│   │   ├── CHEMPROT_guidelines_v6.pdf
│   │   └── GPRO_guidelines.pdf
│   └── Readme.pdf
├── chemprot_test_gs
│   ├── chemprot_test_abstracts_gs.tsv
│   ├── chemprot_test_entities_gs.tsv
│   ├── chemprot_test_gold_standard.tsv
│   ├── chemprot_test_relations_gs.tsv
│   └── readme_test_gs.pdf
└── chemprot_training
    ├── chemprot_training_abstracts.tsv
    ├── chemprot_training_entities.tsv
    ├── chemprot_training_gold_standard.tsv
    ├── chemprot_training_relations.tsv
    └── Readme.pdf

5 directories, 25 files

Example

PMID

10471277

Title

Probing the salmeterol binding site on the beta 2-adrenergic receptor using a novel photoaffinity ligand,[(125)I]iodoazidosalmeterol.

Abstract

Salmeterol is a long-acting beta2-adrenergic receptor (beta 2AR) agonist used clinically to treat asthma. In addition to binding at the active agonist site, it has been proposed that salmeterol also binds with very high affinity at a second site, termed the "exosite", and that this exosite contributes to the long duration of action of salmeterol. To determine the position of the phenyl ring of the aralkyloxyalkyl side chain of salmeterol in the beta 2AR binding site, we designed and synthesized the agonist photoaffinity label [(125)I]iodoazidosalmeterol ([125I]IAS). In direct adenylyl cyclase activation, in effects on adenylyl cyclase after pretreatment of intact cells, and in guinea pig tracheal relaxation assays, IAS and the parent drug salmeterol behave essentially the same. Significantly, the photoreactive azide of IAS is positioned on the phenyl ring at the end of the molecule which is thought to be involved in exosite binding. Carrier-free radioiodinated [125I]IAS was used to photolabel epitope-tagged human beta 2AR in membranes prepared from stably transfected HEK 293 cells. Labeling with [(125)I]IAS was blocked by 10 microM (-)-alprenolol and inhibited by addition of GTP gamma S, and [125I]IAS migrated at the same position on an SDS-PAGE gel as the beta 2AR labeled by the antagonist photoaffinity label [125I]iodoazidobenzylpindolol ([125I]IABP). The labeled receptor was purified on a nickel affinity column and cleaved with factor Xa protease at a specific sequence in the large loop between transmembrane segments 5 and 6, yielding two peptides. While the control antagonist photoaffinity label [125I]IABP labeled both the large N-terminal fragment [containing transmembranes (TMs) 1-5] and the smaller C-terminal fragment (containing TMs 6 and 7), essentially all of the [125I]IAS labeling was on the smaller C-terminal peptide containing TMs 6 and 7. This direct biochemical evidence demonstrates that when salmeterol binds to the receptor, its hydrophobic aryloxyalkyl tail is positioned near TM 6 and/or TM 7. A model of IAS binding to the beta 2AR is proposed.

Entity mention annotations

PMIDEntityTypeStartEndText
10471277T1CHEMICAL135145Salmeterol
10471277T2CHEMICAL12481259[(125)I]IAS
10471277T3CHEMICAL12851299(-)-alprenolol
10471277T4CHEMICAL13291332GTP
10471277T5CHEMICAL13461355[125I]IAS
10471277T6CHEMICAL14671496[125I]iodoazidobenzylpindolol
10471277T7CHEMICAL14981508[125I]IABP
10471277T8CHEMICAL15501556nickel
10471277T9CHEMICAL17621772[125I]IABP
10471277T10CHEMICAL17961797N
10471277T11CHEMICAL18701871C
10471277T12CHEMICAL19391948[125I]IAS
10471277T13CHEMICAL318328salmeterol
10471277T14CHEMICAL19771978C
10471277T15CHEMICAL20762086salmeterol
10471277T16CHEMICAL21262138aryloxyalkyl
10471277T17CHEMICAL21922195IAS
10471277T18CHEMICAL472482salmeterol
10471277T19CHEMICAL517523phenyl
10471277T20CHEMICAL566576salmeterol
10471277T21CHEMICAL667694[(125)I]iodoazidosalmeterol
10471277T22CHEMICAL696705[125I]IAS
10471277T23CHEMICAL718726adenylyl
10471277T24CHEMICAL761769adenylyl
10471277T25CHEMICAL860863IAS
10471277T26CHEMICAL884894salmeterol
10471277T27CHEMICAL957962azide
10471277T28CHEMICAL966969IAS
10471277T29CHEMICAL991997phenyl
10471277T30CHEMICAL11101119[125I]IAS
10471277T31CHEMICAL106133[(125)I]iodoazidosalmeterol
10471277T33GENE-N11581172human beta 2AR
10471277T34GENE-N14121420beta 2AR
10471277T35GENE-N15901599factor Xa
10471277T36GENE-N22112219beta 2AR
10471277T37GENE-N163188beta2-adrenergic receptor
10471277T38GENE-N584592beta 2AR
10471277T39GENE-N190198beta 2AR
10471277T40GENE-N4369beta 2-adrenergic receptor
10471277T32CHEMICAL1222salmeterol
10471277T41CHEMICAL536551aralkyloxyalkyl
10471277T42GENE-Y718734adenylyl cyclase
10471277T44GENE-Y761777adenylyl cyclase

CHEMPROT detailed relation annotations

PMIDCPR GroupEvaluation TypeCPRArg1Arg2
10471277CPR:2NDIRECT-REGULATORArg1:T30Arg2:T33
10471277CPR:2NDIRECT-REGULATORArg1:T31Arg2:T40
10471277CPR:2NDIRECT-REGULATORArg1:T32Arg2:T40
10471277CPR:5YAGONISTArg1:T1Arg2:T37
10471277CPR:5YAGONISTArg1:T1Arg2:T39
10471277CPR:2NDIRECT-REGULATORArg1:T20Arg2:T38
10471277CPR:2NDIRECT-REGULATORArg1:T41Arg2:T38
10471277CPR:2NDIRECT-REGULATORArg1:T19Arg2:T38
10471277CPR:5YAGONISTArg1:T21Arg2:T38
10471277CPR:5YAGONISTArg1:T22Arg2:T38
10471277CPR:3YUPREGULATORArg1:T25Arg2:T42
10471277CPR:3YUPREGULATORArg1:T26Arg2:T42
10471277CPR:3YUPREGULATORArg1:T25Arg2:T44
10471277CPR:3YUPREGULATORArg1:T26Arg2:T44
10471277CPR:6YANTAGONISTArg1:T6Arg2:T34
10471277CPR:6YANTAGONISTArg1:T7Arg2:T34
10471277CPR:2NDIRECT-REGULATORArg1:T5Arg2:T34
10471277CPR:2NDIRECT-REGULATORArg1:T2Arg2:T34
10471277CPR:2NDIRECT-REGULATORArg1:T17Arg2:T36
10471277CPR:2NDIRECT-REGULATORArg1:T3Arg2:T34

CHEMPROT task Gold Standard data and predictions

  • File: chemprot_sample_gold_standard.tsv
PMIDCPR GroupArg1Arg2
10471277CPR:5Arg1:T1Arg2:T37
10471277CPR:5Arg1:T1Arg2:T39
10471277CPR:5Arg1:T21Arg2:T38
10471277CPR:5Arg1:T22Arg2:T38
10471277CPR:3Arg1:T25Arg2:T42
10471277CPR:3Arg1:T26Arg2:T42
10471277CPR:3Arg1:T25Arg2:T44
10471277CPR:3Arg1:T26Arg2:T44
10471277CPR:6Arg1:T6Arg2:T34
10471277CPR:6Arg1:T7Arg2:T34
  • File: chemprot_sample_predictions.tsv
PMIDCPR GroupArg1Arg2
10471277CPR:5Arg1:T17Arg2:T36
10471277CPR:5Arg1:T1Arg2:T37
10471277CPR:5Arg1:T1Arg2:T39
10471277CPR:5Arg1:T20Arg2:T38
10471277CPR:5Arg1:T21Arg2:T38
10471277CPR:5Arg1:T22Arg2:T38
10471277CPR:5Arg1:T31Arg2:T40
10471277CPR:5Arg1:T32Arg2:T40
10471277CPR:6Arg1:T6Arg2:T34
10471277CPR:6Arg1:T7Arg2:T34

Knowledge Graph results

kG

Resources

BLUE, the Biomedical Language Understanding Evaluation benchmark

Chemical-protein Interaction Extraction via Gaussian Probability Distribution and External Biomedical Knowledge

Scibert

Model file of scibert_scivocab_uncased:

wget "https://s3-us-west-2.amazonaws.com/ai2-s2-research/scibert/pytorch_models/scibert_scivocab_uncased.tar"

info

dist info

SetPMIDPosNeg
train35001695648621
dev75037129690
test107500220315
NONE                      48197
INHIBITOR                  5326
DIRECT-REGULATOR           2153
SUBSTRATE                  1988
ACTIVATOR                  1381
INDIRECT-UPREGULATOR       1337
INDIRECT-DOWNREGULATOR     1316
ANTAGONIST                  937
PRODUCT-OF                  917
PART-OF                     877
AGONIST                     646
AGONIST-ACTIVATOR            28
SUBSTRATE_PRODUCT-OF         24
AGONIST-INHIBITOR            10
Name: relation, dtype: int64

NONE                      9683
INHIBITOR                 1136
SUBSTRATE                  494
DIRECT-REGULATOR           452
INDIRECT-DOWNREGULATOR     329
INDIRECT-UPREGULATOR       298
PART-OF                    254
ACTIVATOR                  234
ANTAGONIST                 214
PRODUCT-OF                 156
AGONIST                    126
AGONIST-ACTIVATOR           10
SUBSTRATE_PRODUCT-OF         3
AGONIST-INHIBITOR            2
Name: relation, dtype: int64

LM Results

Base Language ModelAnnotation StyleNegative Samples in Training SetF1 Score on Development Set
bert_base_uncasednoneno0.552
scibert_uncasedscibertno0.739
biobert_base_v1.1biobertno0.727
bert_base_uncasednoneyes0.348
scibert_uncasedscibertyes0.586
biobert_base_v1.1biobertyes0.581

New Results

Base Language ModelAnnotation StyleKnowledge Graph ModelF1 Score on Development Set
bert_base_uncasedrawnone0.765
bert_base_uncasedSciBertnone0.864
bert_base_uncasedBioBertnone0.863
scibert_scivocab_uncasedrawnone0.776
scibert_scivocab_uncasedSciBertnone0.876
scibert_scivocab_uncasedBioBertnone0.869
biobert-base-cased-v1.1rawnone0.776
biobert-base-cased-v1.1SciBertnone0.876
biobert-base-cased-v1.1BioBertnone0.875
bert_base_uncasedrawTransE0.768
bert_base_uncasedSciBertTransE0.858
bert_base_uncasedBioBertTransE0.856
scibert_scivocab_uncasedrawTransE0.776
scibert_scivocab_uncasedSciBertTransE0.872
scibert_scivocab_uncasedBioBertTransE0.868
biobert-base-cased-v1.1rawTransE0.778
biobert-base-cased-v1.1SciBertTransE0.873
biobert-base-cased-v1.1BioBertTransE0.863
bert_base_uncasedrawPairRE0.769
bert_base_uncasedSciBertPairRE0.858
bert_base_uncasedBioBertPairRE0.857
scibert_scivocab_uncasedrawPairRE0.779
scibert_scivocab_uncasedSciBertPairRE0.871
scibert_scivocab_uncasedBioBertPairRE0.867
biobert-base-cased-v1.1rawPairRE0.778
biobert-base-cased-v1.1SciBertPairRE0.872
biobert-base-cased-v1.1BioBertPairRE0.862

Code structure

drugprot_preprocess_data

2021-09-28 02:08:03,003 - INFO - model: bert-base-uncased, annotation: raw
2021-09-28 02:08:03,003 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:04,139 - INFO - f1 score: 0.765
2021-09-28 02:08:04,139 - INFO - model: bert-base-uncased, annotation: sci
2021-09-28 02:08:04,139 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:04,939 - INFO - f1 score: 0.864
2021-09-28 02:08:04,940 - INFO - model: bert-base-uncased, annotation: bio
2021-09-28 02:08:04,940 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:05,764 - INFO - f1 score: 0.863
2021-09-28 02:08:05,764 - INFO - model: allenai/scibert_scivocab_uncased, annotation: raw
2021-09-28 02:08:05,764 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:06,553 - INFO - f1 score: 0.776
2021-09-28 02:08:06,553 - INFO - model: allenai/scibert_scivocab_uncased, annotation: sci
2021-09-28 02:08:06,553 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:07,344 - INFO - f1 score: 0.876
2021-09-28 02:08:07,344 - INFO - model: allenai/scibert_scivocab_uncased, annotation: bio
2021-09-28 02:08:07,344 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:08,177 - INFO - f1 score: 0.869
2021-09-28 02:08:08,177 - INFO - model: dmis-lab/biobert-base-cased-v1.1, annotation: raw
2021-09-28 02:08:08,177 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:08,964 - INFO - f1 score: 0.776
2021-09-28 02:08:08,964 - INFO - model: dmis-lab/biobert-base-cased-v1.1, annotation: sci
2021-09-28 02:08:08,964 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:09,750 - INFO - f1 score: 0.876
2021-09-28 02:08:09,750 - INFO - model: dmis-lab/biobert-base-cased-v1.1, annotation: bio
2021-09-28 02:08:09,750 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:10,545 - INFO - f1 score: 0.875
2021-09-28 02:08:10,545 - INFO - kg_model: TransE, lm_model: bert-base-uncased, annotation: raw
2021-09-28 02:08:10,545 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:11,366 - INFO - f1 score: 0.768
2021-09-28 02:08:11,366 - INFO - kg_model: TransE, lm_model: bert-base-uncased, annotation: sci
2021-09-28 02:08:11,367 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:12,151 - INFO - f1 score: 0.858
2021-09-28 02:08:12,151 - INFO - kg_model: TransE, lm_model: bert-base-uncased, annotation: bio
2021-09-28 02:08:12,151 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:12,936 - INFO - f1 score: 0.856
2021-09-28 02:08:12,937 - INFO - kg_model: TransE, lm_model: allenai/scibert_scivocab_uncased, annotation: raw
2021-09-28 02:08:12,937 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:13,723 - INFO - f1 score: 0.776
2021-09-28 02:08:13,723 - INFO - kg_model: TransE, lm_model: allenai/scibert_scivocab_uncased, annotation: sci
2021-09-28 02:08:13,723 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:14,541 - INFO - f1 score: 0.872
2021-09-28 02:08:14,541 - INFO - kg_model: TransE, lm_model: allenai/scibert_scivocab_uncased, annotation: bio
2021-09-28 02:08:14,541 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:15,326 - INFO - f1 score: 0.868
2021-09-28 02:08:15,326 - INFO - kg_model: TransE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: raw
2021-09-28 02:08:15,326 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:16,110 - INFO - f1 score: 0.778
2021-09-28 02:08:16,110 - INFO - kg_model: TransE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: sci
2021-09-28 02:08:16,110 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:16,894 - INFO - f1 score: 0.873
2021-09-28 02:08:16,894 - INFO - kg_model: TransE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: bio
2021-09-28 02:08:16,894 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:17,712 - INFO - f1 score: 0.863
2021-09-28 02:08:17,712 - INFO - kg_model: PairRE, lm_model: bert-base-uncased, annotation: raw
2021-09-28 02:08:17,712 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:18,495 - INFO - f1 score: 0.769
2021-09-28 02:08:18,495 - INFO - kg_model: PairRE, lm_model: bert-base-uncased, annotation: sci
2021-09-28 02:08:18,495 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:19,280 - INFO - f1 score: 0.858
2021-09-28 02:08:19,280 - INFO - kg_model: PairRE, lm_model: bert-base-uncased, annotation: bio
2021-09-28 02:08:19,280 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:20,099 - INFO - f1 score: 0.857
2021-09-28 02:08:20,099 - INFO - kg_model: PairRE, lm_model: allenai/scibert_scivocab_uncased, annotation: raw
2021-09-28 02:08:20,099 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:20,883 - INFO - f1 score: 0.779
2021-09-28 02:08:20,883 - INFO - kg_model: PairRE, lm_model: allenai/scibert_scivocab_uncased, annotation: sci
2021-09-28 02:08:20,883 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:21,667 - INFO - f1 score: 0.871
2021-09-28 02:08:21,667 - INFO - kg_model: PairRE, lm_model: allenai/scibert_scivocab_uncased, annotation: bio
2021-09-28 02:08:21,667 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:22,450 - INFO - f1 score: 0.867
2021-09-28 02:08:22,450 - INFO - kg_model: PairRE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: raw
2021-09-28 02:08:22,450 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:23,268 - INFO - f1 score: 0.778
2021-09-28 02:08:23,268 - INFO - kg_model: PairRE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: sci
2021-09-28 02:08:23,268 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:24,051 - INFO - f1 score: 0.872
2021-09-28 02:08:24,052 - INFO - kg_model: PairRE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: bio
2021-09-28 02:08:24,052 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:24,841 - INFO - f1 score: 0.862

Contributors

wenyuan-wu

74 commits

wenyuan-wu/chemprot-drugprot_testing_ground

1

stars

74

commits

Jupyter Notebook

primary language

Sep 28, 2021

updated

README

ChemProt/DrugProt Testing Ground

Testing ground for task 5 from BioCreative VI

Task 5: Text mining chemical-protein interactions (CHEMPROT)

The aim of the CHEMPROT task of BioCreative VI is to promote the development and evaluation of systems that are able to automatically detect in running text (PubMed abstracts) relations between chemical compounds/drug and genes/proteins. We will therefore release a manually annotated corpus, the CHEMPROT corpus, where domain experts have exhaustively labeled: (a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (CHEMPROT relation classes).

Working Environment

  • Ubuntu 20.04
  • CUDA 11.2
  • PyTorch 1.8.1+cu111

Data Structure

data/
├── chemprot_development
│   ├── chemprot_development_abstracts.tsv
│   ├── chemprot_development_entities.tsv
│   ├── chemprot_development_gold_standard.tsv
│   ├── chemprot_development_relations.tsv
│   └── Readme.pdf
├── chemprot_sample
│   ├── chemprot_sample_abstracts.tsv
│   ├── chemprot_sample_entities.tsv
│   ├── chemprot_sample_gold_standard.tsv
│   ├── chemprot_sample_predictions_eval.txt
│   ├── chemprot_sample_predictions.tsv
│   ├── chemprot_sample_relations.tsv
│   ├── guidelines
│   │   ├── CEM_guidelines.pdf
│   │   ├── CHEMPROT_guidelines_v6.pdf
│   │   └── GPRO_guidelines.pdf
│   └── Readme.pdf
├── chemprot_test_gs
│   ├── chemprot_test_abstracts_gs.tsv
│   ├── chemprot_test_entities_gs.tsv
│   ├── chemprot_test_gold_standard.tsv
│   ├── chemprot_test_relations_gs.tsv
│   └── readme_test_gs.pdf
└── chemprot_training
    ├── chemprot_training_abstracts.tsv
    ├── chemprot_training_entities.tsv
    ├── chemprot_training_gold_standard.tsv
    ├── chemprot_training_relations.tsv
    └── Readme.pdf

5 directories, 25 files

Example

PMID

10471277

Title

Probing the salmeterol binding site on the beta 2-adrenergic receptor using a novel photoaffinity ligand,[(125)I]iodoazidosalmeterol.

Abstract

Salmeterol is a long-acting beta2-adrenergic receptor (beta 2AR) agonist used clinically to treat asthma. In addition to binding at the active agonist site, it has been proposed that salmeterol also binds with very high affinity at a second site, termed the "exosite", and that this exosite contributes to the long duration of action of salmeterol. To determine the position of the phenyl ring of the aralkyloxyalkyl side chain of salmeterol in the beta 2AR binding site, we designed and synthesized the agonist photoaffinity label [(125)I]iodoazidosalmeterol ([125I]IAS). In direct adenylyl cyclase activation, in effects on adenylyl cyclase after pretreatment of intact cells, and in guinea pig tracheal relaxation assays, IAS and the parent drug salmeterol behave essentially the same. Significantly, the photoreactive azide of IAS is positioned on the phenyl ring at the end of the molecule which is thought to be involved in exosite binding. Carrier-free radioiodinated [125I]IAS was used to photolabel epitope-tagged human beta 2AR in membranes prepared from stably transfected HEK 293 cells. Labeling with [(125)I]IAS was blocked by 10 microM (-)-alprenolol and inhibited by addition of GTP gamma S, and [125I]IAS migrated at the same position on an SDS-PAGE gel as the beta 2AR labeled by the antagonist photoaffinity label [125I]iodoazidobenzylpindolol ([125I]IABP). The labeled receptor was purified on a nickel affinity column and cleaved with factor Xa protease at a specific sequence in the large loop between transmembrane segments 5 and 6, yielding two peptides. While the control antagonist photoaffinity label [125I]IABP labeled both the large N-terminal fragment [containing transmembranes (TMs) 1-5] and the smaller C-terminal fragment (containing TMs 6 and 7), essentially all of the [125I]IAS labeling was on the smaller C-terminal peptide containing TMs 6 and 7. This direct biochemical evidence demonstrates that when salmeterol binds to the receptor, its hydrophobic aryloxyalkyl tail is positioned near TM 6 and/or TM 7. A model of IAS binding to the beta 2AR is proposed.

Entity mention annotations

PMIDEntityTypeStartEndText
10471277T1CHEMICAL135145Salmeterol
10471277T2CHEMICAL12481259[(125)I]IAS
10471277T3CHEMICAL12851299(-)-alprenolol
10471277T4CHEMICAL13291332GTP
10471277T5CHEMICAL13461355[125I]IAS
10471277T6CHEMICAL14671496[125I]iodoazidobenzylpindolol
10471277T7CHEMICAL14981508[125I]IABP
10471277T8CHEMICAL15501556nickel
10471277T9CHEMICAL17621772[125I]IABP
10471277T10CHEMICAL17961797N
10471277T11CHEMICAL18701871C
10471277T12CHEMICAL19391948[125I]IAS
10471277T13CHEMICAL318328salmeterol
10471277T14CHEMICAL19771978C
10471277T15CHEMICAL20762086salmeterol
10471277T16CHEMICAL21262138aryloxyalkyl
10471277T17CHEMICAL21922195IAS
10471277T18CHEMICAL472482salmeterol
10471277T19CHEMICAL517523phenyl
10471277T20CHEMICAL566576salmeterol
10471277T21CHEMICAL667694[(125)I]iodoazidosalmeterol
10471277T22CHEMICAL696705[125I]IAS
10471277T23CHEMICAL718726adenylyl
10471277T24CHEMICAL761769adenylyl
10471277T25CHEMICAL860863IAS
10471277T26CHEMICAL884894salmeterol
10471277T27CHEMICAL957962azide
10471277T28CHEMICAL966969IAS
10471277T29CHEMICAL991997phenyl
10471277T30CHEMICAL11101119[125I]IAS
10471277T31CHEMICAL106133[(125)I]iodoazidosalmeterol
10471277T33GENE-N11581172human beta 2AR
10471277T34GENE-N14121420beta 2AR
10471277T35GENE-N15901599factor Xa
10471277T36GENE-N22112219beta 2AR
10471277T37GENE-N163188beta2-adrenergic receptor
10471277T38GENE-N584592beta 2AR
10471277T39GENE-N190198beta 2AR
10471277T40GENE-N4369beta 2-adrenergic receptor
10471277T32CHEMICAL1222salmeterol
10471277T41CHEMICAL536551aralkyloxyalkyl
10471277T42GENE-Y718734adenylyl cyclase
10471277T44GENE-Y761777adenylyl cyclase

CHEMPROT detailed relation annotations

PMIDCPR GroupEvaluation TypeCPRArg1Arg2
10471277CPR:2NDIRECT-REGULATORArg1:T30Arg2:T33
10471277CPR:2NDIRECT-REGULATORArg1:T31Arg2:T40
10471277CPR:2NDIRECT-REGULATORArg1:T32Arg2:T40
10471277CPR:5YAGONISTArg1:T1Arg2:T37
10471277CPR:5YAGONISTArg1:T1Arg2:T39
10471277CPR:2NDIRECT-REGULATORArg1:T20Arg2:T38
10471277CPR:2NDIRECT-REGULATORArg1:T41Arg2:T38
10471277CPR:2NDIRECT-REGULATORArg1:T19Arg2:T38
10471277CPR:5YAGONISTArg1:T21Arg2:T38
10471277CPR:5YAGONISTArg1:T22Arg2:T38
10471277CPR:3YUPREGULATORArg1:T25Arg2:T42
10471277CPR:3YUPREGULATORArg1:T26Arg2:T42
10471277CPR:3YUPREGULATORArg1:T25Arg2:T44
10471277CPR:3YUPREGULATORArg1:T26Arg2:T44
10471277CPR:6YANTAGONISTArg1:T6Arg2:T34
10471277CPR:6YANTAGONISTArg1:T7Arg2:T34
10471277CPR:2NDIRECT-REGULATORArg1:T5Arg2:T34
10471277CPR:2NDIRECT-REGULATORArg1:T2Arg2:T34
10471277CPR:2NDIRECT-REGULATORArg1:T17Arg2:T36
10471277CPR:2NDIRECT-REGULATORArg1:T3Arg2:T34

CHEMPROT task Gold Standard data and predictions

  • File: chemprot_sample_gold_standard.tsv
PMIDCPR GroupArg1Arg2
10471277CPR:5Arg1:T1Arg2:T37
10471277CPR:5Arg1:T1Arg2:T39
10471277CPR:5Arg1:T21Arg2:T38
10471277CPR:5Arg1:T22Arg2:T38
10471277CPR:3Arg1:T25Arg2:T42
10471277CPR:3Arg1:T26Arg2:T42
10471277CPR:3Arg1:T25Arg2:T44
10471277CPR:3Arg1:T26Arg2:T44
10471277CPR:6Arg1:T6Arg2:T34
10471277CPR:6Arg1:T7Arg2:T34
  • File: chemprot_sample_predictions.tsv
PMIDCPR GroupArg1Arg2
10471277CPR:5Arg1:T17Arg2:T36
10471277CPR:5Arg1:T1Arg2:T37
10471277CPR:5Arg1:T1Arg2:T39
10471277CPR:5Arg1:T20Arg2:T38
10471277CPR:5Arg1:T21Arg2:T38
10471277CPR:5Arg1:T22Arg2:T38
10471277CPR:5Arg1:T31Arg2:T40
10471277CPR:5Arg1:T32Arg2:T40
10471277CPR:6Arg1:T6Arg2:T34
10471277CPR:6Arg1:T7Arg2:T34

Knowledge Graph results

kG

Resources

BLUE, the Biomedical Language Understanding Evaluation benchmark

Chemical-protein Interaction Extraction via Gaussian Probability Distribution and External Biomedical Knowledge

Scibert

Model file of scibert_scivocab_uncased:

wget "https://s3-us-west-2.amazonaws.com/ai2-s2-research/scibert/pytorch_models/scibert_scivocab_uncased.tar"

info

dist info

SetPMIDPosNeg
train35001695648621
dev75037129690
test107500220315
NONE                      48197
INHIBITOR                  5326
DIRECT-REGULATOR           2153
SUBSTRATE                  1988
ACTIVATOR                  1381
INDIRECT-UPREGULATOR       1337
INDIRECT-DOWNREGULATOR     1316
ANTAGONIST                  937
PRODUCT-OF                  917
PART-OF                     877
AGONIST                     646
AGONIST-ACTIVATOR            28
SUBSTRATE_PRODUCT-OF         24
AGONIST-INHIBITOR            10
Name: relation, dtype: int64

NONE                      9683
INHIBITOR                 1136
SUBSTRATE                  494
DIRECT-REGULATOR           452
INDIRECT-DOWNREGULATOR     329
INDIRECT-UPREGULATOR       298
PART-OF                    254
ACTIVATOR                  234
ANTAGONIST                 214
PRODUCT-OF                 156
AGONIST                    126
AGONIST-ACTIVATOR           10
SUBSTRATE_PRODUCT-OF         3
AGONIST-INHIBITOR            2
Name: relation, dtype: int64

LM Results

Base Language ModelAnnotation StyleNegative Samples in Training SetF1 Score on Development Set
bert_base_uncasednoneno0.552
scibert_uncasedscibertno0.739
biobert_base_v1.1biobertno0.727
bert_base_uncasednoneyes0.348
scibert_uncasedscibertyes0.586
biobert_base_v1.1biobertyes0.581

New Results

Base Language ModelAnnotation StyleKnowledge Graph ModelF1 Score on Development Set
bert_base_uncasedrawnone0.765
bert_base_uncasedSciBertnone0.864
bert_base_uncasedBioBertnone0.863
scibert_scivocab_uncasedrawnone0.776
scibert_scivocab_uncasedSciBertnone0.876
scibert_scivocab_uncasedBioBertnone0.869
biobert-base-cased-v1.1rawnone0.776
biobert-base-cased-v1.1SciBertnone0.876
biobert-base-cased-v1.1BioBertnone0.875
bert_base_uncasedrawTransE0.768
bert_base_uncasedSciBertTransE0.858
bert_base_uncasedBioBertTransE0.856
scibert_scivocab_uncasedrawTransE0.776
scibert_scivocab_uncasedSciBertTransE0.872
scibert_scivocab_uncasedBioBertTransE0.868
biobert-base-cased-v1.1rawTransE0.778
biobert-base-cased-v1.1SciBertTransE0.873
biobert-base-cased-v1.1BioBertTransE0.863
bert_base_uncasedrawPairRE0.769
bert_base_uncasedSciBertPairRE0.858
bert_base_uncasedBioBertPairRE0.857
scibert_scivocab_uncasedrawPairRE0.779
scibert_scivocab_uncasedSciBertPairRE0.871
scibert_scivocab_uncasedBioBertPairRE0.867
biobert-base-cased-v1.1rawPairRE0.778
biobert-base-cased-v1.1SciBertPairRE0.872
biobert-base-cased-v1.1BioBertPairRE0.862

Code structure

drugprot_preprocess_data

2021-09-28 02:08:03,003 - INFO - model: bert-base-uncased, annotation: raw
2021-09-28 02:08:03,003 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:04,139 - INFO - f1 score: 0.765
2021-09-28 02:08:04,139 - INFO - model: bert-base-uncased, annotation: sci
2021-09-28 02:08:04,139 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:04,939 - INFO - f1 score: 0.864
2021-09-28 02:08:04,940 - INFO - model: bert-base-uncased, annotation: bio
2021-09-28 02:08:04,940 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:05,764 - INFO - f1 score: 0.863
2021-09-28 02:08:05,764 - INFO - model: allenai/scibert_scivocab_uncased, annotation: raw
2021-09-28 02:08:05,764 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:06,553 - INFO - f1 score: 0.776
2021-09-28 02:08:06,553 - INFO - model: allenai/scibert_scivocab_uncased, annotation: sci
2021-09-28 02:08:06,553 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:07,344 - INFO - f1 score: 0.876
2021-09-28 02:08:07,344 - INFO - model: allenai/scibert_scivocab_uncased, annotation: bio
2021-09-28 02:08:07,344 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:08,177 - INFO - f1 score: 0.869
2021-09-28 02:08:08,177 - INFO - model: dmis-lab/biobert-base-cased-v1.1, annotation: raw
2021-09-28 02:08:08,177 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:08,964 - INFO - f1 score: 0.776
2021-09-28 02:08:08,964 - INFO - model: dmis-lab/biobert-base-cased-v1.1, annotation: sci
2021-09-28 02:08:08,964 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:09,750 - INFO - f1 score: 0.876
2021-09-28 02:08:09,750 - INFO - model: dmis-lab/biobert-base-cased-v1.1, annotation: bio
2021-09-28 02:08:09,750 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:10,545 - INFO - f1 score: 0.875
2021-09-28 02:08:10,545 - INFO - kg_model: TransE, lm_model: bert-base-uncased, annotation: raw
2021-09-28 02:08:10,545 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:11,366 - INFO - f1 score: 0.768
2021-09-28 02:08:11,366 - INFO - kg_model: TransE, lm_model: bert-base-uncased, annotation: sci
2021-09-28 02:08:11,367 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:12,151 - INFO - f1 score: 0.858
2021-09-28 02:08:12,151 - INFO - kg_model: TransE, lm_model: bert-base-uncased, annotation: bio
2021-09-28 02:08:12,151 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:12,936 - INFO - f1 score: 0.856
2021-09-28 02:08:12,937 - INFO - kg_model: TransE, lm_model: allenai/scibert_scivocab_uncased, annotation: raw
2021-09-28 02:08:12,937 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:13,723 - INFO - f1 score: 0.776
2021-09-28 02:08:13,723 - INFO - kg_model: TransE, lm_model: allenai/scibert_scivocab_uncased, annotation: sci
2021-09-28 02:08:13,723 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:14,541 - INFO - f1 score: 0.872
2021-09-28 02:08:14,541 - INFO - kg_model: TransE, lm_model: allenai/scibert_scivocab_uncased, annotation: bio
2021-09-28 02:08:14,541 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:15,326 - INFO - f1 score: 0.868
2021-09-28 02:08:15,326 - INFO - kg_model: TransE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: raw
2021-09-28 02:08:15,326 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:16,110 - INFO - f1 score: 0.778
2021-09-28 02:08:16,110 - INFO - kg_model: TransE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: sci
2021-09-28 02:08:16,110 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:16,894 - INFO - f1 score: 0.873
2021-09-28 02:08:16,894 - INFO - kg_model: TransE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: bio
2021-09-28 02:08:16,894 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:17,712 - INFO - f1 score: 0.863
2021-09-28 02:08:17,712 - INFO - kg_model: PairRE, lm_model: bert-base-uncased, annotation: raw
2021-09-28 02:08:17,712 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:18,495 - INFO - f1 score: 0.769
2021-09-28 02:08:18,495 - INFO - kg_model: PairRE, lm_model: bert-base-uncased, annotation: sci
2021-09-28 02:08:18,495 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:19,280 - INFO - f1 score: 0.858
2021-09-28 02:08:19,280 - INFO - kg_model: PairRE, lm_model: bert-base-uncased, annotation: bio
2021-09-28 02:08:19,280 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:20,099 - INFO - f1 score: 0.857
2021-09-28 02:08:20,099 - INFO - kg_model: PairRE, lm_model: allenai/scibert_scivocab_uncased, annotation: raw
2021-09-28 02:08:20,099 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:20,883 - INFO - f1 score: 0.779
2021-09-28 02:08:20,883 - INFO - kg_model: PairRE, lm_model: allenai/scibert_scivocab_uncased, annotation: sci
2021-09-28 02:08:20,883 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:21,667 - INFO - f1 score: 0.871
2021-09-28 02:08:21,667 - INFO - kg_model: PairRE, lm_model: allenai/scibert_scivocab_uncased, annotation: bio
2021-09-28 02:08:21,667 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:22,450 - INFO - f1 score: 0.867
2021-09-28 02:08:22,450 - INFO - kg_model: PairRE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: raw
2021-09-28 02:08:22,450 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:23,268 - INFO - f1 score: 0.778
2021-09-28 02:08:23,268 - INFO - kg_model: PairRE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: sci
2021-09-28 02:08:23,268 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:24,051 - INFO - f1 score: 0.872
2021-09-28 02:08:24,052 - INFO - kg_model: PairRE, lm_model: dmis-lab/biobert-base-cased-v1.1, annotation: bio
2021-09-28 02:08:24,052 - INFO - Loading file from data/drugprot_preprocessed/bin/development
2021-09-28 02:08:24,841 - INFO - f1 score: 0.862

Contributors

wenyuan-wu

74 commits

Languages

Jupyter Notebook

56.2%

Python

43.6%