The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
We note that this is an updated version of this benchmark after the paper has been through peer-review. We highly encourage to move to this version in detriment of the older version.
Keypoints about the updated datasets:
The different datasets are collected from the following resources:
H3K4me3 (ENCFF706WUF), H3K27ac (ENCFF544LXB), H3K27me3 (ENCFF323WOT), H3K4me1 (ENCFF135ZLM), H3K36me3 (ENCFF561OUZ), H3K9me3 (ENCFF963GZJ), H3K9ac (ENCFF891CHI), H3K4me2 (ENCFF749KLQ), H4K20me1 (ENCFF909RKY), H2AFZ (ENCFF213OTI). For each dataset, we selected 1kb genomic sequences containing peaks as positive examples and all 1kb sequences not overlapping peaks as negative examples.Enhancer) and a multi-label prediction task with labels being tissue-specific enhancer, tissue-invariant enhancer or none (Enhancer types).Promoter all), a promoter with a TATA-box motif (Promoter TATA) or a promoter without a TATA-box motif (Promoter no-TATA).| Task | Number of train sequences | Number of test sequences | Number of labels | Sequence length |
| --------------------- | ------------------------- | ------------------------ | ---------------- | --------------- |
| promoter_all | 30,000 | 1,584 | 2 | 300 |
| promoter_tata | 5,062 | 212 | 2 | 300 |
| promoter_no_tata | 30,000 | 1,372 | 2 | 300 |
| enhancers | 30,000 | 3,000 | 2 | 400 |
| enhancers_types | 30,000 | 3,000 | 3 | 400 |
| splice_sites_all | 30,000 | 3,000 | 3 | 600 |
| splice_sites_acceptor | 30,000 | 3,000 | 2 | 600 |
| splice_sites_donor | 30,000 | 3,000 | 2 | 600 |
| H2AFZ | 30,000 | 3,000 | 2 | 1,000 |
| H3K27ac | 30,000 | 1,616 | 2 | 1,000 |
| H3K27me3 | 30,000 | 3,000 | 2 | 1,000 |
| H3K36me3 | 30,000 | 3,000 | 2 | 1,000 |
| H3K4me1 | 30,000 | 3,000 | 2 | 1,000 |
| H3K4me2 | 30,000 | 2,138 | 2 | 1,000 |
| H3K4me3 | 30,000 | 776 | 2 | 1,000 |
| H3K9ac | 23,274 | 1,004 | 2 | 1,000 |
| H3K9me3 | 27,438 | 850 | 2 | 1,000 |
| H4K20me1 | 30,000 | 2,270 | 2 | 1,000 |
from datasets import load_dataset
# Import the dataset
dataset = load_dataset("InstaDeepAI/nucleotide_transformer_downstream_tasks_revised")
# It loads all the tasks
print(dataset)
# You can filter for a given task, if needed
enhancers_dataset = dataset.filter(lambda example: example["task"] == "enhancers")
print(enhancers_dataset)
The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
We note that this is an updated version of this benchmark after the paper has been through peer-review. We highly encourage to move to this version in detriment of the older version.
Keypoints about the updated datasets:
The different datasets are collected from the following resources:
H3K4me3 (ENCFF706WUF), H3K27ac (ENCFF544LXB), H3K27me3 (ENCFF323WOT), H3K4me1 (ENCFF135ZLM), H3K36me3 (ENCFF561OUZ), H3K9me3 (ENCFF963GZJ), H3K9ac (ENCFF891CHI), H3K4me2 (ENCFF749KLQ), H4K20me1 (ENCFF909RKY), H2AFZ (ENCFF213OTI). For each dataset, we selected 1kb genomic sequences containing peaks as positive examples and all 1kb sequences not overlapping peaks as negative examples.Enhancer) and a multi-label prediction task with labels being tissue-specific enhancer, tissue-invariant enhancer or none (Enhancer types).Promoter all), a promoter with a TATA-box motif (Promoter TATA) or a promoter without a TATA-box motif (Promoter no-TATA).| Task | Number of train sequences | Number of test sequences | Number of labels | Sequence length |
| --------------------- | ------------------------- | ------------------------ | ---------------- | --------------- |
| promoter_all | 30,000 | 1,584 | 2 | 300 |
| promoter_tata | 5,062 | 212 | 2 | 300 |
| promoter_no_tata | 30,000 | 1,372 | 2 | 300 |
| enhancers | 30,000 | 3,000 | 2 | 400 |
| enhancers_types | 30,000 | 3,000 | 3 | 400 |
| splice_sites_all | 30,000 | 3,000 | 3 | 600 |
| splice_sites_acceptor | 30,000 | 3,000 | 2 | 600 |
| splice_sites_donor | 30,000 | 3,000 | 2 | 600 |
| H2AFZ | 30,000 | 3,000 | 2 | 1,000 |
| H3K27ac | 30,000 | 1,616 | 2 | 1,000 |
| H3K27me3 | 30,000 | 3,000 | 2 | 1,000 |
| H3K36me3 | 30,000 | 3,000 | 2 | 1,000 |
| H3K4me1 | 30,000 | 3,000 | 2 | 1,000 |
| H3K4me2 | 30,000 | 2,138 | 2 | 1,000 |
| H3K4me3 | 30,000 | 776 | 2 | 1,000 |
| H3K9ac | 23,274 | 1,004 | 2 | 1,000 |
| H3K9me3 | 27,438 | 850 | 2 | 1,000 |
| H4K20me1 | 30,000 | 2,270 | 2 | 1,000 |
from datasets import load_dataset
# Import the dataset
dataset = load_dataset("InstaDeepAI/nucleotide_transformer_downstream_tasks_revised")
# It loads all the tasks
print(dataset)
# You can filter for a given task, if needed
enhancers_dataset = dataset.filter(lambda example: example["task"] == "enhancers")
print(enhancers_dataset)