The OpenGenome dataset is organized in 2 stages, where stage 1 has context length 8k and stage 2 has context length 131k. Each stage has their own datasplits.
- stage1
- train
- validation
- test
- stage2
- train
- validation
- test
You can load a dataset using HF's API, with an example below.
from datasets import load_dataset
stage1_data = load_dataset("LongSafari/open-genome", 'stage1')
# access just the train data
stage_1_train_data = stage1_data['train']
Note: stage 1 training dataset is sharded into separate files due to it's large size.
We also provide a small dataset sample to test out the pipeline if you prefer.
sample_data = load_dataset("LongSafari/open-genome", 'sample')['validation']
The OpenGenome dataset is organized in 2 stages, where stage 1 has context length 8k and stage 2 has context length 131k. Each stage has their own datasplits.
- stage1
- train
- validation
- test
- stage2
- train
- validation
- test
You can load a dataset using HF's API, with an example below.
from datasets import load_dataset
stage1_data = load_dataset("LongSafari/open-genome", 'stage1')
# access just the train data
stage_1_train_data = stage1_data['train']
Note: stage 1 training dataset is sharded into separate files due to it's large size.
We also provide a small dataset sample to test out the pipeline if you prefer.
sample_data = load_dataset("LongSafari/open-genome", 'sample')['validation']