LongSafari/open-genome

Dataset

Dataset organization

31

14 commits

1 linked in READMEs

updated Jul 10, 2024

See the code

README

Dataset organization

The OpenGenome dataset is organized in 2 stages, where stage 1 has context length 8k and stage 2 has context length 131k. Each stage has their own datasplits.

- stage1
  - train
  - validation
  - test

- stage2
  - train
  - validation
  - test

Instructions to download

You can load a dataset using HF's API, with an example below.

from datasets import load_dataset

stage1_data = load_dataset("LongSafari/open-genome", 'stage1')

# access just the train data
stage_1_train_data = stage1_data['train']

Note: stage 1 training dataset is sharded into separate files due to it's large size.

We also provide a small dataset sample to test out the pipeline if you prefer.

sample_data = load_dataset("LongSafari/open-genome", 'sample')['validation']

biology
deep signal processing
genomics
hybrid
long context
stripedhyena

LongSafari/open-genome

Dataset

Dataset organization

31

14 commits

1 linked in READMEs

updated Jul 10, 2024

See the code

README

Dataset organization

The OpenGenome dataset is organized in 2 stages, where stage 1 has context length 8k and stage 2 has context length 131k. Each stage has their own datasplits.

- stage1
  - train
  - validation
  - test

- stage2
  - train
  - validation
  - test

Instructions to download

You can load a dataset using HF's API, with an example below.

from datasets import load_dataset

stage1_data = load_dataset("LongSafari/open-genome", 'stage1')

# access just the train data
stage_1_train_data = stage1_data['train']

Note: stage 1 training dataset is sharded into separate files due to it's large size.

We also provide a small dataset sample to test out the pipeline if you prefer.

sample_data = load_dataset("LongSafari/open-genome", 'sample')['validation']

biology
deep signal processing
genomics
hybrid
long context
stripedhyena