swj0419/WikiMIA

Dataset

πŸ“˜ WikiMIA Datasets

20

37 commits

2 linked in READMEs

updated Nov 3, 2023

See the code

README

πŸ“˜ WikiMIA Datasets

The WikiMIA datasets serve as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from extensive large language models.

πŸ“Œ Applicability

The datasets can be applied to various models released between 2017 to 2023:

  • LLaMA1/2
  • GPT-Neo
  • OPT
  • Pythia
  • text-davinci-001
  • text-davinci-002
  • ... and more.

Loading the datasets

To load the dataset:

from datasets import load_dataset

LENGTH = 64
dataset = load_dataset("swj0419/WikiMIA", split=f"WikiMIA_length{LENGTH}")
  • Available Text Lengths: 32, 64, 128, 256.
  • Label 0: Refers to the unseen data during pretraining. Label 1: Refers to the seen data.

πŸ› οΈ Codebase

For evaluating MIA methods on our datasets, visit our GitHub repository.

⭐ Citing our Work

If you find our codebase and datasets beneficial, kindly cite our work:

@misc{shi2023detecting,
    title={Detecting Pretraining Data from Large Language Models},
    author={Weijia Shi and Anirudh Ajith and Mengzhou Xia and Yangsibo Huang and Daogao Liu and Terra Blevins and Danqi Chen and Luke Zettlemoyer},
    year={2023},
    eprint={2310.16789},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

swj0419/WikiMIA

Dataset

πŸ“˜ WikiMIA Datasets

20

37 commits

2 linked in READMEs

updated Nov 3, 2023

See the code

README

πŸ“˜ WikiMIA Datasets

The WikiMIA datasets serve as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from extensive large language models.

πŸ“Œ Applicability

The datasets can be applied to various models released between 2017 to 2023:

  • LLaMA1/2
  • GPT-Neo
  • OPT
  • Pythia
  • text-davinci-001
  • text-davinci-002
  • ... and more.

Loading the datasets

To load the dataset:

from datasets import load_dataset

LENGTH = 64
dataset = load_dataset("swj0419/WikiMIA", split=f"WikiMIA_length{LENGTH}")
  • Available Text Lengths: 32, 64, 128, 256.
  • Label 0: Refers to the unseen data during pretraining. Label 1: Refers to the seen data.

πŸ› οΈ Codebase

For evaluating MIA methods on our datasets, visit our GitHub repository.

⭐ Citing our Work

If you find our codebase and datasets beneficial, kindly cite our work:

@misc{shi2023detecting,
    title={Detecting Pretraining Data from Large Language Models},
    author={Weijia Shi and Anirudh Ajith and Mengzhou Xia and Yangsibo Huang and Daogao Liu and Terra Blevins and Danqi Chen and Luke Zettlemoyer},
    year={2023},
    eprint={2310.16789},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}