WikiMIA-25 is an evaluation-only dataset used to assess membership inference attacks (MIA) on language models, with a particular focus on recent models.
The dataset follows the dataset construction methodology introduced in WikiMIA-24, which includes newer non-member data to support evaluation under more recent training cutoff assumptions.
Each example contains the following fields:
input (string): Wikipedia text contentlabel (int64): Membership label
1: member0: non-member| Split | Description | Number of Examples |
|---|---|---|
| WikiMIA_length32 | Approximately 32 tokens per instance | 1518 |
| WikiMIA_length64 | Approximately 64 tokens per instance | 1750 |
| WikiMIA_length128 | Approximately 128 tokens per instance | 1058 |
| paper_subset | Balanced subset sampled from WikiMIA_length32 | 240 |
All splits remain balanced in terms of label distribution.
Specifically, paper_subset is the subset to reproduce results reported in our paper.
Non-member samples in WikiMIA-25 are selected from March 2025 – September 2025, therefore WikiMIA-25 can safely evaluate language models language models with a knowledge cutoff earlier than March 2025.
To load the dataset:
from datasets import load_dataset
block_size = 64
dataset = load_dataset("SimMIA/WikiMIA-25", split=f"WikiMIA_length{block_size}")
This dataset is intended only for evaluation, including:
It is not intended for training or fine-tuning language models.
This dataset is intended to study privacy risks in language models. All content is sourced from publicly available Wikipedia articles and does not include private or personally identifiable information.
If you use this dataset, please cite our paper:
@misc{yi2026membership,
title={Membership Inference on LLMs in the Wild},
author={Jiatong Yi and Yanyang Li},
year={2026},
eprint={2601.11314},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.11314},
}
Our dataset is constructed based on the WikiMIA-24 repo, please also cite:
@inproceedings{fu2024membership,
title={{MIA}-Tuner: Adapting Large Language Models as Pre-training Text Detector},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
author = {Fu, Wenjie and Wang, Huandong and Gao, Chen and Liu, Guanghua and Li, Yong and Jiang, Tao},
year = {2025},
address = {Philadelphia, Pennsylvania, USA}
}
Distributed under the MIT License.
WikiMIA-25 is an evaluation-only dataset used to assess membership inference attacks (MIA) on language models, with a particular focus on recent models.
The dataset follows the dataset construction methodology introduced in WikiMIA-24, which includes newer non-member data to support evaluation under more recent training cutoff assumptions.
Each example contains the following fields:
input (string): Wikipedia text contentlabel (int64): Membership label
1: member0: non-member| Split | Description | Number of Examples |
|---|---|---|
| WikiMIA_length32 | Approximately 32 tokens per instance | 1518 |
| WikiMIA_length64 | Approximately 64 tokens per instance | 1750 |
| WikiMIA_length128 | Approximately 128 tokens per instance | 1058 |
| paper_subset | Balanced subset sampled from WikiMIA_length32 | 240 |
All splits remain balanced in terms of label distribution.
Specifically, paper_subset is the subset to reproduce results reported in our paper.
Non-member samples in WikiMIA-25 are selected from March 2025 – September 2025, therefore WikiMIA-25 can safely evaluate language models language models with a knowledge cutoff earlier than March 2025.
To load the dataset:
from datasets import load_dataset
block_size = 64
dataset = load_dataset("SimMIA/WikiMIA-25", split=f"WikiMIA_length{block_size}")
This dataset is intended only for evaluation, including:
It is not intended for training or fine-tuning language models.
This dataset is intended to study privacy risks in language models. All content is sourced from publicly available Wikipedia articles and does not include private or personally identifiable information.
If you use this dataset, please cite our paper:
@misc{yi2026membership,
title={Membership Inference on LLMs in the Wild},
author={Jiatong Yi and Yanyang Li},
year={2026},
eprint={2601.11314},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.11314},
}
Our dataset is constructed based on the WikiMIA-24 repo, please also cite:
@inproceedings{fu2024membership,
title={{MIA}-Tuner: Adapting Large Language Models as Pre-training Text Detector},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
author = {Fu, Wenjie and Wang, Huandong and Gao, Chen and Liu, Guanghua and Li, Yong and Jiang, Tao},
year = {2025},
address = {Philadelphia, Pennsylvania, USA}
}
Distributed under the MIT License.