This repository provides a cleaned dataset, which is intended to be used for text classification, language modeling, and AI-generated content detection tasks. The dataset covers various fields such as STEM, Social Sciences, and Humanities, and contains datasets from different categories, each of which has been processed and cleaned for easy use. Move to our codebase fro more information (github)
Official Repo for On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing
The dataset is organized into three major categories:
STEM (Science, Technology, Engineering, Mathematics)
Social Sciences
Humanities
The datasets were cleaned to ensure high-quality human content by implementing the following detailed steps:
Removal of Short and Incomplete Texts:
Truncation to Ensure Text Completeness:
Keyword-Based Filtering:
Format-Based Filtering:
&, $, ====, **, ##) were entirely removed if they appeared more than 50 times or in specific machine-generated formats like markdown tags.Consistency and Structure:
If you find this repo and dataset useful, please consider cite our work
@inproceedings{he2024mgtbench,
author = {He, Xinlei and Shen, Xinyue and Chen, Zeyuan and Backes, Michael and Zhang, Yang},
title = {{Mgtbench: Benchmarking machine-generated text detection}},
booktitle = {{ACM SIGSAC Conference on Computer and Communications Security (CCS)}},
pages = {},
publisher = {ACM},
year = {2024}
}
@misc{liu2025generalizationadaptationabilitymachinegenerated,
title={On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing},
author={Yule Liu and Zhiyuan Zhong and Yifan Liao and Zhen Sun and Jingyi Zheng and Jiaheng Wei and Qingyuan Gong and Fenghua Tong and Yang Chen and Yang Zhang and Xinlei He},
year={2025},
eprint={2412.17242},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2412.17242},
}
if you have any questions, please contact:
Xinlei He: xinleihe@hkust-gz.edu.cn
This repository provides a cleaned dataset, which is intended to be used for text classification, language modeling, and AI-generated content detection tasks. The dataset covers various fields such as STEM, Social Sciences, and Humanities, and contains datasets from different categories, each of which has been processed and cleaned for easy use. Move to our codebase fro more information (github)
Official Repo for On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing
The dataset is organized into three major categories:
STEM (Science, Technology, Engineering, Mathematics)
Social Sciences
Humanities
The datasets were cleaned to ensure high-quality human content by implementing the following detailed steps:
Removal of Short and Incomplete Texts:
Truncation to Ensure Text Completeness:
Keyword-Based Filtering:
Format-Based Filtering:
&, $, ====, **, ##) were entirely removed if they appeared more than 50 times or in specific machine-generated formats like markdown tags.Consistency and Structure:
If you find this repo and dataset useful, please consider cite our work
@inproceedings{he2024mgtbench,
author = {He, Xinlei and Shen, Xinyue and Chen, Zeyuan and Backes, Michael and Zhang, Yang},
title = {{Mgtbench: Benchmarking machine-generated text detection}},
booktitle = {{ACM SIGSAC Conference on Computer and Communications Security (CCS)}},
pages = {},
publisher = {ACM},
year = {2024}
}
@misc{liu2025generalizationadaptationabilitymachinegenerated,
title={On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing},
author={Yule Liu and Zhiyuan Zhong and Yifan Liao and Zhen Sun and Jingyi Zheng and Jiaheng Wei and Qingyuan Gong and Fenghua Tong and Yang Chen and Yang Zhang and Xinlei He},
year={2025},
eprint={2412.17242},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2412.17242},
}
if you have any questions, please contact:
Xinlei He: xinleihe@hkust-gz.edu.cn