A survey and reflection on the latest research breakthroughs in LLM-generated Text detection, including data, detectors, metrics, current issues and future directions.
See the code
The powerful ability of large language models (LLMs) to understand, follow, and generate complex languages has enabled LLM-generated texts to flood many areas of our daily lives at an incredible rate, with potentially negative impacts and risks on society and academia. As LLMs continue to expand, how can we detect LLM-generated texts to help minimize the threat posed by the misuse of LLMs?
¹ Junchao Wu, ¹ Shu Yang, ¹ Runzhe Zhan, ¹ ² Yulin Yuan, ¹ Derek Fai Wong, ¹ Lidia Sam Chao
¹ University of Macau, ² Peking University
A survey and reflection on the latest research breakthroughs in LLM-generated Text detection, including data, detectors, metrics, current issues and future directions. Please refer to our article/paper for more details.
| Benchmarks / Datasets | Venue | Date | Use | Human | LLMs |
|---|---|---|---|---|---|
| HC3 | arXiv | 2023-01 | train | 58k | 26k |
| HC3-Chinese | arXiv | 2023-01 | train | 22k | 17k |
| CHEAT | arXiv | 2023-04 | train | 15k | 35k |
| GROVER Dataset | NeurIPS 2019 | 2019-05 | train valid test | 5k 2k 8k | 5k 1k 4k |
| TweepFake | PLoS ONE | 2020-07 | train | 12k | 12k |
| GPT-2 Output Dataset | GitHub | - | train | 250k | 250k |
| TuringBench | EMNLP 2021 Findings | 2021 | train | 10k | 190k |
| MGTBench | arXiv | 2023-03 | train test | 2k 563 | 13k 3k |
| ArguGPT | arXiv | 2023-04 | train valid test | 3k 350 350 | 3k 350 350 |
| DeepfakeText-Dataset | ACL 2024 | 2023-05 | train valid test | 95k 29k 29k | 236k 29k 28k |
| M4 | arXiv | 2023-05 | train valid test | 122k 500 500 | 122k 500 500 |
| GPABenchmark | arXiv | 2023-06 | train | 600k | 600k |
| Scientific-articles Benchmark | TrustNLP 2023 | 2023 | train test | 8k 4k | 8k 4k |
| DetectRL | NeurIPS 2024 D&B | 2024-10 | train test | 101k | 134k (4 LLMs, 4 domains, 4 attacks) |
| DetectRL-X | ACL 2026 | 2026-05 | test | 3.46M (8 langs, 6 domains, 4 LLMs, 8 attacks) | |
| MULTITuDE | EMNLP 2023 | 2023-10 | test | - | 56k (7 langs, 8 LLMs) |
| RAID | ACL 2024 | 2024-05 | test | - | 6M (11 models) |
| M4GT-Bench | ACL 2024 | 2024-02 | test | - | - (4 domains, multi-lingual) |
| MultiSocial | ACL 2025 | 2025-07 | test | 58k | 414k (22 langs, 5 platforms, 7 LLMs) |
| Detecting the Machine | arXiv | 2026-03 | test | 23k | 15k (HC3 + ELI5, multiple LLMs) |
| Tasks | Datasets |
|---|---|
| Questions Answering | PubMedQA, Children book corpus (CBT), ELI5, TruthfulQA, NarrativeQA |
| Scientific writing | Peer Read, arXiv, TOEFL11 |
| Story generation | WritingPrompts |
| News Article writing | XSum |
| Web Text | Wiki40b, WebText, Avax tweets dataset, Climate Change Tweets Ids |
| Opinion statements | r/ChangeMyView (CMV) Reddit subcommunity, Yelp , IMDB Dataset |
| Comprehension and Reasoning | SciGen, ROCStories Corpora, HellaSwag, SQuAD |
If our research helps you, please kindly cite our paper.
@article{wu2025survey,
title={A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions},
author={Junchao Wu and Shu Yang and Runzhe Zhan and Yulin Yuan and Lidia Sam Chao and Derek Fai Wong},
journal = {Computational Linguistics},
volume = {51},
number = {1},
year = {2025},
pages = {275--338},
url = {https://aclanthology.org/2025.cl-1.8/},
}
@inproceedings{wu2025GECScore,
title={Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore},
author={Junchao Wu and Runzhe Zhan and Derek F. Wong and Shu Yang and Xuebo Liu and Lidia S. Chao and Min Zhang},
booktitle = {Proceedings of the 31st International Conference on Computational Linguistics},
year = {2025},
pages = {10275--10292},
url = {https://aclanthology.org/2025.coling-main.684/},
}
@article{chen2025RepreGuard,
title={RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns},
author={Xin Chen and Junchao Wu and Shu Yang and Runzhe Zhan and Zeyu Wu and Ziyang Luo and Di Wang and Min Yang and Lidia S. Chao and Derek F. Wong},
journal = {Transactions of the Association for Computational Linguistics},
volume = {13},
year = {2025},
pages = {1812--1831},
url = {https://aclanthology.org/2025.tacl-1.81/},
}
@inproceedings{wu2026DetectRLX,
title={DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection},
author={Junchao Wu and Yefeng Liu and Chenyu Zhu and Hao Zhang and Zeyu Wu and Tianqi Shi and Yichao Du and Longyue Wang and Weihua Luo and Jinsong Su and Derek F. Wong},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
year = {2026},
pages = {38247--38294},
url = {https://aclanthology.org/2026.acl-long.1773/},
}
@inproceedings{wu2024DetectRL,
title={DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios},
author={Junchao Wu and Runzhe Zhan and Derek F. Wong and Shu Yang and Xinyi Yang and Yulin Yuan and Lidia S. Chao},
booktitle = {Advances in Neural Information Processing Systems 37 (NeurIPS 2024) Datasets and Benchmarks Track},
year = {2024},
url = {https://proceedings.neurips.cc/paper_files/paper/2024/hash/b61bdf7e9f64c04ec75a26e781e2ad51-Abstract-Datasets_and_Benchmarks_Track.html},
}
Contributions are welcome! If you have any ideas, suggestions, or bug reports, please open an issue or submit a pull request. We appreciate your contributions to making LLM-generated Text Detection work even better.
A survey and reflection on the latest research breakthroughs in LLM-generated Text detection, including data, detectors, metrics, current issues and future directions.
See the code
The powerful ability of large language models (LLMs) to understand, follow, and generate complex languages has enabled LLM-generated texts to flood many areas of our daily lives at an incredible rate, with potentially negative impacts and risks on society and academia. As LLMs continue to expand, how can we detect LLM-generated texts to help minimize the threat posed by the misuse of LLMs?
¹ Junchao Wu, ¹ Shu Yang, ¹ Runzhe Zhan, ¹ ² Yulin Yuan, ¹ Derek Fai Wong, ¹ Lidia Sam Chao
¹ University of Macau, ² Peking University
A survey and reflection on the latest research breakthroughs in LLM-generated Text detection, including data, detectors, metrics, current issues and future directions. Please refer to our article/paper for more details.
| Benchmarks / Datasets | Venue | Date | Use | Human | LLMs |
|---|---|---|---|---|---|
| HC3 | arXiv | 2023-01 | train | 58k | 26k |
| HC3-Chinese | arXiv | 2023-01 | train | 22k | 17k |
| CHEAT | arXiv | 2023-04 | train | 15k | 35k |
| GROVER Dataset | NeurIPS 2019 | 2019-05 | train valid test | 5k 2k 8k | 5k 1k 4k |
| TweepFake | PLoS ONE | 2020-07 | train | 12k | 12k |
| GPT-2 Output Dataset | GitHub | - | train | 250k | 250k |
| TuringBench | EMNLP 2021 Findings | 2021 | train | 10k | 190k |
| MGTBench | arXiv | 2023-03 | train test | 2k 563 | 13k 3k |
| ArguGPT | arXiv | 2023-04 | train valid test | 3k 350 350 | 3k 350 350 |
| DeepfakeText-Dataset | ACL 2024 | 2023-05 | train valid test | 95k 29k 29k | 236k 29k 28k |
| M4 | arXiv | 2023-05 | train valid test | 122k 500 500 | 122k 500 500 |
| GPABenchmark | arXiv | 2023-06 | train | 600k | 600k |
| Scientific-articles Benchmark | TrustNLP 2023 | 2023 | train test | 8k 4k | 8k 4k |
| DetectRL | NeurIPS 2024 D&B | 2024-10 | train test | 101k | 134k (4 LLMs, 4 domains, 4 attacks) |
| DetectRL-X | ACL 2026 | 2026-05 | test | 3.46M (8 langs, 6 domains, 4 LLMs, 8 attacks) | |
| MULTITuDE | EMNLP 2023 | 2023-10 | test | - | 56k (7 langs, 8 LLMs) |
| RAID | ACL 2024 | 2024-05 | test | - | 6M (11 models) |
| M4GT-Bench | ACL 2024 | 2024-02 | test | - | - (4 domains, multi-lingual) |
| MultiSocial | ACL 2025 | 2025-07 | test | 58k | 414k (22 langs, 5 platforms, 7 LLMs) |
| Detecting the Machine | arXiv | 2026-03 | test | 23k | 15k (HC3 + ELI5, multiple LLMs) |
| Tasks | Datasets |
|---|---|
| Questions Answering | PubMedQA, Children book corpus (CBT), ELI5, TruthfulQA, NarrativeQA |
| Scientific writing | Peer Read, arXiv, TOEFL11 |
| Story generation | WritingPrompts |
| News Article writing | XSum |
| Web Text | Wiki40b, WebText, Avax tweets dataset, Climate Change Tweets Ids |
| Opinion statements | r/ChangeMyView (CMV) Reddit subcommunity, Yelp , IMDB Dataset |
| Comprehension and Reasoning | SciGen, ROCStories Corpora, HellaSwag, SQuAD |
If our research helps you, please kindly cite our paper.
@article{wu2025survey,
title={A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions},
author={Junchao Wu and Shu Yang and Runzhe Zhan and Yulin Yuan and Lidia Sam Chao and Derek Fai Wong},
journal = {Computational Linguistics},
volume = {51},
number = {1},
year = {2025},
pages = {275--338},
url = {https://aclanthology.org/2025.cl-1.8/},
}
@inproceedings{wu2025GECScore,
title={Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore},
author={Junchao Wu and Runzhe Zhan and Derek F. Wong and Shu Yang and Xuebo Liu and Lidia S. Chao and Min Zhang},
booktitle = {Proceedings of the 31st International Conference on Computational Linguistics},
year = {2025},
pages = {10275--10292},
url = {https://aclanthology.org/2025.coling-main.684/},
}
@article{chen2025RepreGuard,
title={RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns},
author={Xin Chen and Junchao Wu and Shu Yang and Runzhe Zhan and Zeyu Wu and Ziyang Luo and Di Wang and Min Yang and Lidia S. Chao and Derek F. Wong},
journal = {Transactions of the Association for Computational Linguistics},
volume = {13},
year = {2025},
pages = {1812--1831},
url = {https://aclanthology.org/2025.tacl-1.81/},
}
@inproceedings{wu2026DetectRLX,
title={DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection},
author={Junchao Wu and Yefeng Liu and Chenyu Zhu and Hao Zhang and Zeyu Wu and Tianqi Shi and Yichao Du and Longyue Wang and Weihua Luo and Jinsong Su and Derek F. Wong},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
year = {2026},
pages = {38247--38294},
url = {https://aclanthology.org/2026.acl-long.1773/},
}
@inproceedings{wu2024DetectRL,
title={DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios},
author={Junchao Wu and Runzhe Zhan and Derek F. Wong and Shu Yang and Xinyi Yang and Yulin Yuan and Lidia S. Chao},
booktitle = {Advances in Neural Information Processing Systems 37 (NeurIPS 2024) Datasets and Benchmarks Track},
year = {2024},
url = {https://proceedings.neurips.cc/paper_files/paper/2024/hash/b61bdf7e9f64c04ec75a26e781e2ad51-Abstract-Datasets_and_Benchmarks_Track.html},
}
Contributions are welcome! If you have any ideas, suggestions, or bug reports, please open an issue or submit a pull request. We appreciate your contributions to making LLM-generated Text Detection work even better.