EdPuth/LLMs-based-Fuzzer-Survey

This repo list the core literature in the field of fuzzing test, large language model, and LLM-based fuzzer. Most of papers are selected from authoritative platform such as google schlor, and was published recently. It will be helpful for the researchers who wants to develop LLMs-based fuzzer. Feel free to send a pull request.

55

36 commits

updated Feb 20, 2024

See the code

README

LLMs-based-Fuzzer-Survey

Our survey paper : https://arxiv.org/abs/2402.00350

Survey Paper Update Logs

  • 2024.02.07 - Paper v2 released: More figures and contents added.
  • 2024.02.01 - Paper v1 released: Initial version.

This repo list the core literature in the field of fuzzing tests, large language models, and LLM-based fuzzer. Most of the papers are selected from authoritative platforms such as Google Scholar, and were published recently. It will be helpful for researchers who want to develop LLMs-based fuzzer.

Feel free to send a pull request!!!

1. Figures

A General overview of LLMs-based Fuzzer

2. Papers

2.1 14 Papers of LLM based fuzzer
  1. Fuzzing-based hard-label black-box attacks against machine learning models [pdf]

  2. Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models[pdf]

  3. Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT [pdf]

  4. ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP [pdf]

  5. Large Language Models for Fuzzing Parsers (Registered Report) [pdf]

  6. Understanding Large Language Model Based Fuzz Driver Generation [pdf]

  7. Fuzz4All: Universal Fuzzing with Large Language Models [pdf]

  8. White-box Compiler Fuzzing Empowered by Large Language Models [pdf]

  9. AI-Powered Fuzzing: Breaking the Bug Hunting Barrier [Web]

  10. Large Language Model guided Protocol Fuzzing [pdf]

  11. Testing the Limits: Unusual Text Inputs Generation for Mobile App Crash Detection with Large Language Model [pdf]

  12. Smart Fuzzing of 5G Wireless Software Implementation [pdf]

  13. Augmenting Greybox Fuzzing with Generative AI [pdf]

  14. CHEMFUZZ: Large Language Models-assisted Fuzzing for Quantum Chemistry Software Bug Detection [pdf]

2.2 Reference List used in the survey paper
  1. Large Language Models for Fuzzing Parsers (Registered Report) [pdf]
  2. Claude-2 [web]
  3. Coverage-based Greybox Fuzzing as Markov Chain [pdf]
  4. Directed Greybox Fuzzing [pdf]
  5. Language Models are Few-Shot Learners [pdf]
  6. A systematic review of fuzzing techniques [pdf]
  7. Evaluating Large Language Models Trained on Code [pdf]
  8. Fuzzing Deep-Learning Libraries via Automated Relational API Inference [pdf]
  9. Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models [pdf]
  10. Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT [pdf]
  11. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding [pdf]
  12. InCoder: A Generative Model for Code Infilling and Synthesis [pdf]
  13. Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder [pdf]
  14. AI-Powered Fuzzing: Breaking the Bug Hunting Barrier [web]
  15. GPF vdalabs [web]
  16. Muffin: Testing Deep Learning Libraries via Neural Architecture Fuzzing [pdf]
  17. Augmenting Greybox Fuzzing with Generative AI [pdf]
  18. BertRLFuzzer: A BERT and Reinforcement Learning Based Fuzzer [pdf]
  19. Challenges and Applications of Large Language Models [pdf]
  20. Evaluating Fuzz Testing [pdf]
  21. FairFuzz: a targeted mutation strategy for increasing greybox fuzz testing coverage [pdf]
  22. Fuzzing: State of the Art [pdf]
  23. Identifying Insufficient Data Coverage in Databases with Multiple Relations [pdf]
  24. Testing the Limits: Unusual Text Inputs Generation for Mobile App Crash Detection with Large Language Model [pdf]
  25. Demystify the Fuzzing Methods: A Comprehensive Survey [pdf]
  26. The Art, Science, and Engineering of Fuzzing: A Survey [pdf]
  27. Large Language Model guided Protocol Fuzzing [pdf]
  28. A Comprehensive Overview of Large Language Models [pdf]
  29. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis [pdf]
  30. Language Models as Knowledge Bases? [pdf]
  31. AFLNET: A Greybox Fuzzer for Network Protocols [pdf]
  32. NSFuzz: Towards Efficient and State-Aware Network Service Fuzzing [pdf]
  33. Chemfuzz: Large language models-assisted fuzzing for quantum chemistry software bug detection [pdf]
  34. Improving Language Understanding by Generative Pre-Training [pdf]
  35. honggfuzz [web]
  36. StarCoder: A State-of-the-Art LLM for Code [pdf]
  37. Superion: Grammar-Aware Greybox Fuzzing [pdf]
  38. Deep learning library testing via effective model generation [pdf]
  39. Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open Source [pdf]
  40. Smart Fuzzing of 5G Wireless Software Implementation [pdf]
  41. Fuzz4All: Universal Fuzzing with Large Language Models [pdf]
  42. ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP [pdf]
  43. White-box Compiler Fuzzing Empowered by Large Language Models [pdf]
  44. american fuzzy lop [pdf]
  45. Understanding Large Language Model Based Fuzz Driver Generation [pdf]

EdPuth/LLMs-based-Fuzzer-Survey

This repo list the core literature in the field of fuzzing test, large language model, and LLM-based fuzzer. Most of papers are selected from authoritative platform such as google schlor, and was published recently. It will be helpful for the researchers who wants to develop LLMs-based fuzzer. Feel free to send a pull request.

55

36 commits

updated Feb 20, 2024

See the code

README

LLMs-based-Fuzzer-Survey

Our survey paper : https://arxiv.org/abs/2402.00350

Survey Paper Update Logs

  • 2024.02.07 - Paper v2 released: More figures and contents added.
  • 2024.02.01 - Paper v1 released: Initial version.

This repo list the core literature in the field of fuzzing tests, large language models, and LLM-based fuzzer. Most of the papers are selected from authoritative platforms such as Google Scholar, and were published recently. It will be helpful for researchers who want to develop LLMs-based fuzzer.

Feel free to send a pull request!!!

1. Figures

A General overview of LLMs-based Fuzzer

2. Papers

2.1 14 Papers of LLM based fuzzer
  1. Fuzzing-based hard-label black-box attacks against machine learning models [pdf]

  2. Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models[pdf]

  3. Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT [pdf]

  4. ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP [pdf]

  5. Large Language Models for Fuzzing Parsers (Registered Report) [pdf]

  6. Understanding Large Language Model Based Fuzz Driver Generation [pdf]

  7. Fuzz4All: Universal Fuzzing with Large Language Models [pdf]

  8. White-box Compiler Fuzzing Empowered by Large Language Models [pdf]

  9. AI-Powered Fuzzing: Breaking the Bug Hunting Barrier [Web]

  10. Large Language Model guided Protocol Fuzzing [pdf]

  11. Testing the Limits: Unusual Text Inputs Generation for Mobile App Crash Detection with Large Language Model [pdf]

  12. Smart Fuzzing of 5G Wireless Software Implementation [pdf]

  13. Augmenting Greybox Fuzzing with Generative AI [pdf]

  14. CHEMFUZZ: Large Language Models-assisted Fuzzing for Quantum Chemistry Software Bug Detection [pdf]

2.2 Reference List used in the survey paper
  1. Large Language Models for Fuzzing Parsers (Registered Report) [pdf]
  2. Claude-2 [web]
  3. Coverage-based Greybox Fuzzing as Markov Chain [pdf]
  4. Directed Greybox Fuzzing [pdf]
  5. Language Models are Few-Shot Learners [pdf]
  6. A systematic review of fuzzing techniques [pdf]
  7. Evaluating Large Language Models Trained on Code [pdf]
  8. Fuzzing Deep-Learning Libraries via Automated Relational API Inference [pdf]
  9. Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models [pdf]
  10. Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT [pdf]
  11. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding [pdf]
  12. InCoder: A Generative Model for Code Infilling and Synthesis [pdf]
  13. Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder [pdf]
  14. AI-Powered Fuzzing: Breaking the Bug Hunting Barrier [web]
  15. GPF vdalabs [web]
  16. Muffin: Testing Deep Learning Libraries via Neural Architecture Fuzzing [pdf]
  17. Augmenting Greybox Fuzzing with Generative AI [pdf]
  18. BertRLFuzzer: A BERT and Reinforcement Learning Based Fuzzer [pdf]
  19. Challenges and Applications of Large Language Models [pdf]
  20. Evaluating Fuzz Testing [pdf]
  21. FairFuzz: a targeted mutation strategy for increasing greybox fuzz testing coverage [pdf]
  22. Fuzzing: State of the Art [pdf]
  23. Identifying Insufficient Data Coverage in Databases with Multiple Relations [pdf]
  24. Testing the Limits: Unusual Text Inputs Generation for Mobile App Crash Detection with Large Language Model [pdf]
  25. Demystify the Fuzzing Methods: A Comprehensive Survey [pdf]
  26. The Art, Science, and Engineering of Fuzzing: A Survey [pdf]
  27. Large Language Model guided Protocol Fuzzing [pdf]
  28. A Comprehensive Overview of Large Language Models [pdf]
  29. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis [pdf]
  30. Language Models as Knowledge Bases? [pdf]
  31. AFLNET: A Greybox Fuzzer for Network Protocols [pdf]
  32. NSFuzz: Towards Efficient and State-Aware Network Service Fuzzing [pdf]
  33. Chemfuzz: Large language models-assisted fuzzing for quantum chemistry software bug detection [pdf]
  34. Improving Language Understanding by Generative Pre-Training [pdf]
  35. honggfuzz [web]
  36. StarCoder: A State-of-the-Art LLM for Code [pdf]
  37. Superion: Grammar-Aware Greybox Fuzzing [pdf]
  38. Deep learning library testing via effective model generation [pdf]
  39. Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open Source [pdf]
  40. Smart Fuzzing of 5G Wireless Software Implementation [pdf]
  41. Fuzz4All: Universal Fuzzing with Large Language Models [pdf]
  42. ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP [pdf]
  43. White-box Compiler Fuzzing Empowered by Large Language Models [pdf]
  44. american fuzzy lop [pdf]
  45. Understanding Large Language Model Based Fuzz Driver Generation [pdf]