Benchmarking Large Language Modelsβ Resistance to Malicious Code Generation Prompts
βββββββ ββββ ββββ ββββββββββββββ ββββββββββββ βββ ββββββββββ βββ
βββββββββββββ ββββββββββββββββββββββββββββββββββ ββββββββββββββ βββ
ββββββββββββββββββββββ ββββββββββββββ ββββββ ββββββ ββββββββ
ββββββββββββββββββββββ ββββββββββββββ βββββββββββββ ββββββββ
βββ ββββββ βββ ββββββββββββββββββββββββββββββ βββββββββββββββββ βββ
βββ ββββββ βββ ββββββββββββββ βββββββββββ βββββ ββββββββββ βββ
The large language models be used to generate malicious code!!!
This is a hidden danger to the security of LLMs content.
Through RMCBench, you can conduct a comprehensive evaluation of LLMs' ability to resist malicious code generation.
GOODοΌLLMs refuse to generate malicious code
| LLM | GOOD(%) | BAD(%) | UNCLEAR(%) |
|---|---|---|---|
| llama-2-13b-chat-hf | 48.84 | 49.26 | 1.90 |
| deepseek-coder-7b-instruct-v1.5 | 44.19 | 55.81 | 0.00 |
| Meta-Llama-3-8B-Instruct | 43.55 | 56.24 | 0.21 |
| mpt-7b-chat | 39.96 | 57.08 | 2.96 |
| llama-2-7b-chat-hf | 38.27 | 59.20 | 2.54 |
| gpt-4 | 35.73 | 64.27 | 0.00 |
| CodeLlama-13b-Instruct-hf | 30.66 | 68.92 | 0.42 |
| gpt-3.5-turbo | 18.39 | 81.18 | 0.42 |
| zephyr-7b-beta | 8.46 | 90.70 | 0.85 |
| vicuna-7b-v1.3 | 4.86 | 84.14 | 10.99 |
| tulu-2-13b | 2.96 | 90.27 | 6.77 |
| Average | 28.71 | 68.83 | 2.46 |
The following is an explanation of all the fields in the dataset.
πArxiv πACM Digital Library
π Github π€ Hugging Face
@inproceedings{10.1145/3691620.3695480,
author = {Chen, Jiachi and Zhong, Qingyuan and Wang, Yanlin and Ning, Kaiwen and Liu, Yongkun and Xu, Zenan and Zhao, Zhe and Chen, Ting and Zheng, Zibin},
title = {RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code},
year = {2024},
isbn = {9798400712487},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3691620.3695480},
doi = {10.1145/3691620.3695480},
numpages = {12},
keywords = {large language models, malicious code, code generation},
location = {Sacramento, CA, USA},
series = {ASE '24}
}
4 commits
Benchmarking Large Language Modelsβ Resistance to Malicious Code Generation Prompts
βββββββ ββββ ββββ ββββββββββββββ ββββββββββββ βββ ββββββββββ βββ
βββββββββββββ ββββββββββββββββββββββββββββββββββ ββββββββββββββ βββ
ββββββββββββββββββββββ ββββββββββββββ ββββββ ββββββ ββββββββ
ββββββββββββββββββββββ ββββββββββββββ βββββββββββββ ββββββββ
βββ ββββββ βββ ββββββββββββββββββββββββββββββ βββββββββββββββββ βββ
βββ ββββββ βββ ββββββββββββββ βββββββββββ βββββ ββββββββββ βββ
The large language models be used to generate malicious code!!!
This is a hidden danger to the security of LLMs content.
Through RMCBench, you can conduct a comprehensive evaluation of LLMs' ability to resist malicious code generation.
GOODοΌLLMs refuse to generate malicious code
| LLM | GOOD(%) | BAD(%) | UNCLEAR(%) |
|---|---|---|---|
| llama-2-13b-chat-hf | 48.84 | 49.26 | 1.90 |
| deepseek-coder-7b-instruct-v1.5 | 44.19 | 55.81 | 0.00 |
| Meta-Llama-3-8B-Instruct | 43.55 | 56.24 | 0.21 |
| mpt-7b-chat | 39.96 | 57.08 | 2.96 |
| llama-2-7b-chat-hf | 38.27 | 59.20 | 2.54 |
| gpt-4 | 35.73 | 64.27 | 0.00 |
| CodeLlama-13b-Instruct-hf | 30.66 | 68.92 | 0.42 |
| gpt-3.5-turbo | 18.39 | 81.18 | 0.42 |
| zephyr-7b-beta | 8.46 | 90.70 | 0.85 |
| vicuna-7b-v1.3 | 4.86 | 84.14 | 10.99 |
| tulu-2-13b | 2.96 | 90.27 | 6.77 |
| Average | 28.71 | 68.83 | 2.46 |
The following is an explanation of all the fields in the dataset.
πArxiv πACM Digital Library
π Github π€ Hugging Face
@inproceedings{10.1145/3691620.3695480,
author = {Chen, Jiachi and Zhong, Qingyuan and Wang, Yanlin and Ning, Kaiwen and Liu, Yongkun and Xu, Zenan and Zhao, Zhe and Chen, Ting and Zheng, Zibin},
title = {RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code},
year = {2024},
isbn = {9798400712487},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3691620.3695480},
doi = {10.1145/3691620.3695480},
numpages = {12},
keywords = {large language models, malicious code, code generation},
location = {Sacramento, CA, USA},
series = {ASE '24}
}
4 commits