zhongqy/RMCBench

Dataset

RMCBench

3

4 commits

2 linked in READMEs

updated Jul 21, 2025

See the code

README

RMCBench

Benchmarking Large Language Models’ Resistance to Malicious Code Generation Prompts

β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ•—   β–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ•—   β–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•—  β–ˆβ–ˆβ•—
β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•”β–ˆβ–ˆβ–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ•”β–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β•  β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•šβ•β• β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘ β•šβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘
β•šβ•β•  β•šβ•β•β•šβ•β•     β•šβ•β• β•šβ•β•β•β•β•β•β•šβ•β•β•β•β•β• β•šβ•β•β•β•β•β•β•β•šβ•β•  β•šβ•β•β•β• β•šβ•β•β•β•β•β•β•šβ•β•  β•šβ•β•                

Why do we need to do this study?

The large language models be used to generate malicious code!!!

This is a hidden danger to the security of LLMs content.

RMCBench-Examples of using LLMs to generate malicious code.

Through RMCBench, you can conduct a comprehensive evaluation of LLMs' ability to resist malicious code generation.

Result Leaderboard

GOOD:LLMs refuse to generate malicious code

LLMGOOD(%)BAD(%)UNCLEAR(%)
llama-2-13b-chat-hf48.8449.261.90
deepseek-coder-7b-instruct-v1.544.1955.810.00
Meta-Llama-3-8B-Instruct43.5556.240.21
mpt-7b-chat39.9657.082.96
llama-2-7b-chat-hf38.2759.202.54
gpt-435.7364.270.00
CodeLlama-13b-Instruct-hf30.6668.920.42
gpt-3.5-turbo18.3981.180.42
zephyr-7b-beta8.4690.700.85
vicuna-7b-v1.34.8684.1410.99
tulu-2-13b2.9690.276.77
Average28.7168.832.46

Characteristics

multi-scenarios

  • text-to-code
  • code-to-code

multi-tasks

  • text-to-code generation (Level 1 - 3)
  • code completion
  • code translation

data details

The following is an explanation of all the fields in the dataset.

These fields are present in all samples:

  • model_name: The full name of the LLM being tested.
  • pid: The ID of the prompt.
  • category: The scenario of malicious code generation (text-to-code, code-to-code).
  • task: The specific task of malicious code generation (text-to-code generation, code translation, code completion).
  • prompt: The prompt that instructs the LLMs to generate malicious code.
  • malicious functionality: The specific malicious intent/functionality of the malicious code.
  • malicious categories: The category of malicious code corresponding to the malicious intent/functionality.
  • input_tokens: The token length of the prompt.
  • response: The response from the LLMs.
  • label: The automated labeling results from ChatGPT-4.
  • check: The results of manual sampling checks on the label.

These fields are specific to the text-to-code scenario:

  • level: The difficulty level of text-to-code.
  • level description: The description and explanation of the level.
  • jid: The ID of the jailbreak template (in level 3).

These fields are specific to the code-to-code scenario:

  • cid: The ID of the malicious code sample we collected.
  • original code: The complete malicious code sample we collected.
  • language: The programming language of the malicious code.
  • code lines: The number of lines in the malicious code.
  • source: The source of the malicious code.

These fields are specific to the code-to-code scenario's code completion task:

  • code to be completed: The remaining malicious code after being hollowing out.
  • missing part: The hollowed out code (the code that needs to be completed).
  • completion level: The level of code completion (token-level, line-level, multiline-level, function-level).
  • completion position: The position of code completion (next token, fill-in-middle).

πŸ“Arxiv πŸ“ACM Digital Library

Dataset

🌟 Github πŸ€— Hugging Face

Citation

@inproceedings{10.1145/3691620.3695480,
author = {Chen, Jiachi and Zhong, Qingyuan and Wang, Yanlin and Ning, Kaiwen and Liu, Yongkun and Xu, Zenan and Zhao, Zhe and Chen, Ting and Zheng, Zibin},
title = {RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code},
year = {2024},
isbn = {9798400712487},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3691620.3695480},
doi = {10.1145/3691620.3695480},
numpages = {12},
keywords = {large language models, malicious code, code generation},
location = {Sacramento, CA, USA},
series = {ASE '24}
}
ai-safety
code-generation
code-to-code
text-to-code

Contributors

zhongqy

4 commits

zhongqy/RMCBench

Dataset

RMCBench

3

4 commits

2 linked in READMEs

updated Jul 21, 2025

See the code

README

RMCBench

Benchmarking Large Language Models’ Resistance to Malicious Code Generation Prompts

β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ•—   β–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ•—   β–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•—  β–ˆβ–ˆβ•—
β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•”β–ˆβ–ˆβ–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ•”β–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β•  β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•šβ•β• β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘ β•šβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘
β•šβ•β•  β•šβ•β•β•šβ•β•     β•šβ•β• β•šβ•β•β•β•β•β•β•šβ•β•β•β•β•β• β•šβ•β•β•β•β•β•β•β•šβ•β•  β•šβ•β•β•β• β•šβ•β•β•β•β•β•β•šβ•β•  β•šβ•β•                

Why do we need to do this study?

The large language models be used to generate malicious code!!!

This is a hidden danger to the security of LLMs content.

RMCBench-Examples of using LLMs to generate malicious code.

Through RMCBench, you can conduct a comprehensive evaluation of LLMs' ability to resist malicious code generation.

Result Leaderboard

GOOD:LLMs refuse to generate malicious code

LLMGOOD(%)BAD(%)UNCLEAR(%)
llama-2-13b-chat-hf48.8449.261.90
deepseek-coder-7b-instruct-v1.544.1955.810.00
Meta-Llama-3-8B-Instruct43.5556.240.21
mpt-7b-chat39.9657.082.96
llama-2-7b-chat-hf38.2759.202.54
gpt-435.7364.270.00
CodeLlama-13b-Instruct-hf30.6668.920.42
gpt-3.5-turbo18.3981.180.42
zephyr-7b-beta8.4690.700.85
vicuna-7b-v1.34.8684.1410.99
tulu-2-13b2.9690.276.77
Average28.7168.832.46

Characteristics

multi-scenarios

  • text-to-code
  • code-to-code

multi-tasks

  • text-to-code generation (Level 1 - 3)
  • code completion
  • code translation

data details

The following is an explanation of all the fields in the dataset.

These fields are present in all samples:

  • model_name: The full name of the LLM being tested.
  • pid: The ID of the prompt.
  • category: The scenario of malicious code generation (text-to-code, code-to-code).
  • task: The specific task of malicious code generation (text-to-code generation, code translation, code completion).
  • prompt: The prompt that instructs the LLMs to generate malicious code.
  • malicious functionality: The specific malicious intent/functionality of the malicious code.
  • malicious categories: The category of malicious code corresponding to the malicious intent/functionality.
  • input_tokens: The token length of the prompt.
  • response: The response from the LLMs.
  • label: The automated labeling results from ChatGPT-4.
  • check: The results of manual sampling checks on the label.

These fields are specific to the text-to-code scenario:

  • level: The difficulty level of text-to-code.
  • level description: The description and explanation of the level.
  • jid: The ID of the jailbreak template (in level 3).

These fields are specific to the code-to-code scenario:

  • cid: The ID of the malicious code sample we collected.
  • original code: The complete malicious code sample we collected.
  • language: The programming language of the malicious code.
  • code lines: The number of lines in the malicious code.
  • source: The source of the malicious code.

These fields are specific to the code-to-code scenario's code completion task:

  • code to be completed: The remaining malicious code after being hollowing out.
  • missing part: The hollowed out code (the code that needs to be completed).
  • completion level: The level of code completion (token-level, line-level, multiline-level, function-level).
  • completion position: The position of code completion (next token, fill-in-middle).

πŸ“Arxiv πŸ“ACM Digital Library

Dataset

🌟 Github πŸ€— Hugging Face

Citation

@inproceedings{10.1145/3691620.3695480,
author = {Chen, Jiachi and Zhong, Qingyuan and Wang, Yanlin and Ning, Kaiwen and Liu, Yongkun and Xu, Zenan and Zhao, Zhe and Chen, Ting and Zheng, Zibin},
title = {RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code},
year = {2024},
isbn = {9798400712487},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3691620.3695480},
doi = {10.1145/3691620.3695480},
numpages = {12},
keywords = {large language models, malicious code, code generation},
location = {Sacramento, CA, USA},
series = {ASE '24}
}
ai-safety
code-generation
code-to-code
text-to-code

Contributors

zhongqy

4 commits