Gradient similarity. It is equal to Virus with $lambda$=0. The gradient of this data should resemble the original mixing harmful data.
Guardrail jailbreak. It is equal to Virus with $lambda$=1. This dataset should not be detected as harmful by the llama guard2 model.
Virus. It is equal to Virus with $lambda$=0.1. This dataset is produced by dual goal optimization, such that i) its gradient resembles the orginal harmful gradient, ii) it can bypass llama guard2 detection.
Gradient similarity. It is equal to Virus with $lambda$=0. The gradient of this data should resemble the original mixing harmful data.
Guardrail jailbreak. It is equal to Virus with $lambda$=1. This dataset should not be detected as harmful by the llama guard2 model.
Virus. It is equal to Virus with $lambda$=0.1. This dataset is produced by dual goal optimization, such that i) its gradient resembles the orginal harmful gradient, ii) it can bypass llama guard2 detection.