anonymous4486/Virus

Dataset

2

stars

18

commits

1

linked in READMEs

Jan 18, 2025

updated

README

There are four datasets avaialble

  • Benign data. It is the original GSM8K data.
  • Gradient similarity. It is equal to Virus with $lambda$=0. The gradient of this data should resemble the original mixing harmful data.
  • Guardrail jailbreak. It is equal to Virus with $lambda$=1. This dataset should not be detected as harmful by the llama guard2 model.
  • Virus. It is equal to Virus with $lambda$=0.1. This dataset is produced by dual goal optimization, such that i) its gradient resembles the orginal harmful gradient, ii) it can bypass llama guard2 detection.

Contributors

anonymous4486

18 commits

anonymous4486/Virus

Dataset

2

stars

18

commits

1

linked in READMEs

Jan 18, 2025

updated

README

There are four datasets avaialble

  • Benign data. It is the original GSM8K data.
  • Gradient similarity. It is equal to Virus with $lambda$=0. The gradient of this data should resemble the original mixing harmful data.
  • Guardrail jailbreak. It is equal to Virus with $lambda$=1. This dataset should not be detected as harmful by the llama guard2 model.
  • Virus. It is equal to Virus with $lambda$=0.1. This dataset is produced by dual goal optimization, such that i) its gradient resembles the orginal harmful gradient, ii) it can bypass llama guard2 detection.

Contributors

anonymous4486

18 commits