This is a repository containing our method only.
We now support vLLM through customized model registered through plugin system. After installing vllm, run
cd vllm_steered_plugins
pip install -e .
You can use transferSteer2Safetensors.py to transform Probes + HF Model into a repo that can be loaded by vLLM (and also Transformers). We have uploaded jailbroken Circuit Breaker. You can try it with
vllm serve FTK11558/Llama-3-8B-Instruct-RR-Broken-APS-vllm --port YOUR_PORT
vllm chat --url YOUR_URL
We provide some probes at HuggingFace. You can directly load them with from_pretrained.
You can generate probes by running iterSCAV.py. For example, run:
CUDA_VISIBLE_DEVICES="0" python iterSCAV.py --judge srf --bs 50 --maxIter 50 --trainL 512 --layer -2 --model 'GraySwanAI/Mistral-7B-Instruct-RR' --saveDir ./iterSCAVWeight --thres 0.05 0.8 --linearC cuSVC --evalPT "1" --pt "mean" --embType last --train "rd" --val "rdVal"
After training the probes, you can evaluate the probe-based steering by running eval.py. For example, run
CUDA_VISIBLE_DEVICES="7" python eval.py --evalData sr --bs 25 --maxL 512 --model 'GraySwanAI/Mistral-7B-Instruct-RR' --evalPT "1" --clfP "path_to_probe --csvP myRes.csv --evalClfr 'best' --judge "srf" "hb" "qwen-plus https://dashscope.aliyuncs.com/compatible-mode/v1 sk-xxxxx"
Which component is important for a model extraction (or active learning) algorithm?
First, the judge should be accurate. Without an accurate judge, how can the probe, which approximates the judge, be accurate? You can add some new powerful judges into myJudge.py. Of course, you can also utilize commercial big LLMs if you are ok with their randomness induced by the lack of batch invariance. We use SRF, which is finetuned from an old gemma-2b, by default because it is deterministic and saves token charge.
Second, you may develop some better and adaptive steering strength schemes for sampling hidden states for the next iteration. We set strength="abs0", which steers the hidden state to the boundary, because classic model extraction (or active learning) algorithms do so. We have proven that increasing the strength to "50%" can slightly improve our method, as shown in Table 7. So, don't let the default setting limit you (Yet, it is ok to use the default setting if you want to beat our method in your new paper).
Third, explore which part of the hidden state can best reflect the LLM's text output. We use the hidden state located at the first response token position by default because baselines do so and because we do not want to bargain with reviewers about the fairness of comparison. We have also proven that using the averaged hidden states from all response token positions can improve our method against some LLMs in Table 7. So, if you do not give a shit about the troubling peer review, do explore the alignment between the hidden states and the text output.
Fourth, why not use more prompts? We include 50 harmful prompts during the model extraction because we have limited computation resources and have to benchmark the experiment, again, to deal with the peer review. If you get plenty of daily inference that can be annotated, you may store the hidden states to train the probe iteratively. I don't believe there will be some awful reviewers jumping out of nowhere and arguing that this scheme is not fair.
42 commits
Python
98.9%
Shell
1.1%
This is a repository containing our method only.
We now support vLLM through customized model registered through plugin system. After installing vllm, run
cd vllm_steered_plugins
pip install -e .
You can use transferSteer2Safetensors.py to transform Probes + HF Model into a repo that can be loaded by vLLM (and also Transformers). We have uploaded jailbroken Circuit Breaker. You can try it with
vllm serve FTK11558/Llama-3-8B-Instruct-RR-Broken-APS-vllm --port YOUR_PORT
vllm chat --url YOUR_URL
We provide some probes at HuggingFace. You can directly load them with from_pretrained.
You can generate probes by running iterSCAV.py. For example, run:
CUDA_VISIBLE_DEVICES="0" python iterSCAV.py --judge srf --bs 50 --maxIter 50 --trainL 512 --layer -2 --model 'GraySwanAI/Mistral-7B-Instruct-RR' --saveDir ./iterSCAVWeight --thres 0.05 0.8 --linearC cuSVC --evalPT "1" --pt "mean" --embType last --train "rd" --val "rdVal"
After training the probes, you can evaluate the probe-based steering by running eval.py. For example, run
CUDA_VISIBLE_DEVICES="7" python eval.py --evalData sr --bs 25 --maxL 512 --model 'GraySwanAI/Mistral-7B-Instruct-RR' --evalPT "1" --clfP "path_to_probe --csvP myRes.csv --evalClfr 'best' --judge "srf" "hb" "qwen-plus https://dashscope.aliyuncs.com/compatible-mode/v1 sk-xxxxx"
Which component is important for a model extraction (or active learning) algorithm?
First, the judge should be accurate. Without an accurate judge, how can the probe, which approximates the judge, be accurate? You can add some new powerful judges into myJudge.py. Of course, you can also utilize commercial big LLMs if you are ok with their randomness induced by the lack of batch invariance. We use SRF, which is finetuned from an old gemma-2b, by default because it is deterministic and saves token charge.
Second, you may develop some better and adaptive steering strength schemes for sampling hidden states for the next iteration. We set strength="abs0", which steers the hidden state to the boundary, because classic model extraction (or active learning) algorithms do so. We have proven that increasing the strength to "50%" can slightly improve our method, as shown in Table 7. So, don't let the default setting limit you (Yet, it is ok to use the default setting if you want to beat our method in your new paper).
Third, explore which part of the hidden state can best reflect the LLM's text output. We use the hidden state located at the first response token position by default because baselines do so and because we do not want to bargain with reviewers about the fairness of comparison. We have also proven that using the averaged hidden states from all response token positions can improve our method against some LLMs in Table 7. So, if you do not give a shit about the troubling peer review, do explore the alignment between the hidden states and the text output.
Fourth, why not use more prompts? We include 50 harmful prompts during the model extraction because we have limited computation resources and have to benchmark the experiment, again, to deal with the peer review. If you get plenty of daily inference that can be annotated, you may store the hidden states to train the probe iteratively. I don't believe there will be some awful reviewers jumping out of nowhere and arguing that this scheme is not fair.
42 commits
Python
98.9%
Shell
1.1%