cygu/sampling-distill-train-data-kgw-k1-gamma0.25-delta1

Dataset

Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1"

0

5 commits

1 linked in READMEs

updated May 22, 2024

See the code

README

Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1"

Training data for sampling-based watermark distillation using the KGW \(k=1, \gamma=0.25, \delta=1\) watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7B with decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included in the samples).

cygu/sampling-distill-train-data-kgw-k1-gamma0.25-delta1

Dataset

Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1"

0

5 commits

1 linked in READMEs

updated May 22, 2024

See the code

README

Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1"

Training data for sampling-based watermark distillation using the KGW \(k=1, \gamma=0.25, \delta=1\) watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7B with decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included in the samples).