yyu/arxiv-attrprompt

Dataset

This is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias.

1

7 commits

1 linked in READMEs

updated Sep 13, 2023

See the code

README

This is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias. See the paper: https://arxiv.org/abs/2306.15895 for details.

  • label.txt: the label name for each class
  • train.jsonl: The original training set.
  • valid.jsonl: The original validation set.
  • test.jsonl: The original test set.
  • simprompt.jsonl: The training data generated by the simple prompt.
  • attrprompt.jsonl: The training data generated by the attributed prompt.

Note: Different than the other datasets, the labels for training/validation/test data are all a list instead of an integer as it is a multi-label classification dataset.

arxiv
multilabel_classification
scientific_papers

yyu/arxiv-attrprompt

Dataset

This is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias.

1

7 commits

1 linked in READMEs

updated Sep 13, 2023

See the code

README

This is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias. See the paper: https://arxiv.org/abs/2306.15895 for details.

  • label.txt: the label name for each class
  • train.jsonl: The original training set.
  • valid.jsonl: The original validation set.
  • test.jsonl: The original test set.
  • simprompt.jsonl: The training data generated by the simple prompt.
  • attrprompt.jsonl: The training data generated by the attributed prompt.

Note: Different than the other datasets, the labels for training/validation/test data are all a list instead of an integer as it is a multi-label classification dataset.

arxiv
multilabel_classification
scientific_papers