HaoyeZhang/RLHF-V-Dataset

Dataset

Dataset Card for RLHF-V-Dataset

74

24 commits

2 linked in READMEs

updated May 28, 2024

See the code

README

Dataset Card for RLHF-V-Dataset

Project Page | Paper | GitHub

Updates

  • [2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now!
  • [2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now!
  • [2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that exhibits comparable trustworthiness with GPT-4V!
  • [2024.01.06] 🔥 A larger, more diverse set of fine-grained human correction data is available now! 🔥 The newly released data has about 5.7k of fine-grained human correction data that covers the output of more powerful models (Qwen-VL-Chat, InstructBLIP, etc.). We also expand the image types from everyday scenes to diverse styles and themes (WikiArt, landmarks, scene texts, etc.).
  • [2024.01.05] 🔧 We reformat our dataset and now it is more convenient to preview and use our data! The dataset now supports the load_dataset function, and the data content can be easily previewed online.
  • [2023.12.15] We incorporated a new annotation subset with an additional 1065 fine-grained annotations into our dataset !

Dataset Summary

RLHF-V-Dataset is the human preference data used in "RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback".

We collected a large amount of fine-grained segment-level human corrections on diverse instructions, including detailed descriptions and question-answering instructions. The dataset contains a total of 5,733 preference pairs.

fig1

Utilizing our dataset can dramatically reduce model hallucinations by 34.8% while keeping informativeness.

fig2

Usage

from datasets import load_dataset

data = load_dataset("HaoyeZhang/RLHF-V-Dataset")

Data fields

KeyDescription
0ds_nameDataset name.
1imageDict contains path and bytes. If loaded by load_dataset, it can be automatically converted into a PIL Image.
2textPreference data. Each data item contains a dict with the keys "question", "chosen", and "rejected".
3origin_datasetOriginal dataset for annotation, which is not used in training.
4origin_splitMeta information for each data item, including the name of the model we use to generate the original answer, and the question type ("detailed description" or "question answering")
5idxData index.
6image_pathImage path.

Citation

If you find this dataset helpful, please consider cite our papers 📝:

@article{yu2023rlhf,
  title={Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback},
  author={Yu, Tianyu and Yao, Yuan and Zhang, Haoye and He, Taiwen and Han, Yifeng and Cui, Ganqu and Hu, Jinyi and Liu, Zhiyuan and Zheng, Hai-Tao and Sun, Maosong and others},
  journal={arXiv preprint arXiv:2312.00849},
  year={2023}
}

@article{yu2024rlaifv,
  title={RLAIF-V: Aligning MLLMs through Open-Source AI Feedback for Super GPT-4V Trustworthiness}, 
  author={Yu, Tianyu and Zhang, Haoye and Yao, Yuan and Dang, Yunkai and Chen, Da and Lu, Xiaoman and Cui, Ganqu and He, Taiwen and Liu, Zhiyuan and Chua, Tat-Seng and Sun, Maosong},
  journal={arXiv preprint arXiv:2405.17220},
  year={2024},
}

Contributors

HaoyeZhang

20 commits

Yirany

4 commits

HaoyeZhang/RLHF-V-Dataset

Dataset

Dataset Card for RLHF-V-Dataset

74

24 commits

2 linked in READMEs

updated May 28, 2024

See the code

README

Dataset Card for RLHF-V-Dataset

Project Page | Paper | GitHub

Updates

  • [2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now!
  • [2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now!
  • [2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that exhibits comparable trustworthiness with GPT-4V!
  • [2024.01.06] 🔥 A larger, more diverse set of fine-grained human correction data is available now! 🔥 The newly released data has about 5.7k of fine-grained human correction data that covers the output of more powerful models (Qwen-VL-Chat, InstructBLIP, etc.). We also expand the image types from everyday scenes to diverse styles and themes (WikiArt, landmarks, scene texts, etc.).
  • [2024.01.05] 🔧 We reformat our dataset and now it is more convenient to preview and use our data! The dataset now supports the load_dataset function, and the data content can be easily previewed online.
  • [2023.12.15] We incorporated a new annotation subset with an additional 1065 fine-grained annotations into our dataset !

Dataset Summary

RLHF-V-Dataset is the human preference data used in "RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback".

We collected a large amount of fine-grained segment-level human corrections on diverse instructions, including detailed descriptions and question-answering instructions. The dataset contains a total of 5,733 preference pairs.

fig1

Utilizing our dataset can dramatically reduce model hallucinations by 34.8% while keeping informativeness.

fig2

Usage

from datasets import load_dataset

data = load_dataset("HaoyeZhang/RLHF-V-Dataset")

Data fields

KeyDescription
0ds_nameDataset name.
1imageDict contains path and bytes. If loaded by load_dataset, it can be automatically converted into a PIL Image.
2textPreference data. Each data item contains a dict with the keys "question", "chosen", and "rejected".
3origin_datasetOriginal dataset for annotation, which is not used in training.
4origin_splitMeta information for each data item, including the name of the model we use to generate the original answer, and the question type ("detailed description" or "question answering")
5idxData index.
6image_pathImage path.

Citation

If you find this dataset helpful, please consider cite our papers 📝:

@article{yu2023rlhf,
  title={Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback},
  author={Yu, Tianyu and Yao, Yuan and Zhang, Haoye and He, Taiwen and Han, Yifeng and Cui, Ganqu and Hu, Jinyi and Liu, Zhiyuan and Zheng, Hai-Tao and Sun, Maosong and others},
  journal={arXiv preprint arXiv:2312.00849},
  year={2023}
}

@article{yu2024rlaifv,
  title={RLAIF-V: Aligning MLLMs through Open-Source AI Feedback for Super GPT-4V Trustworthiness}, 
  author={Yu, Tianyu and Zhang, Haoye and Yao, Yuan and Dang, Yunkai and Chen, Da and Lu, Xiaoman and Cui, Ganqu and He, Taiwen and Liu, Zhiyuan and Chua, Tat-Seng and Sun, Maosong},
  journal={arXiv preprint arXiv:2405.17220},
  year={2024},
}

Contributors

HaoyeZhang

20 commits

Yirany

4 commits