VLFeedback is a large-scale vision-language preference dataset, annotated by GPT-4V. It consists of 80k multi-modal instructions from various souces that encompass various capabilities of LVLMs.
We build a model pool of 12 LVLMs and each data sample contains 4 responses from different models. Each response is annotated in three aspects: helpfulness, visual faithfulness, and ethical considerations. The resulting preference dataset contains more than 380k comparison pairs.
@article{2023vlfeedback,
author = {Lei Li and Zhihui Xie and Mukai Li and Shunian Chen and Peiyi Wang and Liang Chen and Yazheng Yang and Benyou Wang and Lingpeng Kong},
title = {Silkie: Preference Distillation for Large Visual Language Models},
publisher = {arXiv:2312.10665},
year = {2023}
}
VLFeedback is a large-scale vision-language preference dataset, annotated by GPT-4V. It consists of 80k multi-modal instructions from various souces that encompass various capabilities of LVLMs.
We build a model pool of 12 LVLMs and each data sample contains 4 responses from different models. Each response is annotated in three aspects: helpfulness, visual faithfulness, and ethical considerations. The resulting preference dataset contains more than 380k comparison pairs.
@article{2023vlfeedback,
author = {Lei Li and Zhihui Xie and Mukai Li and Shunian Chen and Peiyi Wang and Liang Chen and Yazheng Yang and Benyou Wang and Lingpeng Kong},
title = {Silkie: Preference Distillation for Large Visual Language Models},
publisher = {arXiv:2312.10665},
year = {2023}
}