mucai/ViP-LLaVA-Instruct

Dataset

ViP-LLaVA Instruct Dataset Card

10

4 commits

2 linked in READMEs

updated Feb 26, 2024

See the code

README

ViP-LLaVA Instruct Dataset Card

Dataset details

Dataset type: ViP-LLaVA Instruct is composed of a mixture of LLaVA-1.5 instruction data and the region-level visual prompting data. It is constructed for visual instruction tuning and for building large multimodal towards GPT-4 level regional understanding capability.

Specifically, we use 1.2M data for stage 2 finetuning, and use 26K data for the optional stage 3 finetuning.

Dataset date: ViP-LLaVA Instruct was collected in November 2023, by using a mixture of academic dataset and GPT-4/GPT-4V instructed dataset.

Paper or resources for more information: https://vip-llava.github.io/

License: Apache-2.0; and it should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use

Where to send questions or comments about the model: https://github.com/mu-cai/ViP-LLaVA/issues

Intended use

Primary intended uses: The primary use of ViP-LLaVA is research on large multimodal models and chatbots.

Primary intended users: The primary intended users of the model are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.

Contributors

mucai

4 commits

mucai/ViP-LLaVA-Instruct

Dataset

ViP-LLaVA Instruct Dataset Card

10

4 commits

2 linked in READMEs

updated Feb 26, 2024

See the code

README

ViP-LLaVA Instruct Dataset Card

Dataset details

Dataset type: ViP-LLaVA Instruct is composed of a mixture of LLaVA-1.5 instruction data and the region-level visual prompting data. It is constructed for visual instruction tuning and for building large multimodal towards GPT-4 level regional understanding capability.

Specifically, we use 1.2M data for stage 2 finetuning, and use 26K data for the optional stage 3 finetuning.

Dataset date: ViP-LLaVA Instruct was collected in November 2023, by using a mixture of academic dataset and GPT-4/GPT-4V instructed dataset.

Paper or resources for more information: https://vip-llava.github.io/

License: Apache-2.0; and it should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use

Where to send questions or comments about the model: https://github.com/mu-cai/ViP-LLaVA/issues

Intended use

Primary intended uses: The primary use of ViP-LLaVA is research on large multimodal models and chatbots.

Primary intended users: The primary intended users of the model are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.

Contributors

mucai

4 commits