Dataset type: M4-Instruct is a set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API. It is constructed for training LMMs for their interleaved multi-image capbilities, e.g., LLaVA-NeXT-Interleave.
Dataset date: M4-Instruct was collected in April 2024, and released in June 2024.
Paper or resources for more information: Blog: https://llava-vl.github.io/blog/2024-06-16-llava-next-interleave/
Data Statistics:
Note that we only released the multi-image, multi-frame (video), and multi-view (3D) data of M4-Instruct in this repo.
Please also download videos from LLaVA-Hound, and organize them according to their instructions.
To reproduce LLaVA-NeXT-Interleave, you can complement the multi-patch (single-image) data by randomly sampling 307K of the stage-2 SFT data of LLaVA-1.5.
Data Content:
License: Creative Commons Attribution 4.0 International; and it should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Where to send questions or comments about the model: fliay@connect.ust.hk 1700012927@pku.edu.cn
Primary intended uses: The primary use of M4-Instruct Data is research on large multimodal models and chatbots.
Primary intended users: The primary intended users of the model are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.
28 commits
2 commits
Dataset type: M4-Instruct is a set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API. It is constructed for training LMMs for their interleaved multi-image capbilities, e.g., LLaVA-NeXT-Interleave.
Dataset date: M4-Instruct was collected in April 2024, and released in June 2024.
Paper or resources for more information: Blog: https://llava-vl.github.io/blog/2024-06-16-llava-next-interleave/
Data Statistics:
Note that we only released the multi-image, multi-frame (video), and multi-view (3D) data of M4-Instruct in this repo.
Please also download videos from LLaVA-Hound, and organize them according to their instructions.
To reproduce LLaVA-NeXT-Interleave, you can complement the multi-patch (single-image) data by randomly sampling 307K of the stage-2 SFT data of LLaVA-1.5.
Data Content:
License: Creative Commons Attribution 4.0 International; and it should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Where to send questions or comments about the model: fliay@connect.ust.hk 1700012927@pku.edu.cn
Primary intended uses: The primary use of M4-Instruct Data is research on large multimodal models and chatbots.
Primary intended users: The primary intended users of the model are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.
28 commits
2 commits