FoundationVision/groma_instruct

Dataset

Groma Instruct contains 30k visually grounded conversations for instruction finetuning, which is the first grounded chat dataset constructed with both visual and textual prompts, leveraging the powerful GPT-4V for data generation.

8

4 commits

2 linked in READMEs

updated May 2, 2024

See the code

README

Groma Instruct contains 30k visually grounded conversations for instruction finetuning, which is the first grounded chat dataset constructed with both visual and textual prompts, leveraging the powerful GPT-4V for data generation. For more information about this dataset and its usage, please take a look at our paper and github page.

Contributors

machuofan

4 commits

FoundationVision/groma_instruct

Dataset

Groma Instruct contains 30k visually grounded conversations for instruction finetuning, which is the first grounded chat dataset constructed with both visual and textual prompts, leveraging the powerful GPT-4V for data generation.

8

4 commits

2 linked in READMEs

updated May 2, 2024

See the code

README

Groma Instruct contains 30k visually grounded conversations for instruction finetuning, which is the first grounded chat dataset constructed with both visual and textual prompts, leveraging the powerful GPT-4V for data generation. For more information about this dataset and its usage, please take a look at our paper and github page.

Contributors

machuofan

4 commits