ColorfulAI/LLaVA-NeXT-Speech

Dataset

we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.

0

3 commits

5 linked in READMEs

updated Apr 1, 2025

See the code

README

we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.

ColorfulAI/LLaVA-NeXT-Speech

Dataset

we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.

0

3 commits

5 linked in READMEs

updated Apr 1, 2025

See the code

README

we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.