we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.
0
3 commits
5 linked in READMEs
updated Apr 1, 2025
we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.
we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.
0
3 commits
5 linked in READMEs
updated Apr 1, 2025
we randomly select 300k visual instructions from LLaVA-NeXT, and convert the user queries into audio using CosyVoice with randomly selected audio prompts from VoiceAssistant-400K.