We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks.
paper: https://arxiv.org/abs/2408.03695
code: https://github.com/YeLuoSuiYou/openstorypp
2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected from YouTube. Due to copyright issues, we have not released the raw images, but we will provide the method of organizing the dataset later.
2024/6/22 we release the high-quality subset of our dataset’s unique data, which contains about 15M sample.
We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks.
paper: https://arxiv.org/abs/2408.03695
code: https://github.com/YeLuoSuiYou/openstorypp
2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected from YouTube. Due to copyright issues, we have not released the raw images, but we will provide the method of organizing the dataset later.
2024/6/22 we release the high-quality subset of our dataset’s unique data, which contains about 15M sample.