CuMo contains three-stage training:
Stage 1: Pre-Training
In this stage, we use LLaVA-558K to pretrain MLP.
Stage 2: Pre-FineTuning
For pre-finetuning, we use the ALLaVA caption data, you may use the original one or the cumo_pft_allava.json in this repo.
Stage 3: Visual Instruction Tuning
Please download these datasets following the instructions and cumo_vit_1649K.json for visual instruction tuning.
CuMo utilizes these datasets that are subject to their respective original licenses. Users must comply with all terms and conditions specified in these original licenses.
CuMo contains three-stage training:
Stage 1: Pre-Training
In this stage, we use LLaVA-558K to pretrain MLP.
Stage 2: Pre-FineTuning
For pre-finetuning, we use the ALLaVA caption data, you may use the original one or the cumo_pft_allava.json in this repo.
Stage 3: Visual Instruction Tuning
Please download these datasets following the instructions and cumo_vit_1649K.json for visual instruction tuning.
CuMo utilizes these datasets that are subject to their respective original licenses. Users must comply with all terms and conditions specified in these original licenses.