HuanjinYao/Mulberry-SFT

Dataset

11

stars

6

commits

4

linked in READMEs

Jan 26, 2025

updated

MLLM

README

Please check our GitHub for more details.: https://github.com/HJYao00/Mulberry

Training

We use LLaMA-Factory to fine-tune the Mulberry models. We provide the training instructions and configs here.

First, install LLaMA-Factory according to the official_instruction.

Then, refer here and update the following customized dataset into dataset_info.json in LLaMA-Factory.

"mulberry": {
    "file_name": "./mulberry_sft.json",
    "formatting": "sharegpt",
    "columns": {
      "messages": "messages",
      "images": "images"
    },
    "tags": {
      "role_tag": "role",
      "content_tag": "content",
      "user_tag": "user",
      "assistant_tag": "assistant"
    }
  },

Finally, you can use the following command to train the models.

llamafactory-cli train examples/train_full/mulberry_llava_8b_full_sft.yaml

Citation

@article{yao2024mulberry,
  title={Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search},
  author={Yao, Huanjin and Huang, Jiaxing and Wu, Wenhao and Zhang, Jingyi and Wang, Yibo and Liu, Shunyu and Wang, Yingjie and Song, Yuxin and Feng, Haocheng and Shen, Li and others},
  journal={arXiv preprint arXiv:2412.18319},
  year={2024}
}

Contributors

HuanjinYao

6 commits

HuanjinYao/Mulberry-SFT

Dataset

11

stars

6

commits

4

linked in READMEs

Jan 26, 2025

updated

MLLM

README

Please check our GitHub for more details.: https://github.com/HJYao00/Mulberry

Training

We use LLaMA-Factory to fine-tune the Mulberry models. We provide the training instructions and configs here.

First, install LLaMA-Factory according to the official_instruction.

Then, refer here and update the following customized dataset into dataset_info.json in LLaMA-Factory.

"mulberry": {
    "file_name": "./mulberry_sft.json",
    "formatting": "sharegpt",
    "columns": {
      "messages": "messages",
      "images": "images"
    },
    "tags": {
      "role_tag": "role",
      "content_tag": "content",
      "user_tag": "user",
      "assistant_tag": "assistant"
    }
  },

Finally, you can use the following command to train the models.

llamafactory-cli train examples/train_full/mulberry_llava_8b_full_sft.yaml

Citation

@article{yao2024mulberry,
  title={Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search},
  author={Yao, Huanjin and Huang, Jiaxing and Wu, Wenhao and Zhang, Jingyi and Wang, Yibo and Liu, Shunyu and Wang, Yingjie and Song, Yuxin and Feng, Haocheng and Shen, Li and others},
  journal={arXiv preprint arXiv:2412.18319},
  year={2024}
}

Contributors

HuanjinYao

6 commits