66
stars
30
commits
2
linked in READMEs
Jan 5, 2025
updated
π Homepage | π€ MAmmoTH-VL-8B | π» Code | π Arxiv | π PDF | π₯οΈ Demo
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.


@article{guo2024mammothvlelicitingmultimodalreasoning,
title={MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale},
author={Jarvis Guo and Tuney Zheng and Yuelin Bai and Bo Li and Yubo Wang and King Zhu and Yizhi Li and Graham Neubig and Wenhu Chen and Xiang Yue},
year={2024},
eprint={2412.05237},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.05237},
}
66
stars
30
commits
2
linked in READMEs
Jan 5, 2025
updated
π Homepage | π€ MAmmoTH-VL-8B | π» Code | π Arxiv | π PDF | π₯οΈ Demo
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.


@article{guo2024mammothvlelicitingmultimodalreasoning,
title={MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale},
author={Jarvis Guo and Tuney Zheng and Yuelin Bai and Bo Li and Yubo Wang and King Zhu and Yizhi Li and Graham Neubig and Wenhu Chen and Xiang Yue},
year={2024},
eprint={2412.05237},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.05237},
}