Welcome to Xiaomi MiMo-VL-Miloco — the first open-source multimodal model built to actually understand what’s happening at home!
We use a carefully tuned two-stage pipeline to nail home-scene skills without sacrificing general abilities.
This stage focuses on boosting the model’s core capabilities in home scenarios. Even with a limited training set, we strike a good balance between sample-efficient learning and fast inference:
Building on fine-tuning, this stage introduces GRPO-based reinforcement learning to enhance the model’s overall performance:
In short: Xiaomi MiMo-VL-Miloco is your friendly, sharp-eyed model roommate—great at recognizing what’s going on around the house, and still ready for the wider world.
Both versions of the MiMo-VL-Miloco-7B model are now open-sourced:
In household scene understanding, we prioritize video and image perception alongside the model’s reasoning ability.
@misc{xiaomimimovlmiloco,
author = {Jiaze Li, Yuxun Qu, Jingyang Chen, Shijie Xu, Zhenru Lin, Junyou Zhu, Boshen Xu, Wenhui Tan, Pei Fu, JianZhong Ju, Zhenbo Luo, Jian Luan},
title = {Xiaomi MiMo-VL-Miloco},
year = {2025},
howpublished = {\url{https://github.com/XiaoMi/xiaomi-mimo-vl-miloco}},
}
Please contact us at milm-plus@xiaomi.com or open an issue if you have any questions.
7 commits
Welcome to Xiaomi MiMo-VL-Miloco — the first open-source multimodal model built to actually understand what’s happening at home!
We use a carefully tuned two-stage pipeline to nail home-scene skills without sacrificing general abilities.
This stage focuses on boosting the model’s core capabilities in home scenarios. Even with a limited training set, we strike a good balance between sample-efficient learning and fast inference:
Building on fine-tuning, this stage introduces GRPO-based reinforcement learning to enhance the model’s overall performance:
In short: Xiaomi MiMo-VL-Miloco is your friendly, sharp-eyed model roommate—great at recognizing what’s going on around the house, and still ready for the wider world.
Both versions of the MiMo-VL-Miloco-7B model are now open-sourced:
In household scene understanding, we prioritize video and image perception alongside the model’s reasoning ability.
@misc{xiaomimimovlmiloco,
author = {Jiaze Li, Yuxun Qu, Jingyang Chen, Shijie Xu, Zhenru Lin, Junyou Zhu, Boshen Xu, Wenhui Tan, Pei Fu, JianZhong Ju, Zhenbo Luo, Jian Luan},
title = {Xiaomi MiMo-VL-Miloco},
year = {2025},
howpublished = {\url{https://github.com/XiaoMi/xiaomi-mimo-vl-miloco}},
}
Please contact us at milm-plus@xiaomi.com or open an issue if you have any questions.
7 commits