An interactive demo for MOSS-VL-Instruct-0408, an 11B-parameter instruction-tuned vision-language model developed by the OpenMOSS Team. Built on MOSS-VL-Base-0408 through supervised fine-tuning, it serves as a high-performance offline multimodal engine with particular strength in video understanding.
MOSS-VL adopts a cross-attention-based architecture that decouples visual encoding from cognitive reasoning:
Note: The model weights (~22 GB) may take a few minutes to load on first use (cold start). Subsequent requests will be faster.
@misc{moss_vl_2026,
title = {{MOSS-VL Technical Report}},
author = {OpenMOSS Team},
year = {2026},
howpublished = {\url{https://github.com/OpenMOSS/MOSS-VL}},
note = {GitHub repository}
}
1 commits
An interactive demo for MOSS-VL-Instruct-0408, an 11B-parameter instruction-tuned vision-language model developed by the OpenMOSS Team. Built on MOSS-VL-Base-0408 through supervised fine-tuning, it serves as a high-performance offline multimodal engine with particular strength in video understanding.
MOSS-VL adopts a cross-attention-based architecture that decouples visual encoding from cognitive reasoning:
Note: The model weights (~22 GB) may take a few minutes to load on first use (cold start). Subsequent requests will be faster.
@misc{moss_vl_2026,
title = {{MOSS-VL Technical Report}},
author = {OpenMOSS Team},
year = {2026},
howpublished = {\url{https://github.com/OpenMOSS/MOSS-VL}},
note = {GitHub repository}
}
1 commits