0
stars
3
commits
Jupyter Notebook
primary language
Jan 14, 2024
updated
We present MobileVLM, a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comprises a set of language models at the scale of 1.4B and 2.7B parameters, trained from scratch, a multimodal vision model that is pre-trained in the CLIP fashion, cross-modality interaction via an efficient projector. We evaluate MobileVLM on several typical VLM benchmarks. Our models demonstrate on par performance compared with a few much larger models. More importantly, we measure the inference speed on both a Qualcomm Snapdragon 888 CPU and an NVIDIA Jeston Orin GPU, and we obtain state-of-the-art performance of 21.5 tokens and 65.3 tokens per second, respectively.

The MobileVLM architecture (right) utilizes MobileLLaMA as its language model, intakes $\mathbf{X}_v$ and $\mathbf{X}_q$ which are image and language instructions as respective inputs and gives $\mathbf{Y}_a$ as the output language response. LDP refers to a lightweight downsample projector (left).
Jan. 11st, 2024: The training and evaluation codes of MobileVLM are available now! Follow these step-by-step instructions below to easily train your own mobileVLM in 5 hours β‘οΈ !Dec. 31st, 2023: Our MobileVLM weights are uploaded on the HuggingFace website. We also provide inference examples for the MobileLLaMA/MobileVLM model so that anyone can enjoy them early.Dec. 29th, 2023: Our MobileLLaMA weights are uploaded on the HuggingFace website. Enjoy them !Dec. 28th, 2023: π₯π₯π₯ We release MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices on arxiv. Refer to our paper for more details !π Usage and License Notices: This project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses. This project is licensed permissively under the Apache 2.0 license and does not impose any additional constraints. LLaVA
Clone this repository and navigate to MobileVLM folder
git clone https://github.com/Meituan-AutoML/MobileVLM.git
cd MobileVLM
Install Package
conda create -n mobilevlm python=3.10 -y
conda activate mobilevlm
pip install --upgrade pip
pip install -r requirements.txt
import torch
from transformers import LlamaTokenizer, LlamaForCausalLM
model_path = 'mtgv/MobileLLaMA-1.4B-Chat'
tokenizer = LlamaTokenizer.from_pretrained(model_path)
model = LlamaForCausalLM.from_pretrained(
model_path, torch_dtype=torch.float16, device_map='auto',
)
prompt = 'Q: What is the largest animal?\nA:'
input_ids = tokenizer(prompt, return_tensors="pt").input_ids
generation_output = model.generate(
input_ids=input_ids, max_new_tokens=32
)
print(tokenizer.decode(generation_output[0]))
from scripts.inference import inference_once
model_path = "mtgv/MobileVLM-1.7B"
image_file = "assets/samples/demo.jpg"
prompt_str = "Who is the author of this book?\nAnswer the question using a single word or phrase."
# (or) What is the title of this book?
# (or) Is this book related to Education & Teaching?
args = type('Args', (), {
"model_path": model_path,
"image_file": image_file,
"prompt": prompt_str,
"conv_mode": "v1",
"temperature": 0,
"top_p": None,
"num_beams": 1,
"max_new_tokens": 512,
"load_8bit": False,
"load_4bit": False,
})()
inference_once(args)
The training process of MobileVLM is divided into two stages:
Note: To train on fewer GPU memory or cards, you can reduce the per_device_train_batch_size and increase the gradient_accumulation_steps accordingly. Always keep the global batch size the same: per_device_train_batch_size x gradient_accumulation_steps x num_gpus.
Download MobileLLaMA chatbot checkpoints from huggingface website (π€ 1.7B, 2.7B). Please note that this is optional (it depends on your working environment), run the training script we provide below and the model will be automatically downloaded by the transformers library.
For convenience, assume your working directory /path/to/project/mobilevlm as work_dir:
cd ${work_dir} && mkdir -p data/pretrain_data data/finetune_data data/benchmark_dataprepare alignment pre-training data
cd ${work_dir}/data/pretrain_dataprepare instruction tuning data
prepare benchmark data
We evaluate models on a diverse set of 6 benchmarks, i.e. GQA, MMBench, MME, POPE, SQA, TextVQA. We do not evaluate using beam search to make the inference process consistent with the chat demo of real-time outputs. You should follow these instructions to manage the datasets.
unzip benchmark_data.zip && cd benchmark_databmk_dir=${work_dir}/data/benchmark_datacd ${bmk_dir}/gqa && ln -s /path/to/gqa/images imagescd ${bmk_dir}/mme && ln -s /path/to/MME/MME_Benchmark_release_version imagescd ${bmk_dir}/pope && ln -s /path/to/pope/coco coco && ln -s /path/to/coco/val2014 val2014data/scienceqa folder of the ScienceQA repo.cd ${bmk_dir}/sqa && ln -s /path/to/sqa/images imagescd ${bmk_dir}/textvqa && ln -s /path/to/textvqa/train_images train_imagesorganize the data directory as follows after downloading all of them:
.
βββ benchmark_data
βΒ Β βββ gqa
βΒ Β βΒ Β βββ convert_gqa_for_eval.py
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ images -> /path/to/your/gqa/images
βΒ Β βΒ Β βββ llava_gqa_testdev_balanced.jsonl
βΒ Β βΒ Β βββ testdev_balanced_questions.json
βΒ Β βββ mmbench
βΒ Β βΒ Β βββ convert_mmbench_for_submission.py
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ mmbench_dev_en_20231003.tsv
βΒ Β βββ mme
βΒ Β βΒ Β βββ calculation.py
βΒ Β βΒ Β βββ convert_answer_to_mme.py
βΒ Β βΒ Β βββ images -> /path/to/your/MME/MME_Benchmark_release_version
βΒ Β βΒ Β βββ llava_mme.jsonl
βΒ Β βββ pope
βΒ Β βΒ Β βββ coco -> /path/to/your/pope/coco
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ llava_pope_test.jsonl
βΒ Β βΒ Β βββ val2014 -> /path/to/your/coco/val2014
βΒ Β βββ sqa
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ images -> /path/to/your/scienceqa/images
βΒ Β βΒ Β βββ llava_test_CQM-A.json
βΒ Β βΒ Β βββ pid_splits.json
βΒ Β βΒ Β βββ problems.json
βΒ Β βββ textvqa
βΒ Β βββ eval.py
βΒ Β βββ llava_textvqa_val_v051_ocr.jsonl
βΒ Β βββ m4c_evaluator.py
βΒ Β βββ TextVQA_0.5.1_val.json
βΒ Β βββ train_images -> /path/to/your/textvqa/train_images
βββ finetune_data
β βββ llava_v1_5_mix665k.json
β βββ coco
β β βββ train2017
β βββ gqa
β β βββ images
β βββ ocr_vqa
β β βββ images
β βββ textvqa
β β βββ train_images
β βββ vg
β βββ VG_100K
β βββ VG_100K_2
βββ pretrain_data
β βββ images
β βββ blip_laion_cc_sbu_558k.json
LANGUAGE_MODEL=/path/to/your/MobileLLaMA-1.4B-Chat # or 2.7B
VISION_MODEL=/path/to/your/clip-vit-large-patch14-336
bash run.sh mobilevlm1.7b pretrain-finetune-test ${LANGUAGE_MODEL} ${VISION_MODEL}
# (test-only) bash run.sh mobilevlm1.7b test /path/to/your/own/checkpoint
# (3B) bash run.sh mobilevlm3b pretrain-finetune-test ${LANGUAGE_MODEL} ${VISION_MODEL}
run.sh so they can be run with one click for simplification. If you would like to modify some super-parameters to observe their impact, please dive into run.sh to explore.If you find MobileVLM or MobileLLaMA useful in your research or applications, please consider giving a star β and citing using the following BibTeX:
@misc{chu2023mobilevlm,
title={MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices},
author={Xiangxiang Chu and Limeng Qiao and Xinyang Lin and Shuang Xu and Yang Yang and Yiming Hu and Fei Wei and Xinyu Zhang and Bo Zhang and Xiaolin Wei and Chunhua Shen},
year={2023},
eprint={2312.16886},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
3 commits
Jupyter Notebook
60.2%
Python
37.9%
Shell
1.9%
0
stars
3
commits
Jupyter Notebook
primary language
Jan 14, 2024
updated
We present MobileVLM, a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comprises a set of language models at the scale of 1.4B and 2.7B parameters, trained from scratch, a multimodal vision model that is pre-trained in the CLIP fashion, cross-modality interaction via an efficient projector. We evaluate MobileVLM on several typical VLM benchmarks. Our models demonstrate on par performance compared with a few much larger models. More importantly, we measure the inference speed on both a Qualcomm Snapdragon 888 CPU and an NVIDIA Jeston Orin GPU, and we obtain state-of-the-art performance of 21.5 tokens and 65.3 tokens per second, respectively.

The MobileVLM architecture (right) utilizes MobileLLaMA as its language model, intakes $\mathbf{X}_v$ and $\mathbf{X}_q$ which are image and language instructions as respective inputs and gives $\mathbf{Y}_a$ as the output language response. LDP refers to a lightweight downsample projector (left).
Jan. 11st, 2024: The training and evaluation codes of MobileVLM are available now! Follow these step-by-step instructions below to easily train your own mobileVLM in 5 hours β‘οΈ !Dec. 31st, 2023: Our MobileVLM weights are uploaded on the HuggingFace website. We also provide inference examples for the MobileLLaMA/MobileVLM model so that anyone can enjoy them early.Dec. 29th, 2023: Our MobileLLaMA weights are uploaded on the HuggingFace website. Enjoy them !Dec. 28th, 2023: π₯π₯π₯ We release MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices on arxiv. Refer to our paper for more details !π Usage and License Notices: This project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses. This project is licensed permissively under the Apache 2.0 license and does not impose any additional constraints. LLaVA
Clone this repository and navigate to MobileVLM folder
git clone https://github.com/Meituan-AutoML/MobileVLM.git
cd MobileVLM
Install Package
conda create -n mobilevlm python=3.10 -y
conda activate mobilevlm
pip install --upgrade pip
pip install -r requirements.txt
import torch
from transformers import LlamaTokenizer, LlamaForCausalLM
model_path = 'mtgv/MobileLLaMA-1.4B-Chat'
tokenizer = LlamaTokenizer.from_pretrained(model_path)
model = LlamaForCausalLM.from_pretrained(
model_path, torch_dtype=torch.float16, device_map='auto',
)
prompt = 'Q: What is the largest animal?\nA:'
input_ids = tokenizer(prompt, return_tensors="pt").input_ids
generation_output = model.generate(
input_ids=input_ids, max_new_tokens=32
)
print(tokenizer.decode(generation_output[0]))
from scripts.inference import inference_once
model_path = "mtgv/MobileVLM-1.7B"
image_file = "assets/samples/demo.jpg"
prompt_str = "Who is the author of this book?\nAnswer the question using a single word or phrase."
# (or) What is the title of this book?
# (or) Is this book related to Education & Teaching?
args = type('Args', (), {
"model_path": model_path,
"image_file": image_file,
"prompt": prompt_str,
"conv_mode": "v1",
"temperature": 0,
"top_p": None,
"num_beams": 1,
"max_new_tokens": 512,
"load_8bit": False,
"load_4bit": False,
})()
inference_once(args)
The training process of MobileVLM is divided into two stages:
Note: To train on fewer GPU memory or cards, you can reduce the per_device_train_batch_size and increase the gradient_accumulation_steps accordingly. Always keep the global batch size the same: per_device_train_batch_size x gradient_accumulation_steps x num_gpus.
Download MobileLLaMA chatbot checkpoints from huggingface website (π€ 1.7B, 2.7B). Please note that this is optional (it depends on your working environment), run the training script we provide below and the model will be automatically downloaded by the transformers library.
For convenience, assume your working directory /path/to/project/mobilevlm as work_dir:
cd ${work_dir} && mkdir -p data/pretrain_data data/finetune_data data/benchmark_dataprepare alignment pre-training data
cd ${work_dir}/data/pretrain_dataprepare instruction tuning data
prepare benchmark data
We evaluate models on a diverse set of 6 benchmarks, i.e. GQA, MMBench, MME, POPE, SQA, TextVQA. We do not evaluate using beam search to make the inference process consistent with the chat demo of real-time outputs. You should follow these instructions to manage the datasets.
unzip benchmark_data.zip && cd benchmark_databmk_dir=${work_dir}/data/benchmark_datacd ${bmk_dir}/gqa && ln -s /path/to/gqa/images imagescd ${bmk_dir}/mme && ln -s /path/to/MME/MME_Benchmark_release_version imagescd ${bmk_dir}/pope && ln -s /path/to/pope/coco coco && ln -s /path/to/coco/val2014 val2014data/scienceqa folder of the ScienceQA repo.cd ${bmk_dir}/sqa && ln -s /path/to/sqa/images imagescd ${bmk_dir}/textvqa && ln -s /path/to/textvqa/train_images train_imagesorganize the data directory as follows after downloading all of them:
.
βββ benchmark_data
βΒ Β βββ gqa
βΒ Β βΒ Β βββ convert_gqa_for_eval.py
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ images -> /path/to/your/gqa/images
βΒ Β βΒ Β βββ llava_gqa_testdev_balanced.jsonl
βΒ Β βΒ Β βββ testdev_balanced_questions.json
βΒ Β βββ mmbench
βΒ Β βΒ Β βββ convert_mmbench_for_submission.py
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ mmbench_dev_en_20231003.tsv
βΒ Β βββ mme
βΒ Β βΒ Β βββ calculation.py
βΒ Β βΒ Β βββ convert_answer_to_mme.py
βΒ Β βΒ Β βββ images -> /path/to/your/MME/MME_Benchmark_release_version
βΒ Β βΒ Β βββ llava_mme.jsonl
βΒ Β βββ pope
βΒ Β βΒ Β βββ coco -> /path/to/your/pope/coco
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ llava_pope_test.jsonl
βΒ Β βΒ Β βββ val2014 -> /path/to/your/coco/val2014
βΒ Β βββ sqa
βΒ Β βΒ Β βββ eval.py
βΒ Β βΒ Β βββ images -> /path/to/your/scienceqa/images
βΒ Β βΒ Β βββ llava_test_CQM-A.json
βΒ Β βΒ Β βββ pid_splits.json
βΒ Β βΒ Β βββ problems.json
βΒ Β βββ textvqa
βΒ Β βββ eval.py
βΒ Β βββ llava_textvqa_val_v051_ocr.jsonl
βΒ Β βββ m4c_evaluator.py
βΒ Β βββ TextVQA_0.5.1_val.json
βΒ Β βββ train_images -> /path/to/your/textvqa/train_images
βββ finetune_data
β βββ llava_v1_5_mix665k.json
β βββ coco
β β βββ train2017
β βββ gqa
β β βββ images
β βββ ocr_vqa
β β βββ images
β βββ textvqa
β β βββ train_images
β βββ vg
β βββ VG_100K
β βββ VG_100K_2
βββ pretrain_data
β βββ images
β βββ blip_laion_cc_sbu_558k.json
LANGUAGE_MODEL=/path/to/your/MobileLLaMA-1.4B-Chat # or 2.7B
VISION_MODEL=/path/to/your/clip-vit-large-patch14-336
bash run.sh mobilevlm1.7b pretrain-finetune-test ${LANGUAGE_MODEL} ${VISION_MODEL}
# (test-only) bash run.sh mobilevlm1.7b test /path/to/your/own/checkpoint
# (3B) bash run.sh mobilevlm3b pretrain-finetune-test ${LANGUAGE_MODEL} ${VISION_MODEL}
run.sh so they can be run with one click for simplification. If you would like to modify some super-parameters to observe their impact, please dive into run.sh to explore.If you find MobileVLM or MobileLLaMA useful in your research or applications, please consider giving a star β and citing using the following BibTeX:
@misc{chu2023mobilevlm,
title={MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices},
author={Xiangxiang Chu and Limeng Qiao and Xinyang Lin and Shuang Xu and Yang Yang and Yiming Hu and Fei Wei and Xinyu Zhang and Bo Zhang and Xiaolin Wei and Chunhua Shen},
year={2023},
eprint={2312.16886},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
3 commits
Jupyter Notebook
60.2%
Python
37.9%
Shell
1.9%