This is the repository for Leopard, a MLLM that is specifically designed to handle complex vision-language tasks involving multiple text-rich images. In real-world applications, such as presentation slides, scanned documents, and webpage snapshots, understanding the inter-relationships and logical flow across multiple images is crucial.
The code, data, and model checkpoints will be released in one month. Stay tuned!
For evaluation, please refer to the Evaluations section.
We provide the checkpoints of Leopard-LLaVA and Leopard-Idefics2 on Huggingface.
For model training, please refer to the Training section.
1 commits
Python
93.8%
Shell
3.9%
C++
1.7%
This is the repository for Leopard, a MLLM that is specifically designed to handle complex vision-language tasks involving multiple text-rich images. In real-world applications, such as presentation slides, scanned documents, and webpage snapshots, understanding the inter-relationships and logical flow across multiple images is crucial.
The code, data, and model checkpoints will be released in one month. Stay tuned!
For evaluation, please refer to the Evaluations section.
We provide the checkpoints of Leopard-LLaVA and Leopard-Idefics2 on Huggingface.
For model training, please refer to the Training section.
1 commits
Python
93.8%
Shell
3.9%
C++
1.7%