LinWeizheDragon/PreFLMR_ViT-L_ENCN

Model

PreFLMR model card

0

4 commits

2 linked in READMEs

updated Dec 23, 2024

See the code

README

PreFLMR model card

PreFLMR is an open-source model for multimodal knowledge retrieval. It is a transformer-based model that uses a combination of text and image inputs to retrieve relevant documents from a large corpus.

Model Details

PreFLMR_ViT-L_ENCN is based on PreFLMR_ViT-L, and the text_encoder is replaced with bge-m3 for training. The training dataset includes Chinese and English datasets.

Model Description

  • Model type: FLMRModelForRetrieval
  • Language(s) (NLP): English Chinese
  • License: MIT License

Paper and resources for more detail

Uses

Direct Use

This model can be used directly to retrieve documents from a large corpus using a combination of text and image input queries. The retrieval usage can be found in the official implementation.

Downstream Use

This model can be used combined with language models to create a retrieval-augmented language model. The use for Knowledge-based VQA can be found in RAVQA

How to Get Started with the Model

For details of training, indexing, and performing retrieval, please refer to here.

Training datasets

The model is pre-trained on three types of tasks with a total of nine datasets:

  1. Image to Text retrieval: WIT, KVQA, and CC3M
  2. Question to Text retrieval: MSMARCO
  3. Image & Question to Text retrieval: LLaVA, OVEN, OKVQA, Infoseek and E-VQA

These datasets were converted to retrieval format. For details on the dataset split and conversion process, please refer to the paper PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers. We will release the proprocessed datasets soon.

Evaluation datasets

We evaluate our models on WIT, LLaVA, OVEN, KVQA, IGLUE (subset of WIT), Infoseek, E-VQA, OKVQA and MSMARCO.

ModelVision EncoderText EncoderCheckpoint NameNo. Param.WIT(EN)WIT(CN)LLaVA(EN)LLaVA(CN)OVEN(EN)OVEN(CN)KVQA(EN)KVQA(CN)Infoseek(EN)Infoseek(CN)EVQA(EN)EVQA(CN)OKVQA(EN)OKVQA(CN)MSMARCO(EN)MSMARCO(CN)
PreFLMRViT-LBase-v2LinWeizheDragon/PreFLMR_ViT-L543M60.510.971.83.259.86.643.63.257.97.970.82.868.52.178.710.3
PreFLMRVit-L_ENCNbge-m3LinWeizheDragon/PreFLMR_ViT-L_ENCN883M60.883.471.158.960.858.841.137.341.939.758.046.613.913.382.682.3

For the evaluation metrics, WIT uses Recall@10 and all the rest datasets use Recall@5.

Citation

BibTeX:

@article{Lin_Mei_Chen_Byrne_2024, 
        title={PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers}, 
        url={http://arxiv.org/abs/2402.08327}, 
        number={arXiv:2402.08327}, 
        publisher={arXiv}, 
        author={Lin, Weizhe and Mei, Jingbiao and Chen, Jinghong and Byrne, Bill}, 
        year={2024}}
custom_code
feature-extraction
flmr
FLMR
knowledge-based visual question answering
multi-modal
PreFLMR
retrieval
safetensors
transformers

LinWeizheDragon/PreFLMR_ViT-L_ENCN

Model

PreFLMR model card

0

4 commits

2 linked in READMEs

updated Dec 23, 2024

See the code

README

PreFLMR model card

PreFLMR is an open-source model for multimodal knowledge retrieval. It is a transformer-based model that uses a combination of text and image inputs to retrieve relevant documents from a large corpus.

Model Details

PreFLMR_ViT-L_ENCN is based on PreFLMR_ViT-L, and the text_encoder is replaced with bge-m3 for training. The training dataset includes Chinese and English datasets.

Model Description

  • Model type: FLMRModelForRetrieval
  • Language(s) (NLP): English Chinese
  • License: MIT License

Paper and resources for more detail

Uses

Direct Use

This model can be used directly to retrieve documents from a large corpus using a combination of text and image input queries. The retrieval usage can be found in the official implementation.

Downstream Use

This model can be used combined with language models to create a retrieval-augmented language model. The use for Knowledge-based VQA can be found in RAVQA

How to Get Started with the Model

For details of training, indexing, and performing retrieval, please refer to here.

Training datasets

The model is pre-trained on three types of tasks with a total of nine datasets:

  1. Image to Text retrieval: WIT, KVQA, and CC3M
  2. Question to Text retrieval: MSMARCO
  3. Image & Question to Text retrieval: LLaVA, OVEN, OKVQA, Infoseek and E-VQA

These datasets were converted to retrieval format. For details on the dataset split and conversion process, please refer to the paper PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers. We will release the proprocessed datasets soon.

Evaluation datasets

We evaluate our models on WIT, LLaVA, OVEN, KVQA, IGLUE (subset of WIT), Infoseek, E-VQA, OKVQA and MSMARCO.

ModelVision EncoderText EncoderCheckpoint NameNo. Param.WIT(EN)WIT(CN)LLaVA(EN)LLaVA(CN)OVEN(EN)OVEN(CN)KVQA(EN)KVQA(CN)Infoseek(EN)Infoseek(CN)EVQA(EN)EVQA(CN)OKVQA(EN)OKVQA(CN)MSMARCO(EN)MSMARCO(CN)
PreFLMRViT-LBase-v2LinWeizheDragon/PreFLMR_ViT-L543M60.510.971.83.259.86.643.63.257.97.970.82.868.52.178.710.3
PreFLMRVit-L_ENCNbge-m3LinWeizheDragon/PreFLMR_ViT-L_ENCN883M60.883.471.158.960.858.841.137.341.939.758.046.613.913.382.682.3

For the evaluation metrics, WIT uses Recall@10 and all the rest datasets use Recall@5.

Citation

BibTeX:

@article{Lin_Mei_Chen_Byrne_2024, 
        title={PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers}, 
        url={http://arxiv.org/abs/2402.08327}, 
        number={arXiv:2402.08327}, 
        publisher={arXiv}, 
        author={Lin, Weizhe and Mei, Jingbiao and Chen, Jinghong and Byrne, Bill}, 
        year={2024}}
custom_code
feature-extraction
flmr
FLMR
knowledge-based visual question answering
multi-modal
PreFLMR
retrieval
safetensors
transformers