lmms-lab/textvqa

Dataset

24

stars

3

commits

1

linked in READMEs

Mar 8, 2024

updated

Browse cluster: Multimodal Document and Visual Understanding β†’

README

Large-scale Multi-modality Models Evaluation Suite

Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval

🏠 Homepage | πŸ“š Documentation | πŸ€— Huggingface Datasets

This Dataset

This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.

@inproceedings{singh2019towards,
  title={Towards vqa models that can read},
  author={Singh, Amanpreet and Natarajan, Vivek and Shah, Meet and Jiang, Yu and Chen, Xinlei and Batra, Dhruv and Parikh, Devi and Rohrbach, Marcus},
  booktitle={Proceedings of the IEEE/CVF conference on computer vision and pattern recognition},
  pages={8317--8326},
  year={2019}
}

Contributors

pufanyi

2 commits

luodian

1 commits

lmms-lab/textvqa

Dataset

24

stars

3

commits

1

linked in READMEs

Mar 8, 2024

updated

Browse cluster: Multimodal Document and Visual Understanding β†’

README

Large-scale Multi-modality Models Evaluation Suite

Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval

🏠 Homepage | πŸ“š Documentation | πŸ€— Huggingface Datasets

This Dataset

This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.

@inproceedings{singh2019towards,
  title={Towards vqa models that can read},
  author={Singh, Amanpreet and Natarajan, Vivek and Shah, Meet and Jiang, Yu and Chen, Xinlei and Batra, Dhruv and Parikh, Devi and Rohrbach, Marcus},
  booktitle={Proceedings of the IEEE/CVF conference on computer vision and pattern recognition},
  pages={8317--8326},
  year={2019}
}

Contributors

pufanyi

2 commits

luodian

1 commits