Vision-Language Models for VQA & VCR

10 repos

Pre-trained vision-language models (VLE) built on ViLT for visual question answering and visual commonsense reasoning tasks. The cluster centers on PyTorch-based transformer implementations optimized for endpoints-compatible deployment, with model variants spanning base and large architectures tuned for different downstream tasks like VQA and visual commonsense reasoning question-answering.

Python · 1
endpoints_compatible ·437
pytorch ·437
transformers ·437
vilt ·424
visual-question-answering ·424
vle ·209
nlp ·196
language ·196
llm ·196
multi-modal ·196