6 repos
Vision-language models and retrieval-augmented generation (RAG) systems for understanding and indexing documents, particularly focused on processing visual document content. The cluster centers on ColPali, a technology for dense retrieval from document images, with supporting benchmarks, reproducibility studies, and implementations. Developers exploring this area will find tools for building document understanding pipelines, evaluation frameworks, and practical applications of vision-language models in information retrieval workflows.