A curated list of the latest advancements, papers, tools, and datasets for **Multimodal Retrieval-Augmented Generation (RAG)**. Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).
54
13 commits
updated Sep 17, 2026
A curated list of the latest advancements, papers, tools, and datasets for Multimodal Retrieval-Augmented Generation (RAG). Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).
Multimodal RAG is a cutting-edge approach combining the power of information retrieval and generative models to handle multimodal data. By integrating diverse modalities such as text, images, and audio, Multimodal RAG aims to improve retrieval quality, generate contextually rich outputs, and address complex reasoning tasks. This repository summarizes the latest research, datasets, and tools to foster innovation in this exciting area.
Comprehensive reviews and overviews of Multimodal RAG.
Key Papers:
Tutorials:
Contributions are welcome! Please submit a pull request or open an issue to add new papers, datasets, tools, or corrections.
Thanks to the research community for their efforts in advancing Multimodal RAG. If you find this repository useful, please consider starring it!
31 followers Β· starred Mar 2025
A curated list of the latest advancements, papers, tools, and datasets for **Multimodal Retrieval-Augmented Generation (RAG)**. Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).
54
13 commits
updated Sep 17, 2026
A curated list of the latest advancements, papers, tools, and datasets for Multimodal Retrieval-Augmented Generation (RAG). Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).
Multimodal RAG is a cutting-edge approach combining the power of information retrieval and generative models to handle multimodal data. By integrating diverse modalities such as text, images, and audio, Multimodal RAG aims to improve retrieval quality, generate contextually rich outputs, and address complex reasoning tasks. This repository summarizes the latest research, datasets, and tools to foster innovation in this exciting area.
Comprehensive reviews and overviews of Multimodal RAG.
Key Papers:
Tutorials:
Contributions are welcome! Please submit a pull request or open an issue to add new papers, datasets, tools, or corrections.
Thanks to the research community for their efforts in advancing Multimodal RAG. If you find this repository useful, please consider starring it!
31 followers Β· starred Mar 2025