JarvisUSTC/Awesome-Multimodal-RAG

A curated list of the latest advancements, papers, tools, and datasets for **Multimodal Retrieval-Augmented Generation (RAG)**. Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).

54

13 commits

updated Sep 17, 2026

See the code

README

🌟 Awesome Multimodal RAG

A curated list of the latest advancements, papers, tools, and datasets for Multimodal Retrieval-Augmented Generation (RAG). Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).


πŸ“š Contents


✨ Introduction

Multimodal RAG is a cutting-edge approach combining the power of information retrieval and generative models to handle multimodal data. By integrating diverse modalities such as text, images, and audio, Multimodal RAG aims to improve retrieval quality, generate contextually rich outputs, and address complex reasoning tasks. This repository summarizes the latest research, datasets, and tools to foster innovation in this exciting area.


πŸ“ Papers

πŸ“– Surveys and Tutorials

🧠 General Multimodal RAG

πŸ“„ Multimodal Document RAG

πŸ” Domain-Specific Multimodal RAG


πŸ“Š Datasets


πŸ”§ Tools and Frameworks

  • Notable Projects:
    • πŸ”₯ Kiln - Build a RAG in 5 minutes using drag-and-drop. Kiln is a free tool for building production-ready AI systems, supporting RAG pipelines (text, image, audio, video), evaluations, agents, MCP tool-calling, synthetic data generation, and fine-tuning. GitHub ⏰ 2025-11
    • πŸ”¨ Together Cookbook - It is a collection of code and guides designed to help developers build with open source models using Together AI. ⏰ 2024-12
    • πŸ“Š Pixeltable - Declarative multimodal AI data engine supporting document/media chunking, embedding generation, hybrid vector search, and incremental updates for Multimodal RAG. GitHub ⏰ 2024-03

πŸ“ˆ Benchmarks and Metrics

  • Evaluation Metrics:
    • Retrieval Metrics: Precision@k, Recall@k
    • Generation Metrics: BLEU, ROUGE, CIDEr

πŸš€ Open Challenges

  • Key Challenges:
    • Efficient multimodal retrieval at scale
    • Alignment between modalities
    • Handling noisy or incomplete multimodal data
    • Real-time processing for practical applications

🀝 Contributing

Contributions are welcome! Please submit a pull request or open an issue to add new papers, datasets, tools, or corrections.


πŸ™ Acknowledgments

Thanks to the research community for their efforts in advancing Multimodal RAG. If you find this repository useful, please consider starring it!


Significant stargazers

Aojie Zhou

31 followers Β· starred Mar 2025

JarvisUSTC/Awesome-Multimodal-RAG

A curated list of the latest advancements, papers, tools, and datasets for **Multimodal Retrieval-Augmented Generation (RAG)**. Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).

54

13 commits

updated Sep 17, 2026

See the code

README

🌟 Awesome Multimodal RAG

A curated list of the latest advancements, papers, tools, and datasets for Multimodal Retrieval-Augmented Generation (RAG). Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).


πŸ“š Contents


✨ Introduction

Multimodal RAG is a cutting-edge approach combining the power of information retrieval and generative models to handle multimodal data. By integrating diverse modalities such as text, images, and audio, Multimodal RAG aims to improve retrieval quality, generate contextually rich outputs, and address complex reasoning tasks. This repository summarizes the latest research, datasets, and tools to foster innovation in this exciting area.


πŸ“ Papers

πŸ“– Surveys and Tutorials

🧠 General Multimodal RAG

πŸ“„ Multimodal Document RAG

πŸ” Domain-Specific Multimodal RAG


πŸ“Š Datasets


πŸ”§ Tools and Frameworks

  • Notable Projects:
    • πŸ”₯ Kiln - Build a RAG in 5 minutes using drag-and-drop. Kiln is a free tool for building production-ready AI systems, supporting RAG pipelines (text, image, audio, video), evaluations, agents, MCP tool-calling, synthetic data generation, and fine-tuning. GitHub ⏰ 2025-11
    • πŸ”¨ Together Cookbook - It is a collection of code and guides designed to help developers build with open source models using Together AI. ⏰ 2024-12
    • πŸ“Š Pixeltable - Declarative multimodal AI data engine supporting document/media chunking, embedding generation, hybrid vector search, and incremental updates for Multimodal RAG. GitHub ⏰ 2024-03

πŸ“ˆ Benchmarks and Metrics

  • Evaluation Metrics:
    • Retrieval Metrics: Precision@k, Recall@k
    • Generation Metrics: BLEU, ROUGE, CIDEr

πŸš€ Open Challenges

  • Key Challenges:
    • Efficient multimodal retrieval at scale
    • Alignment between modalities
    • Handling noisy or incomplete multimodal data
    • Real-time processing for practical applications

🀝 Contributing

Contributions are welcome! Please submit a pull request or open an issue to add new papers, datasets, tools, or corrections.


πŸ™ Acknowledgments

Thanks to the research community for their efforts in advancing Multimodal RAG. If you find this repository useful, please consider starring it!


Significant stargazers

Aojie Zhou

31 followers Β· starred Mar 2025