Agentic-RAG explores advanced Retrieval-Augmented Generation systems enhanced with AI LLM agents.
See the code
Overview of Agentic RAG
Recent Update (2025-02-04):
Check section 4 in the table of contents in this repo for the new Agentic Workflow Patterns. New images have been added to enhance the Overview of Agentic RAG. The paper is also updated.
Agentic Retrieval-Augmented Generation ( Agentic RAG) represents a transformative leap in artificial intelligence by embedding autonomous agents into the RAG pipeline. This repository complements the survey paper "Agentic Retrieval-Augmented Generation (Agentic RAG): A Survey On Agentic RAG," providing insights into:
This repository serves as a comprehensive resource for researchers and practitioners to explore, implement, and advance the capabilities of Agentic RAG systems.
Retrieval-Augmented Generation (RAG) systems combine the capabilities of large language models (LLMs) with retrieval mechanisms to generate contextually relevant and accurate responses. While traditional RAG systems excel in knowledge retrieval and generation, they often fall short in handling dynamic, multi-step reasoning tasks, adaptability, and orchestration for complex workflows.
Agentic Retrieval-Augmented Generation (Agentic RAG) overcomes these limitations by integrating autonomous AI agents. These agents employ core Agentic Patterns, such as reflection, planning, tool use, and multi-agent collaboration, to dynamically adapt to task-specific requirements and provide superior performance in:
This repository explores the evolution of RAG to Agentic RAG, presenting:
Whether you’re a researcher, developer, or practitioner, this repository offers valuable insights and resources to understand and advance Agentic RAG systems.
Agentic RAG systems derive their intelligence and adaptability from well-defined agentic patterns. These patterns enable agents to handle complex reasoning tasks, adapt to dynamic environments, and collaborate effectively. Below are the key patterns central to Agentic RAG:
Figure 1: Reflection Pattern
Figure 2: Planning Pattern
Figure 3: Tool Use Pattern
Figure 4: Multi-Agent Collaboration Pattern
These patterns form the backbone of Agentic RAG systems, enabling them to:
Agentic workflow patterns help structure LLM-based applications to optimize performance, accuracy, and efficiency. Different approaches are suitable depending on task complexity and processing requirements.
Source: Anthropic Research and LangGraph Workflows
Figure 1: Illustration of Prompt Chaining Workflow
Figure 2: Illustration of Routing Workflow
Figure 3: Illustration of Parallelization Workflow
Figure 4: Illustration of Orchestrator-Workers Workflow
Figure 5: Illustration of Evaluator-Optimizer Workflow
Agentic Retrieval-Augmented Generation (RAG) systems encompass various architectures and workflows, each tailored to specific tasks and levels of complexity. Below is a detailed taxonomy of these systems:
Key Idea: A team of agents collaborates to perform complex retrieval and reasoning tasks.
Workflow:
Advantages:
Limitations:
AgentFlow is a trainable, tool-integrated agentic framework designed to overcome the scalability and generalization limits of today’s tool-augmented reasoning approaches. It coordinates four specialized modules—Planner, Executor, Verifier, Generator—and optimizes the planner in the flow of multi-turn tasks using Flow-GRPO, improving long-horizon credit assignment and tool-use reliability.
Key Features:
🧩 Modular Agentic System – Four specialized agent modules (Planner, Executor, Verifier, Generator) that coordinate via evolving memory and integrated tools across multiple turns.
🔗 Multi-Tool Integration – Seamlessly connect with diverse tool ecosystems, including base_generator, python_coder, google_search, wikipedia_search, web_search, and more.
🎯 Flow-GRPO Algorithm – Enables in-the-flow agent optimization for long-horizon reasoning tasks with sparse rewards.
Graph-based RAG systems extend traditional RAG by integrating graph-based data structures for advanced reasoning.
Agentic Document Workflows (ADW) extend traditional RAG systems by automating document-centric processes with intelligent agents.
Figure 5: Single-Agent RAG Diagram
Figure 6: Multi-Agent RAG Diagram
Figure 7: Hierarchical RAG Workflow
Figure 8: Graph-Based RAG Workflow
Figure 9: ADW Workflow Diagram [Source]
The table below provides a comprehensive comparative analysis of the three architectural frameworks: Traditional RAG, Agentic RAG, and Agentic Document Workflows (ADW). This analysis highlights their respective strengths, weaknesses, and best-fit scenarios, offering valuable insights into their applicability across diverse use cases.
| Feature | Traditional RAG | Agentic RAG | Agentic Document Workflows (ADW) |
|---|---|---|---|
| Focus | Isolated retrieval and generation tasks | Multi-agent collaboration and reasoning | Document-centric end-to-end workflows |
| Context Maintenance | Limited | Enabled through memory modules | Maintains state across multi-step workflows |
| Dynamic Adaptability | Minimal | High | Tailored to document workflows |
| Workflow Orchestration | Absent | Orchestrates multi-agent tasks | Integrates multi-step document processing |
| Use of External Tools/APIs | Basic integration (e.g., retrieval tools) | Extends via tools like APIs and knowledge bases | Deeply integrates business rules and domain-specific tools |
| Scalability | Limited to small datasets or queries | Scalable for multi-agent systems | Scales for multi-domain enterprise workflows |
| Complex Reasoning | Basic (e.g., simple Q&A) | Multi-step reasoning with agents | Structured reasoning across documents |
| Primary Applications | QA systems, knowledge retrieval | Multi-domain knowledge and reasoning | Contract review, invoice processing, claims analysis |
| Strengths | Simplicity, quick setup | High accuracy, collaborative reasoning | End-to-end automation, domain-specific intelligence |
| Challenges | Poor contextual understanding | Coordination complexity | Resource overhead, domain standardization |
Agentic Retrieval-Augmented Generation (RAG) systems have transformative potential across diverse industries, enabling intelligent retrieval, multi-step reasoning, and dynamic adaptation to complex tasks. Below are some key domains where Agentic RAG systems make a significant impact:
While Agentic Retrieval-Augmented Generation (RAG) systems show immense promise, there are several challenges and research opportunities that remain unaddressed:
Coordination Complexity in Multi-Agent Systems:
Ethical and Responsible AI:
Scalability and Latency:
Hybrid Human-Agent Collaboration:
Expanding Multimodal Capabilities:
Enhanced Agentic Orchestration:
Domain-Specific Applications:
Ethical AI and Governance Frameworks:
Efficient Graph-Based Reasoning:
Human-AI Synergy:
| Technique | Tools | Description | Notebooks |
|---|---|---|---|
| Single Agentic RAG | LangChain, FAISS, Athina AI | Uses AI agents to find and generate answers using tools like vectordb and web searches. | View Notebook |
| LlamaIndex, Vertex AI (Vector Store, Text Embedding, LLM), Google Cloud Storage | Demonstrates a single-router Agentic RAG system using LlamaIndex with Vertex AI for context retrieval and response generation. | View Notebook | |
| LangChain, IBM Granite-3-8B-Instruct, Watsonx.ai, Chroma DB, WebBaseLoader | Builds an Agentic RAG system using IBM Granite-3-8B-Instruct model in Watsonx.ai to answer complex queries with external information. | View Notebook | |
| LangGraph, Chroma, NVIDIA Inference Microservices (NIMs), Tavily Search API | This system uses a router-based architecture to determine whether a query should be handled by a RAG pipeline (retrieving from a vector database) or a websearch pipeline. An AI agent evaluates the query's topic and routes it to the appropriate pipeline for information retrieval and response generation, ensuring accurate, relevant, and contextually augmented answers. | View Notebook | |
| LlamaIndex, Redis, Amazon Bedrock, RedisVectorStore, LlamaParse, BedrockEmbedding, SemanticCache | This system implements a ReAct agent-based RAG pipeline where the agent interacts with a Redis-backed index and vector store to retrieve and process data from a PDF document. It utilizes Amazon Bedrock embeddings and LlamaIndex to process the document, build embeddings, and handle retrieval-based augmented generation. Additionally, semantic caching optimizes the system by reducing redundant LLM queries for repeated or similar user questions, improving response times and efficiency. | View Notebook | |
| Multi-Agent Agentic RAG Orchestrator | AutoGen, SQL, AI Search Indexes | This orchestrator utilizes a multi-agent system to facilitate complex task execution through coordinated agent interactions. Using a factory pattern and various predefined strategies (e.g., classic_rag for retrieval-augmented generation and nl2sql for translating natural language to SQL), the system enables flexible, multi-agent collaboration for tasks like database querying and document retrieval. The orchestrator supports agent communication, iterative responses, and customizable strategies, offering a high level of adaptability for diverse use cases. | View Notebook |
| Hierarchical Multi-Agent Agentic RAG | Weaviate, ExaSearch, Groq, crewAI | This approach uses a hierarchical agentic architecture with multiple agents, each responsible for specific tasks or tools. A manager agent coordinates the work of specialized agents (such as WeaviateTool for internal document retrieval, ExaSearchTool for web searches, and Groq for fast AI inference) to handle complex queries. The flexible, task-oriented system can support various use cases such as QA and workflow automation. | View Notebook |
| Corrective RAG | LangChain, LangGraph, Chromadb, Athina AI | Refines relevant documents, removes irrelevant ones or does the web search. | View Notebook |
| LangChain, FAISS, HuggingFace Inference API, SmolAgents, HyDE, Self-Query | This system incorporates query reformulation and self-query strategies to address limitations in traditional RAG systems. It performs iterative retrieval by critiquing the relevance of retrieved documents and re-querying as needed. The agent refines queries to improve semantic similarity and ensure higher accuracy. Self-grading mechanisms assess the quality of retrieved information, enhancing results through iterative improvement. The system aligns with Corrective RAG principles by reducing confabulations and improving retrieval relevance. | View Notebook | |
| Adaptive RAG | LangChain, LangGraph, FAISS, Athina AI | Adjusts retrieval methods based on query type, using indexed data or web search. | View Notebook |
| ReAct RAG | LangChain, LangGraph, FAISS, Athina AI | System combining reasoning and retrieval for context-aware responses | |
| Self RAG | LangChain, LangGraph, FAISS, Athina AI | Reflects on retrieved data to ensure accurate and complete responses. |
DeepLearning.AI: How agents can improve LLM performance. DeepLearning.AI
Weaviate Blog: What is agentic RAG? Weaviate Blog
LangGraph CRAG Tutorial: LangGraph CRAG: Contextualized retrieval-augmented generation tutorial. LangGraph CRAG
LangGraph Adaptive RAG Tutorial: LangGraph adaptive RAG: Adaptive retrieval-augmented generation tutorial. LangGraph Adaptive RAG. Accessed: 2025-01-14.
LlamaIndex Blog: Agentic RAG with LlamaIndex. LlamaIndex Blog
Hugging Face Cookbook. Agentic RAG: Turbocharge your retrieval-augmented generation with query reformulation and self-query. Hugging Face Cookbook
Hugging Face Agentic RAG: https://huggingface.co/docs/smolagents/en/examples/rag
Qdrant Blog. Agentic RAG: Combining RAG with agents for enhanced information retrieval. Qdrant Blog
Semantic Kernel: Semantic Kernel is an open-source SDK by Microsoft that integrates large language models (LLMs) into applications. It supports agentic patterns, enabling the creation of autonomous AI agents for natural language understanding, task automation, and decision-making. It has been used in scenarios like ServiceNow’s P1 incident management to facilitate real-time collaboration, automate task execution, and retrieve contextual information seamlessly.
Below are some noteworthy resources related to Agentic Design Patterns. The first five items are from Andrew Ng’s series at DeepLearning.ai:
Additional Resources
If you find this work useful in your research, please cite:
@misc{singh2025agenticretrievalaugmentedgenerationsurvey,
title={Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG},
author={Aditi Singh and Abul Ehtesham and Saket Kumar and Tala Talaei Khoei},
year={2025},
eprint={2501.09136},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2501.09136},
}
170 followers · starred Aug 2025
124 followers · starred Mar 2025
14 followers · starred Aug 2026
34 followers · starred Jan 2025
Agentic-RAG explores advanced Retrieval-Augmented Generation systems enhanced with AI LLM agents.
See the code
Overview of Agentic RAG
Recent Update (2025-02-04):
Check section 4 in the table of contents in this repo for the new Agentic Workflow Patterns. New images have been added to enhance the Overview of Agentic RAG. The paper is also updated.
Agentic Retrieval-Augmented Generation ( Agentic RAG) represents a transformative leap in artificial intelligence by embedding autonomous agents into the RAG pipeline. This repository complements the survey paper "Agentic Retrieval-Augmented Generation (Agentic RAG): A Survey On Agentic RAG," providing insights into:
This repository serves as a comprehensive resource for researchers and practitioners to explore, implement, and advance the capabilities of Agentic RAG systems.
Retrieval-Augmented Generation (RAG) systems combine the capabilities of large language models (LLMs) with retrieval mechanisms to generate contextually relevant and accurate responses. While traditional RAG systems excel in knowledge retrieval and generation, they often fall short in handling dynamic, multi-step reasoning tasks, adaptability, and orchestration for complex workflows.
Agentic Retrieval-Augmented Generation (Agentic RAG) overcomes these limitations by integrating autonomous AI agents. These agents employ core Agentic Patterns, such as reflection, planning, tool use, and multi-agent collaboration, to dynamically adapt to task-specific requirements and provide superior performance in:
This repository explores the evolution of RAG to Agentic RAG, presenting:
Whether you’re a researcher, developer, or practitioner, this repository offers valuable insights and resources to understand and advance Agentic RAG systems.
Agentic RAG systems derive their intelligence and adaptability from well-defined agentic patterns. These patterns enable agents to handle complex reasoning tasks, adapt to dynamic environments, and collaborate effectively. Below are the key patterns central to Agentic RAG:
Figure 1: Reflection Pattern
Figure 2: Planning Pattern
Figure 3: Tool Use Pattern
Figure 4: Multi-Agent Collaboration Pattern
These patterns form the backbone of Agentic RAG systems, enabling them to:
Agentic workflow patterns help structure LLM-based applications to optimize performance, accuracy, and efficiency. Different approaches are suitable depending on task complexity and processing requirements.
Source: Anthropic Research and LangGraph Workflows
Figure 1: Illustration of Prompt Chaining Workflow
Figure 2: Illustration of Routing Workflow
Figure 3: Illustration of Parallelization Workflow
Figure 4: Illustration of Orchestrator-Workers Workflow
Figure 5: Illustration of Evaluator-Optimizer Workflow
Agentic Retrieval-Augmented Generation (RAG) systems encompass various architectures and workflows, each tailored to specific tasks and levels of complexity. Below is a detailed taxonomy of these systems:
Key Idea: A team of agents collaborates to perform complex retrieval and reasoning tasks.
Workflow:
Advantages:
Limitations:
AgentFlow is a trainable, tool-integrated agentic framework designed to overcome the scalability and generalization limits of today’s tool-augmented reasoning approaches. It coordinates four specialized modules—Planner, Executor, Verifier, Generator—and optimizes the planner in the flow of multi-turn tasks using Flow-GRPO, improving long-horizon credit assignment and tool-use reliability.
Key Features:
🧩 Modular Agentic System – Four specialized agent modules (Planner, Executor, Verifier, Generator) that coordinate via evolving memory and integrated tools across multiple turns.
🔗 Multi-Tool Integration – Seamlessly connect with diverse tool ecosystems, including base_generator, python_coder, google_search, wikipedia_search, web_search, and more.
🎯 Flow-GRPO Algorithm – Enables in-the-flow agent optimization for long-horizon reasoning tasks with sparse rewards.
Graph-based RAG systems extend traditional RAG by integrating graph-based data structures for advanced reasoning.
Agentic Document Workflows (ADW) extend traditional RAG systems by automating document-centric processes with intelligent agents.
Figure 5: Single-Agent RAG Diagram
Figure 6: Multi-Agent RAG Diagram
Figure 7: Hierarchical RAG Workflow
Figure 8: Graph-Based RAG Workflow
Figure 9: ADW Workflow Diagram [Source]
The table below provides a comprehensive comparative analysis of the three architectural frameworks: Traditional RAG, Agentic RAG, and Agentic Document Workflows (ADW). This analysis highlights their respective strengths, weaknesses, and best-fit scenarios, offering valuable insights into their applicability across diverse use cases.
| Feature | Traditional RAG | Agentic RAG | Agentic Document Workflows (ADW) |
|---|---|---|---|
| Focus | Isolated retrieval and generation tasks | Multi-agent collaboration and reasoning | Document-centric end-to-end workflows |
| Context Maintenance | Limited | Enabled through memory modules | Maintains state across multi-step workflows |
| Dynamic Adaptability | Minimal | High | Tailored to document workflows |
| Workflow Orchestration | Absent | Orchestrates multi-agent tasks | Integrates multi-step document processing |
| Use of External Tools/APIs | Basic integration (e.g., retrieval tools) | Extends via tools like APIs and knowledge bases | Deeply integrates business rules and domain-specific tools |
| Scalability | Limited to small datasets or queries | Scalable for multi-agent systems | Scales for multi-domain enterprise workflows |
| Complex Reasoning | Basic (e.g., simple Q&A) | Multi-step reasoning with agents | Structured reasoning across documents |
| Primary Applications | QA systems, knowledge retrieval | Multi-domain knowledge and reasoning | Contract review, invoice processing, claims analysis |
| Strengths | Simplicity, quick setup | High accuracy, collaborative reasoning | End-to-end automation, domain-specific intelligence |
| Challenges | Poor contextual understanding | Coordination complexity | Resource overhead, domain standardization |
Agentic Retrieval-Augmented Generation (RAG) systems have transformative potential across diverse industries, enabling intelligent retrieval, multi-step reasoning, and dynamic adaptation to complex tasks. Below are some key domains where Agentic RAG systems make a significant impact:
While Agentic Retrieval-Augmented Generation (RAG) systems show immense promise, there are several challenges and research opportunities that remain unaddressed:
Coordination Complexity in Multi-Agent Systems:
Ethical and Responsible AI:
Scalability and Latency:
Hybrid Human-Agent Collaboration:
Expanding Multimodal Capabilities:
Enhanced Agentic Orchestration:
Domain-Specific Applications:
Ethical AI and Governance Frameworks:
Efficient Graph-Based Reasoning:
Human-AI Synergy:
| Technique | Tools | Description | Notebooks |
|---|---|---|---|
| Single Agentic RAG | LangChain, FAISS, Athina AI | Uses AI agents to find and generate answers using tools like vectordb and web searches. | View Notebook |
| LlamaIndex, Vertex AI (Vector Store, Text Embedding, LLM), Google Cloud Storage | Demonstrates a single-router Agentic RAG system using LlamaIndex with Vertex AI for context retrieval and response generation. | View Notebook | |
| LangChain, IBM Granite-3-8B-Instruct, Watsonx.ai, Chroma DB, WebBaseLoader | Builds an Agentic RAG system using IBM Granite-3-8B-Instruct model in Watsonx.ai to answer complex queries with external information. | View Notebook | |
| LangGraph, Chroma, NVIDIA Inference Microservices (NIMs), Tavily Search API | This system uses a router-based architecture to determine whether a query should be handled by a RAG pipeline (retrieving from a vector database) or a websearch pipeline. An AI agent evaluates the query's topic and routes it to the appropriate pipeline for information retrieval and response generation, ensuring accurate, relevant, and contextually augmented answers. | View Notebook | |
| LlamaIndex, Redis, Amazon Bedrock, RedisVectorStore, LlamaParse, BedrockEmbedding, SemanticCache | This system implements a ReAct agent-based RAG pipeline where the agent interacts with a Redis-backed index and vector store to retrieve and process data from a PDF document. It utilizes Amazon Bedrock embeddings and LlamaIndex to process the document, build embeddings, and handle retrieval-based augmented generation. Additionally, semantic caching optimizes the system by reducing redundant LLM queries for repeated or similar user questions, improving response times and efficiency. | View Notebook | |
| Multi-Agent Agentic RAG Orchestrator | AutoGen, SQL, AI Search Indexes | This orchestrator utilizes a multi-agent system to facilitate complex task execution through coordinated agent interactions. Using a factory pattern and various predefined strategies (e.g., classic_rag for retrieval-augmented generation and nl2sql for translating natural language to SQL), the system enables flexible, multi-agent collaboration for tasks like database querying and document retrieval. The orchestrator supports agent communication, iterative responses, and customizable strategies, offering a high level of adaptability for diverse use cases. | View Notebook |
| Hierarchical Multi-Agent Agentic RAG | Weaviate, ExaSearch, Groq, crewAI | This approach uses a hierarchical agentic architecture with multiple agents, each responsible for specific tasks or tools. A manager agent coordinates the work of specialized agents (such as WeaviateTool for internal document retrieval, ExaSearchTool for web searches, and Groq for fast AI inference) to handle complex queries. The flexible, task-oriented system can support various use cases such as QA and workflow automation. | View Notebook |
| Corrective RAG | LangChain, LangGraph, Chromadb, Athina AI | Refines relevant documents, removes irrelevant ones or does the web search. | View Notebook |
| LangChain, FAISS, HuggingFace Inference API, SmolAgents, HyDE, Self-Query | This system incorporates query reformulation and self-query strategies to address limitations in traditional RAG systems. It performs iterative retrieval by critiquing the relevance of retrieved documents and re-querying as needed. The agent refines queries to improve semantic similarity and ensure higher accuracy. Self-grading mechanisms assess the quality of retrieved information, enhancing results through iterative improvement. The system aligns with Corrective RAG principles by reducing confabulations and improving retrieval relevance. | View Notebook | |
| Adaptive RAG | LangChain, LangGraph, FAISS, Athina AI | Adjusts retrieval methods based on query type, using indexed data or web search. | View Notebook |
| ReAct RAG | LangChain, LangGraph, FAISS, Athina AI | System combining reasoning and retrieval for context-aware responses | |
| Self RAG | LangChain, LangGraph, FAISS, Athina AI | Reflects on retrieved data to ensure accurate and complete responses. |
DeepLearning.AI: How agents can improve LLM performance. DeepLearning.AI
Weaviate Blog: What is agentic RAG? Weaviate Blog
LangGraph CRAG Tutorial: LangGraph CRAG: Contextualized retrieval-augmented generation tutorial. LangGraph CRAG
LangGraph Adaptive RAG Tutorial: LangGraph adaptive RAG: Adaptive retrieval-augmented generation tutorial. LangGraph Adaptive RAG. Accessed: 2025-01-14.
LlamaIndex Blog: Agentic RAG with LlamaIndex. LlamaIndex Blog
Hugging Face Cookbook. Agentic RAG: Turbocharge your retrieval-augmented generation with query reformulation and self-query. Hugging Face Cookbook
Hugging Face Agentic RAG: https://huggingface.co/docs/smolagents/en/examples/rag
Qdrant Blog. Agentic RAG: Combining RAG with agents for enhanced information retrieval. Qdrant Blog
Semantic Kernel: Semantic Kernel is an open-source SDK by Microsoft that integrates large language models (LLMs) into applications. It supports agentic patterns, enabling the creation of autonomous AI agents for natural language understanding, task automation, and decision-making. It has been used in scenarios like ServiceNow’s P1 incident management to facilitate real-time collaboration, automate task execution, and retrieve contextual information seamlessly.
Below are some noteworthy resources related to Agentic Design Patterns. The first five items are from Andrew Ng’s series at DeepLearning.ai:
Additional Resources
If you find this work useful in your research, please cite:
@misc{singh2025agenticretrievalaugmentedgenerationsurvey,
title={Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG},
author={Aditi Singh and Abul Ehtesham and Saket Kumar and Tala Talaei Khoei},
year={2025},
eprint={2501.09136},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2501.09136},
}
170 followers · starred Aug 2025
124 followers · starred Mar 2025
14 followers · starred Aug 2026
34 followers · starred Jan 2025