{"id":170681,"date":"2025-04-04T09:10:23","date_gmt":"2025-04-04T09:10:23","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/a-comprehensive-guide-to-rag-developer-stack\/"},"modified":"2025-04-04T09:10:23","modified_gmt":"2025-04-04T09:10:23","slug":"a-comprehensive-guide-to-rag-developer-stack","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=170681","title":{"rendered":"A Comprehensive Guide to RAG Developer Stack"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Building a RAG (Retrieval-Augmented Generation) application isn\u2019t just about plugging in a few tools\u2014it\u2019s about choosing the right stack that makes retrieval and generation not just possible but <em>efficient and scalable<\/em>.<\/p>\n<p>Let\u2019s say you\u2019re working on something like <em>\u201cSmart Chat with PDF\u201d<\/em>\u2014an AI app that lets users interact with PDFs conversationally. It\u2019s not as simple as just loading a file and asking questions. You need to:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Extract relevant content<\/strong> from the PDF<\/li>\n<li><strong>Chunk the text<\/strong> into meaningful pieces<\/li>\n<li><strong>Store those chunks<\/strong> in a vector database<\/li>\n<li>Then, when a user asks something, the app runs a <strong>similarity search<\/strong>, fetches the most relevant chunks, and passes them to the language model to generate a coherent and accurate response<\/li>\n<\/ol>\n<p>Sounds like a lot? It is. Working across multiple tools, frameworks, and databases can get overwhelming fast.<\/p>\n<p>That\u2019s exactly why I created the RAG Developer\u2019s Stack\u2014a curated set of tools and frameworks designed to streamline this whole process. From smart data extractors to efficient vector databases and cost-effective generation models, it\u2019s everything you need to build robust, production-ready RAG applications without reinventing the wheel every time.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-why-you-need-rag-developer-stack\">Why You Need RAG Developer Stack?<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"2444\" height=\"2106\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17.png\" alt=\"Rag architecture\" class=\"wp-image-229433\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17.png 2444w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17-300x259.png 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17-768x662.png 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17-1536x1324.png 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17-2048x1765.png 2048w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-17-150x129.png 150w\" sizes=\"(max-width: 2444px) 100vw, 2444px\"\/><figcaption class=\"wp-element-caption\">Source: Hugging Face<\/figcaption><\/figure>\n<\/div>\n<p>Firstly, here is a brief on RAG \u2013 Retrieval-Augmented Generation (RAG) enhance the capabilities of large language models (LLMs) by integrating external information retrieval mechanisms. This approach allows LLMs to generate more accurate, contextually relevant, and factually grounded responses by supplementing their static training data with up-to-date or domain-specific information.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-does-rag-work\">How does RAG work?<\/h3>\n<p>RAG operates in four key stages:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Indexing<\/strong>: Data from external sources (e.g., documents, databases) is converted into vector representations (embeddings) and stored in a vector database. This enables efficient retrieval of relevant information.<\/li>\n<li><strong>Retrieval<\/strong>: When a user submits a query, the system retrieves the most relevant data from the indexed sources using similarity-based search techniques.<\/li>\n<li><strong>Augmentation<\/strong>: The retrieved information is combined with the user\u2019s query through prompt engineering, effectively \u201caugmenting\u201d the input to the LLM.<\/li>\n<li><strong>Generation<\/strong>: The LLM uses both its internal knowledge and the augmented prompt to produce a response. This process ensures that the output is informed by both pre-trained data and real-time, authoritative sources. <\/li>\n<\/ol>\n<p>Now, why do you need a RAG developer stack?<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-why-do-you-need-a-rag-developer-stack\">Why Do You Need a RAG Developer Stack?<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Accelerate Development:<\/strong> Leverage pre-built, ready-to-integrate components to move from prototype to production faster.<\/li>\n<li><strong>Boost Accuracy:<\/strong> Retrieve real-time, context-relevant data to ground responses and reduce hallucinations.<\/li>\n<li><strong>Strengthen Deployment:<\/strong> Built-in tools enhance security, observability, and scalability, making production readiness a smoother ride.<\/li>\n<li><strong>Maximize Flexibility:<\/strong> Modular design lets you mix and match tools, adapting to the unique demands of different industries and use cases.<\/li>\n<li><strong>Customizable by Design:<\/strong> Developers can hand-pick components that fit their workflow, architecture, and performance goals.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-rag-developer-stack-for-your-next-project\">RAG Developer Stack for Your Next Project<\/h2>\n<p>Here are 9 things you should know to develop RAG Projects:<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-1-large-language-models-llms\">1. Large Language Models (LLMs)<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"976\" height=\"476\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-1-5.webp\" alt=\"LLMs\" class=\"wp-image-229527\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-1-5.webp 976w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-1-5-300x146.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-1-5-768x375.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-1-5-150x73.webp 150w\" sizes=\"auto, (max-width: 976px) 100vw, 976px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>LLMs are the brains of RAG systems, leveraging transformer-based architectures to generate coherent and contextually relevant text. These models come in two categories:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Open-source LLMs<\/strong>: Examples include LLaMA, Falcon, Cohere and more, which allow customization and local deployment.<\/li>\n<li><strong>Closed LLMs<\/strong>: Proprietary models like GPT-4 and Bard offer advanced capabilities but are typically accessible via APIs.<\/li>\n<\/ul>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"858\" height=\"732\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-2-9.webp\" alt=\"Large language model\" class=\"wp-image-229529\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-2-9.webp 858w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-2-9-300x256.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-2-9-768x655.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-2-9-150x128.webp 150w\" sizes=\"auto, (max-width: 858px) 100vw, 858px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-example-of-llm-usage-in-rag\">Example of LLM Usage in RAG<\/h3>\n<p>I have already imported the JSON Documents using the JSON Loader and here is the pipeline for understanding how LLM is used in RAG.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-prompt-template\">Prompt Template<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_core.prompts import ChatPromptTemplate\nrag_prompt = \"\"\"You are an assistant who is an expert in question-answering tasks.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Answer the following question using only the following pieces of retrieved context.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0If the answer is not in the context, do not make up answers, just say that you don't know.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Keep the answer detailed and well formatted based on the information from the context.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Question:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{question}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Context:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{context}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Answer:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"\"\"\nrag_prompt_template = ChatPromptTemplate.from_template(rag_prompt)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-pipeline-construction\">Pipeline Construction<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_core.runnables import RunnablePassthrough\nfrom langchain_openai import ChatOpenAI\n# Initialize ChatGPT model\nchatgpt = ChatOpenAI(model_name=\"gpt-4o-mini\", temperature=0)\n# Format documents into a single string\ndef format_docs(docs):\n\u00a0\u00a0\u00a0\u00a0return \"\\n\\n\".join(doc.page_content for doc in docs)\n# Construct the RAG pipeline\nqa_rag_chain = (\n\u00a0\u00a0\u00a0\u00a0{\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"context\": (similarity_retriever | format_docs),\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"question\": RunnablePassthrough()\n\u00a0\u00a0\u00a0\u00a0}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0|\n\u00a0\u00a0\u00a0\u00a0rag_prompt_template\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0|\n\u00a0\u00a0\u00a0\u00a0chatgpt\n)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-example-usage\">Example Usage<\/h4>\n<pre class=\"wp-block-code\"><code>query = \"What is the difference between AI, ML, and DL?\"\nresult = qa_rag_chain.invoke(query)\n# Display the generated answer\nfrom IPython.display import display, Markdown\ndisplay(Markdown(result.content))<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output\">Output<\/h4>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1039\" height=\"482\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-3-5.webp\" alt=\"Output\" class=\"wp-image-229531\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-3-5.webp 1039w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-3-5-300x139.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-3-5-768x356.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-3-5-150x70.webp 150w\" sizes=\"auto, (max-width: 1039px) 100vw, 1039px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-2-llms-used-in-response-generation-for-rag\">2. LLMs Used in Response Generation for RAG<\/h2>\n<p>In <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/09\/retrieval-augmented-generation-rag-in-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">Retrieval-Augmented <\/a>Generation (RAG) systems, the response generation LLM plays an important role as the final decision-maker \u2014 it takes the retrieved documents, user query, and context and synthesizes everything into a coherent, relevant, and often conversational response. While retrieval models bring in potentially useful information, the LLM can reason, summarize, and contextualize, which ensures the output feels intelligent and human-like. <br \/>A strong response model can filter noisy or partial information, infer unstated connections, and deliver answers that align with user intent. This is especially critical in applications like enterprise search, customer support, legal\/medical assistants, and technical Q&amp;A, where users expect precise, grounded, and trustworthy responses.<\/p>\n<p>In a nutshell, without a capable generation model, even the best retrieval stack falls flat \u2014 making this component the core brain of any RAG pipeline.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-commercial-llms\">Commercial LLMs<\/h3>\n<figure class=\"wp-block-table\">\n<table class=\"table table-bordered border-black table-striped\">\n<thead>\n<tr>\n<th><strong>Model<\/strong><\/th>\n<th><strong>Developer<\/strong><\/th>\n<th><strong>Key Strengths<\/strong><\/th>\n<th><strong>Common Use Cases<\/strong><\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>GPT-4.5<\/strong><\/td>\n<td>OpenAI<\/td>\n<td>Advanced text generation, summarization, conversational fluency<\/td>\n<td>Chatbots, customer support, content creation<\/td>\n<\/tr>\n<tr>\n<td><strong>Claude 3.7 Sonnet<\/strong><\/td>\n<td>Anthropic<\/td>\n<td>Real-time conversations, strong reasoning, \u201cextended thinking mode\u201d<\/td>\n<td>Business automation, customer service<\/td>\n<\/tr>\n<tr>\n<td><strong>Gemini 2.0 Pro<\/strong><\/td>\n<td>Google DeepMind<\/td>\n<td>Multimodal (text + image), high performance<\/td>\n<td>Data analysis, enterprise automation, content generation<\/td>\n<\/tr>\n<tr>\n<td><strong>Cohere Command R+<\/strong><\/td>\n<td>Cohere<\/td>\n<td>Retrieval-Augmented Generation (RAG), enterprise-grade design<\/td>\n<td>Knowledge management, support automation, moderation<\/td>\n<\/tr>\n<tr>\n<td><strong>DeepSeek<\/strong><\/td>\n<td>DeepSeek AI<\/td>\n<td>On-premise deployment, secure data handling, high customizability<\/td>\n<td>Finance, healthcare, privacy-sensitive industries<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h3 class=\"wp-block-heading\" id=\"h-open-source-llms\">Open-Source LLMs<\/h3>\n<figure class=\"wp-block-table\">\n<table class=\"table table-bordered border-black table-striped\">\n<thead>\n<tr>\n<th><strong>Model<\/strong><\/th>\n<th><strong>Developer<\/strong><\/th>\n<th><strong>Key Strengths<\/strong><\/th>\n<th><strong>Common Use Cases<\/strong><\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>LLaMA 3<\/strong><\/td>\n<td>Meta<\/td>\n<td>Scalable (up to 405B params), multimodal capabilities<\/td>\n<td>Conversational AI, research, content generation<\/td>\n<\/tr>\n<tr>\n<td><strong>Mistral 7B<\/strong><\/td>\n<td>Mistral AI<\/td>\n<td>Lightweight yet powerful, optimized for code and chat<\/td>\n<td>Code generation, chatbots, content automation<\/td>\n<\/tr>\n<tr>\n<td><strong>Falcon 180B<\/strong><\/td>\n<td>Technology Innovation Institute<\/td>\n<td>Efficient, high-performance, open-access<\/td>\n<td>Real-time applications, science\/research bots<\/td>\n<\/tr>\n<tr>\n<td><strong>DeepSeek R1<\/strong><\/td>\n<td>DeepSeek AI<\/td>\n<td>Strong logic\/reasoning, 128K context window<\/td>\n<td>Math tasks, summarization, complex reasoning<\/td>\n<\/tr>\n<tr>\n<td><strong> Qwen2.5-72B-Instruct<br \/><\/strong><\/td>\n<td>Alibaba Cloud<\/td>\n<td>72.7 billion parameters, supporting long contexts up to 128K tokens.<br \/>coding, mathematical reasoning, and multilingual support.<\/td>\n<td>Generates structured outputs like JSON, making it highly versatile for technical applications in RAG workflows.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-3-frameworks\">3. Frameworks<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1012\" height=\"541\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-4-2.webp\" alt=\"Frameworks\" class=\"wp-image-229534\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-4-2.webp 1012w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-4-2-300x160.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-4-2-768x411.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-4-2-150x80.webp 150w\" sizes=\"auto, (max-width: 1012px) 100vw, 1012px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>The Frameworks simplify the development of <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/08\/self-hosting-rag-applications-on-edge-devices-with-langchain-and-ollama-part-ii\/\" target=\"_blank\" rel=\"noreferrer noopener\">RAG applications<\/a> by providing pre-built components:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong><a href=\"https:\/\/www.langchain.com\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LangChain<\/a><\/strong>: Framework for LLM application development with modular architecture for prompt management, chaining, memory handling, and agent creation. Excels at building RAG pipelines with built-in support for document loaders, retrievers, and vector stores.<\/li>\n<li><strong><a href=\"https:\/\/www.llamaindex.ai\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LlamaIndex<\/a><\/strong>: Specialized framework for data indexing and retrieval, connecting unstructured data with language models through custom indices. Optimized for ingesting, transforming, and querying large datasets for chatbots and knowledge management.<\/li>\n<li><strong><a href=\"https:\/\/www.langchain.com\/langgraph\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LangGraph<\/a><\/strong>: It integrates LLMs with graph-based structures, allowing developers to define application logic using nodes and edges. Ideal for complex workflows with multiple branches and feedback loops, especially in multi-agent systems.<\/li>\n<li><strong><a href=\"https:\/\/ragflow.io\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">RAGFlow<\/a><\/strong>: A Framework specifically for Retrieval-Augmented Generation systems, orchestrating retrievers, rankers, and generators into coherent pipelines. Enhances relevance when pulling from external data sources for search-driven interfaces and Q&amp;A systems.<\/li>\n<\/ul>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1017\" height=\"750\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-5-5.webp\" alt=\"Rag Frameworks\" class=\"wp-image-229536\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-5-5.webp 1017w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-5-5-300x221.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-5-5-768x566.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-5-5-150x111.webp 150w\" sizes=\"auto, (max-width: 1017px) 100vw, 1017px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>Frameworks like LangChain, LangGraph, and LlamaIndex significantly streamline RAG (Retrieval-Augmented Generation) development by offering modular tools for integrating retrieval and generation processes. LangChain simplifies chaining LLM calls, managing prompts, and connecting to vector stores. LangGraph introduces graph-based flow control, enabling dynamic and multi-step RAG workflows. LlamaIndex focuses on data ingestion, indexing, and retrieval, making large datasets queryable by LLMs. Together, they abstract away complex infrastructure, allowing developers to focus on logic and data quality. These tools enable rapid prototyping and robust deployment of RAG applications for tasks like question answering, document search, and knowledge assistance.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-example-of-frameworks-for-rag-building\">Example of Frameworks for RAG Building<\/h3>\n<p>Let\u2019s build a simple RAG using LangChain:<\/p>\n<pre class=\"wp-block-code\"><code>%pip install --quiet --upgrade langchain-text-splitters langchain-community langgraph\n!pip install -qU \"langchain[openai]\"<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-chat-model\">Chat model<\/h4>\n<pre class=\"wp-block-code\"><code>import getpass\nimport os\n\nif not os.environ.get(\"OPENAI_API_KEY\"):\n  os.environ[\"OPENAI_API_KEY\"] = getpass.getpass(\"Enter API key for OpenAI: \")\n\nfrom langchain.chat_models import init_chat_model\n\nllm = init_chat_model(\"gpt-4o-mini\", model_provider=\"openai\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-select-embeddings-model\">Select embeddings model<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_openai import OpenAIEmbeddings\nembeddings = OpenAIEmbeddings(model=\"text-embedding-3-large\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-select-vector-store\">Select vector store<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_core.vectorstores import InMemoryVectorStore\nvector_store = InMemoryVectorStore(embeddings)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-creating-the-indexing-pipeline\">Creating the indexing pipeline<\/h4>\n<pre class=\"wp-block-code\"><code>import bs4\nfrom langchain import hub\nfrom langchain_community.document_loaders import WebBaseLoader\nfrom langchain_core.documents import Document\nfrom langchain_text_splitters import RecursiveCharacterTextSplitter\nfrom langgraph.graph import START, StateGraph\nfrom typing_extensions import List, TypedDict\n\n# Load and chunk contents of the blog\nloader = WebBaseLoader(\n    web_paths=(\"https:\/\/lilianweng.github.io\/posts\/2023-06-23-agent\/\",),\n    bs_kwargs=dict(\n        parse_only=bs4.SoupStrainer(\n            class_=(\"post-content\", \"post-title\", \"post-header\")\n        )\n    ),\n)\ndocs = loader.load()\n\ntext_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)\nall_splits = text_splitter.split_documents(docs)\n\n# Index chunks\n_ = vector_store.add_documents(documents=all_splits)\n\n# Define prompt for question-answering\nprompt = hub.pull(\"rlm\/rag-prompt\")\n\n\n# Define state for application\nclass State(TypedDict):\n    question: str\n    context: List[Document]\n    answer: str\n\n\n# Define application steps\ndef retrieve(state: State):\n    retrieved_docs = vector_store.similarity_search(state[\"question\"])\n    return {\"context\": retrieved_docs}\n\n\ndef generate(state: State):\n    docs_content = \"\\n\\n\".join(doc.page_content for doc in state[\"context\"])\n    messages = prompt.invoke({\"question\": state[\"question\"], \"context\": docs_content})\n    response = llm.invoke(messages)\n    return {\"answer\": response.content}\n\n\n# Compile application and test\ngraph_builder = StateGraph(State).add_sequence([retrieve, generate])\ngraph_builder.add_edge(START, \"retrieve\")\ngraph = graph_builder.compile()<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>response = graph.invoke({\"question\": \"What are Types of Memory?\"})\nprint(response[\"answer\"])<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output-0\">Output<\/h4>\n<pre class=\"wp-block-preformatted\">The types of memory include Sensory Memory, Short-Term Memory (STM), and<br\/>Long-Term Memory (LTM). Sensory Memory retains impressions of sensory<br\/>information for a few seconds, while Short-Term Memory holds currently<br\/>relevant information for 20-30 seconds. Long-Term Memory can store<br\/>information for days to decades and includes explicit (declarative) and<br\/>implicit (procedural) memory.<\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"849\" height=\"476\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-26-1.webp\" alt=\"data extraction\" class=\"wp-image-229537\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-26-1.webp 849w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-26-1-300x168.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-26-1-768x431.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-26-1-150x84.webp 150w\" sizes=\"auto, (max-width: 849px) 100vw, 849px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>If you are extracting the data from other sources, then data extraction tools work very well. RAG applications require robust tools for extracting structured and unstructured data from various sources:<\/p>\n<ul class=\"wp-block-list\">\n<li>Websites, PDFs, Word documents, slides, etc.<\/li>\n<li>Tools like <a href=\"https:\/\/pypi.org\/project\/beautifulsoup4\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">BeautifulSoup<\/a> or <a href=\"https:\/\/pypi.org\/project\/PyPDF2\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">PyPDF2<\/a> can automate this process.<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>pip install -U langchain-community\n%pip install langchain pypdf<\/code><\/pre>\n<pre class=\"wp-block-code\"><code># %pip install langchain pypdf\n\nfrom langchain.document_loaders import PyPDFLoader\n\n# Define the path to your PDF file\npdf_path = \"\/content\/Multimodal Agent Using Agno Framework.pdf\"\n\n# Initialize the PyPDFLoader\nloader = PyPDFLoader(pdf_path)\n\n# Load the PDF and split it into pages\ndocuments = loader.load()\n\n# Print the content of each page\nfor i, doc in enumerate(documents):\n    print(f\"Page {i + 1} Content:\")\n    print(doc.page_content)\n    print(\"\\n\")\n<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output-1\">Output<\/h4>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"949\" height=\"730\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-13-5.webp\" alt=\"Output\" class=\"wp-image-229538\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-13-5.webp 949w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-13-5-300x231.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-13-5-768x591.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-13-5-150x115.webp 150w\" sizes=\"auto, (max-width: 949px) 100vw, 949px\"\/><\/figure>\n<\/div>\n<h2 class=\"wp-block-heading\" id=\"h-5-embeddings\">5. Embeddings<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"715\" height=\"443\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-23-1-1.webp\" alt=\"Embeddings\" class=\"wp-image-229539\" style=\"width:805px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-23-1-1.webp 715w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-23-1-1-300x186.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-23-1-1-150x93.webp 150w\" sizes=\"auto, (max-width: 715px) 100vw, 715px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>Text embeddings transform textual data into numerical vectors for similarity-based retrieval. Beyond text embeddings:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Image embeddings<\/strong>: Used in multimodal RAG applications.<\/li>\n<li><strong>Multi-modal embeddings<\/strong>: Combine text, image, and other data types for complex tasks.<\/li>\n<\/ul>\n<p>Here are the embedding models across providers:<\/p>\n<h3 class=\"wp-block-heading\">OpenAI Embeddings<\/h3>\n<ul class=\"wp-block-list\">\n<li>Latest models: <a href=\"https:\/\/platform.openai.com\/docs\/guides\/embeddings\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">text-embedding-3-small <\/a>(lower cost) and text-embedding-3-large (higher accuracy)<\/li>\n<li>Features: Dynamic dimension adjustment (e.g., 256-3072 dim), multilingual support, optimized for search\/RAG<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\">Cohere Embed v3<\/h3>\n<ul class=\"wp-block-list\">\n<li>Specializes in document quality ranking and noisy data handling<\/li>\n<li>Models: English\/multilingual variants (1024\/384 dim), compression-aware training for cost efficiency<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\">Nomic Embed v2<\/h3>\n<ul class=\"wp-block-list\">\n<li>Open-source <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/01\/how-to-run-mixtral-8x7b-moe-on-colab-for-free\/\" target=\"_blank\" rel=\"noreferrer noopener\">MoE architecture<\/a> (305M active params) with Matryoshka embeddings<\/li>\n<li>Multilingual (100+ languages), outperforms models 2x its size on MTEB\/BEIR benchmarks<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\">Gemini Embedding<\/h3>\n<ul class=\"wp-block-list\">\n<li>Experimental model (gemini-embedding-exp-03-07) with 8K token input and 3K dimensions<\/li>\n<li>MTEB leaderboard leader (68.32 mean score), supports 100+ languages<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\">Ollama Embeddings<\/h3>\n<ul class=\"wp-block-list\">\n<li>Hosts models like mxbai-embed-large and custom variants (e.g., suntray-embedding)<\/li>\n<li>Designed for RAG workflows with local inference and ChromaDB integration<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\">BGE (BAAI)<\/h3>\n<ul class=\"wp-block-list\">\n<li>BERT-based models (large\/base\/small-en-v1.5) for retrieval\/RAG<\/li>\n<li>Open-source, supports instruction tuning (e.g., \u201cRepresent this sentence\u2026\u201d)<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-mixedbread\">Mixedbread<\/h3>\n<ul class=\"wp-block-list\">\n<li>The mxbai-embed-large-v1 model by <a href=\"https:\/\/www.mixedbread.com\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Mixedbread AI <\/a>is a state-of-the-art sentence embedding solution designed for multilingual and multimodal retrieval tasks.<\/li>\n<li>It supports advanced techniques like Matryoshka Representation Learning (MRL) and binary quantization, enabling efficient memory usage and cost reduction at scale. With strong performance across diverse tasks, it rivals larger proprietary models while maintaining open-source accessibility<\/li>\n<\/ul>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"975\" height=\"669\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-12-4.webp\" alt=\"embedding RAG\" class=\"wp-image-229540\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-12-4.webp 975w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-12-4-300x206.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-12-4-768x527.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-12-4-150x103.webp 150w\" sizes=\"auto, (max-width: 975px) 100vw, 975px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-splitting-the-pdf-content-into-chunks\">Splitting the PDF content into chunks<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain.document_loaders import PyMuPDFLoader\nfrom langchain.text_splitter import RecursiveCharacterTextSplitter\ndef create_simple_chunks(file_path, chunk_size=3500, chunk_overlap=200):\n\u00a0\u00a0\u00a0\u00a0loader = PyMuPDFLoader(file_path)\n\u00a0\u00a0\u00a0\u00a0doc_pages = loader.load()\n\u00a0\u00a0\u00a0\u00a0splitter = RecursiveCharacterTextSplitter(chunk_size=chunk_size, chunk_overlap=chunk_overlap)\n\u00a0\u00a0\u00a0\u00a0return splitter.split_documents(doc_pages)\nfrom glob import glob\npdf_files = glob('.\/rag_docs\/*.pdf')\n# Process PDF files\npaper_docs = []\nfor fp in pdf_files:\n\u00a0\u00a0\u00a0\u00a0paper_docs.extend(create_simple_chunks(file_path=fp))<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output-2\">Output<\/h4>\n<pre class=\"wp-block-preformatted\">Loading pages: .\/rag_docs\/cnn_paper.pdf<p>Chunking pages: .\/rag_docs\/cnn_paper.pdf<\/p><p>Finished processing: .\/rag_docs\/cnn_paper.pdf<\/p><p>Loading pages: .\/rag_docs\/attention_paper.pdf<\/p><p>Chunking pages: .\/rag_docs\/attention_paper.pdf<\/p><p>Finished processing: .\/rag_docs\/attention_paper.pdf<\/p><p>Loading pages: .\/rag_docs\/vision_transformer.pdf<\/p><p>Chunking pages: .\/rag_docs\/vision_transformer.pdf<\/p><p>Finished processing: .\/rag_docs\/vision_transformer.pdf<\/p><p>Loading pages: .\/rag_docs\/resnet_paper.pdf<\/p><p>Chunking pages: .\/rag_docs\/resnet_paper.pdf<\/p><p>Finished processing: .\/rag_docs\/resnet_paper.pdf<\/p><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-creating-the-embeddings\">Creating the Embeddings<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_openai import OpenAIEmbeddings\nfrom langchain_chroma import Chroma\n# Initialize embedding model\nopenai_embed_model = OpenAIEmbeddings(model=\"text-embedding-3-small\")\n# Combine documents\ntotal_docs = wiki_docs_processed + paper_docs\n# Create and save vector database\nchroma_db = Chroma.from_documents(documents=total_docs,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0collection_name=\"my_db\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0embedding=openai_embed_model,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0collection_metadata={\"hnsw:space\": \"cosine\"},\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0persist_directory=\".\/my_db\")<\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-6-vector-databases\">6. Vector Databases<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"774\" height=\"478\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-6-4.webp\" alt=\"Vector databases\" class=\"wp-image-229541\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-6-4.webp 774w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-6-4-300x185.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-6-4-768x474.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-6-4-150x93.webp 150w\" sizes=\"auto, (max-width: 774px) 100vw, 774px\"\/><\/figure>\n<\/div>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/12\/top-vector-databases\/\" target=\"_blank\" rel=\"noreferrer noopener\">Vector databases<\/a> store embeddings (numerical representations of text or other data), enabling efficient retrieval of semantically similar chunks. Examples include:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong><a href=\"https:\/\/www.pinecone.io\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Pinecone<\/a>:<\/strong> A managed vector database platform designed for high-performance and scalable applications, enabling efficient storage and retrieval of high-dimensional vector embeddings.<\/li>\n<li><strong><a href=\"https:\/\/www.trychroma.com\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Chroma DB<\/a>:<\/strong> An open-source AI-native embedding database that includes features like vector search, document storage, full-text search, and metadata filtering, facilitating seamless retrieval in AI applications.<\/li>\n<li><strong><a href=\"https:\/\/qdrant.tech\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Qdrant<\/a>:<\/strong> An open-source vector database and search engine written in Rust, offering fast and scalable vector similarity search services with extended filtering support, suitable for neural-network or semantic-based matching.<\/li>\n<li><strong><a href=\"https:\/\/milvus.io\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Milvus DB<\/a>:<\/strong> An open-source vector database built for scalable similarity search, capable of handling large-scale and dynamic vector data, and supporting various index types for efficient retrieval.<\/li>\n<li><strong><a href=\"https:\/\/weaviate.io\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Weaviate<\/a>:<\/strong> An open-source vector database that stores both objects and vectors, allowing for combining vector search with structured filtering, and is modular, cloud-native, and real-time.<\/li>\n<\/ul>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1010\" height=\"708\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-7-3.webp\" alt=\"vector database RAG\" class=\"wp-image-229543\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-7-3.webp 1010w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-7-3-300x210.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-7-3-768x538.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-7-3-150x105.webp 150w\" sizes=\"auto, (max-width: 1010px) 100vw, 1010px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-example-of-vector-database-for-rag-building\">Example of Vector Database for RAG Building<\/h3>\n<p>Note: Above we already did make the embeddings, and now we will store them in the vector database.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-using-chroma-db-to-store-the-embeddings\">Using Chroma db to store the embeddings<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_openai import OpenAIEmbeddings\nfrom langchain_chroma import Chroma\n# Initialize embedding model\nopenai_embed_model = OpenAIEmbeddings(model=\"text-embedding-3-small\")\n# Combine documents\ntotal_docs = wiki_docs_processed + paper_docs\n# Create and save vector database\nchroma_db = Chroma.from_documents(documents=total_docs,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0collection_name=\"my_db\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0embedding=openai_embed_model,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0collection_metadata={\"hnsw:space\": \"cosine\"},\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0persist_directory=\".\/my_db\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-loading-the-vector-database\">Loading the Vector database<\/h4>\n<pre class=\"wp-block-code\"><code>chroma_db = Chroma(persist_directory=\".\/my_db\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0collection_name=\"my_db\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0embedding_function=openai_embed_model)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-retrieving-the-information-and-getting-the-output\">Retrieving the information and getting the output<\/h4>\n<pre class=\"wp-block-code\"><code>similarity_retriever = chroma_db.as_retriever(search_type=\"similarity\", search_kwargs={\"k\": 5})\n# Query for semantic similarity\nquery = \"What is machine learning?\"\ntop_docs = similarity_retriever.invoke(query)\n# Display results\nfrom IPython.display import display, Markdown\ndef display_docs(docs):\n\u00a0\u00a0\u00a0\u00a0for doc in docs:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0print('Metadata:', doc.metadata)\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0print('Content Brief:')\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0display(Markdown(doc.page_content[:1000]))\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0print()\ndisplay_docs(top_docs)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output-3\">Output<\/h4>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"891\" height=\"384\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-8.png\" alt=\"\" class=\"wp-image-229405\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-8.png 891w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-8-300x129.png 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-8-768x331.png 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-8-150x65.png 150w\" sizes=\"auto, (max-width: 891px) 100vw, 891px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-7-rerankers\">7. Rerankers<\/h2>\n<p>Rerankers refine the retrieval process by improving the relevance of retrieved documents:<\/p>\n<p>They operate in a two-stage retrieval pipeline:<\/p>\n<ol class=\"wp-block-list\">\n<li>Initial recall retrieves a broad set of candidates from the vector database.<\/li>\n<li>Rerankers prioritize the most relevant documents based on additional scoring mechanisms like semantic similarity or contextual relevance.<br \/>This approach significantly enhances the precision of <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/02\/rag-systems-with-nomic-embeddings\/\" target=\"_blank\" rel=\"noreferrer noopener\">RAG systems<\/a>.<\/li>\n<\/ol>\n<p>By integrating rerankers into the stack, developers can ensure higher-quality responses tailored to user queries while optimizing retrieval efficiency.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"863\" height=\"476\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-22-5.webp\" alt=\"Rerankers\" class=\"wp-image-229545\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-22-5.webp 863w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-22-5-300x165.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-22-5-768x424.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-22-5-150x83.webp 150w\" sizes=\"auto, (max-width: 863px) 100vw, 863px\"\/><\/figure>\n<p>Also read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/reranker-for-rag\/\" target=\"_blank\" rel=\"noreferrer noopener\">Comprehensive Guide on Reranker for RAG<\/a><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-example-of-rerankers-for-rag-building\">Example of Rerankers for RAG Building<\/h3>\n<pre class=\"wp-block-code\"><code>%pip install --upgrade --quiet  cohere<\/code><\/pre>\n<p>Set up the Cohere and ContextualCompressionRetriever<\/p>\n<pre class=\"wp-block-code\"><code>from langchain.retrievers.contextual_compression import ContextualCompressionRetriever\nfrom langchain_cohere import CohereRerank\nfrom langchain_community.llms import Cohere\nfrom langchain.chains import RetrievalQA\n\nllm = Cohere(temperature=0)\ncompressor = CohereRerank(model=\"rerank-english-v3.0\")\ncompression_retriever = ContextualCompressionRetriever(\n   base_compressor=compressor, base_retriever=retriever\n)\nchain = RetrievalQA.from_chain_type(\n   llm=Cohere(temperature=0), retriever=compression_retriever\n)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output-4\">Output<\/h4>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"751\" height=\"215\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-21.png\" alt=\"\" class=\"wp-image-229446\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-21.png 751w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-21-300x86.png 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-21-150x43.png 150w\" sizes=\"auto, (max-width: 751px) 100vw, 751px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-8-evaluation\">8. Evaluation<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"718\" height=\"401\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-24-3.webp\" alt=\"Evaluation\" class=\"wp-image-229546\" style=\"width:843px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-24-3.webp 718w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-24-3-300x168.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-24-3-150x84.webp 150w\" sizes=\"auto, (max-width: 718px) 100vw, 718px\"\/><\/figure>\n<\/div>\n<p>Evaluation ensures the accuracy and relevance of RAG systems:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Giskard<\/strong>: A library for testing machine learning pipelines.<\/li>\n<li><strong>Ragas<\/strong>: Specifically designed to evaluate RAG pipelines by analyzing retrieval quality and generated outputs.<\/li>\n<li><strong>Arize Phoenix<\/strong>: An open-source observability library for evaluating, troubleshooting, and improving LLM outputs with features like model drift detection and cohort analysis.<\/li>\n<li><strong>Comet Opik<\/strong>: A fully open-source platform for evaluating, testing, and monitoring LLM applications with tools for observability, automated scoring, and unit testing across the development lifecycle<\/li>\n<li><strong>DeepEval:<\/strong> deepeval\u00a0offers three LLM evaluation metrics to evaluate retrievals:\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/docs.confident-ai.com\/docs\/metrics-contextual-precision\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">ContextualPrecisionMetric<\/a>: evaluates whether the\u00a0<strong>reranker<\/strong>\u00a0in your retriever ranks more relevant nodes in your retrieval context higher than irrelevant ones.<\/li>\n<li><a href=\"https:\/\/docs.confident-ai.com\/docs\/metrics-contextual-recall\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">ContextualRecallMetric<\/a>: evaluates whether the\u00a0<strong>embedding model<\/strong>\u00a0in your retriever is able to accurately capture and retrieve relevant information based on the context of the input.<\/li>\n<li><a href=\"https:\/\/docs.confident-ai.com\/docs\/metrics-contextual-relevancy\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">ContextualRelevancyMetric<\/a>: evaluates whether the\u00a0<strong>text chunk size<\/strong>\u00a0and\u00a0<strong>top-K<\/strong>\u00a0of your retriever is able to retrieve information without much irrelevancy.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-example-of-evaluation-for-rag-building\">Example of Evaluation for RAG Building<\/h3>\n<pre class=\"wp-block-code\"><code>from tqdm.notebook import tqdm\nfrom datasets import load_dataset\nfrom qdrant_client import QdrantClient\nfrom tqdm import tqdm\nfrom langchain.docstore.document import Document as LangchainDocument\nfrom langchain_text_splitters import RecursiveCharacterTextSplitter\nfrom openai import OpenAI\nimport deepeval\n\n# Get your key from https:\/\/platform.openai.com\/api-keys\nOPENAI_API_KEY = \"<openai_api_key>\"\n\n# Get your Confident AI API key from https:\/\/app.confident-ai.com\nCONFIDENT_AI_API_KEY = \"<confident_ai_api_key>\"\n\n# Get a FREE forever cluster at https:\/\/cloud.qdrant.io\/\n# More info: https:\/\/qdrant.tech\/documentation\/cloud\/create-cluster\/\nQDRANT_URL = \"<qdrant_url>\"\nQDRANT_API_KEY = \"<qdrant_api_key>\"\nCOLLECTION_NAME = \"qdrant-deepeval\"\n\nEVAL_SIZE = 10\nRETRIEVAL_SIZE = 3\n\ndataset = load_dataset(\"atitaarora\/qdrant_doc\", split=\"train\")\n\nlangchain_docs = [\n    LangchainDocument(\n        page_content=doc[\"text\"], metadata={\"source\": doc[\"source\"]}\n    )\n    for doc in tqdm(dataset)\n]\n\ntext_splitter = RecursiveCharacterTextSplitter(\n    chunk_size=512,\n    chunk_overlap=50,\n    add_start_index=True,\n    separators=[\"\\n\\n\", \"\\n\", \".\", \" \", \"\"],\n)\n\ndocs_processed = []\nfor doc in langchain_docs:\n    docs_processed += text_splitter.split_documents([doc])\n\nclient = QdrantClient(url=QDRANT_URL, api_key=QDRANT_API_KEY)\n\ndocs_contents, docs_metadatas = [], []\n\nfor doc in docs_processed:\n    if hasattr(doc, \"page_content\") and hasattr(doc, \"metadata\"):\n        docs_contents.append(doc.page_content)\n        docs_metadatas.append(doc.metadata)\n    else:\n        print(\n            \"Warning: Some documents do not have 'page_content' or 'metadata' attributes.\"\n        )\n\n# Uses FastEmbed - https:\/\/qdrant.tech\/documentation\/fastembed\/\n# To generate embeddings for the documents\n# The default model is `BAAI\/bge-small-en-v1.5`\nclient.add(\n    collection_name=COLLECTION_NAME,\n    metadata=docs_metadatas,\n    documents=docs_contents,\n)\n\nopenai_client = OpenAI(api_key=OPENAI_API_KEY)\n\n\ndef query_with_context(query, limit):\n\n    search_result = client.query(\n        collection_name=COLLECTION_NAME, query_text=query, limit=limit\n    )\n\n    contexts = [\n        \"document: \" + r.document + \",source: \" + r.metadata[\"source\"]\n        for r in search_result\n    ]\n    prompt_start = \"\"\" You're assisting a user who has a question based on the documentation.\n        Your goal is to provide a clear and concise response that addresses their query while referencing relevant information\n        from the documentation.\n        Remember to:\n        Understand the user's question thoroughly.\n        If the user's query is general (e.g., \"hi,\" \"good morning\"),\n        greet them normally and avoid using the context from the documentation.\n        If the user's query is specific and related to the documentation, locate and extract the pertinent information.\n        Craft a response that directly addresses the user's query and provides accurate information\n        referring the relevant source and page from the 'source' field of fetched context from the documentation to support your answer.\n        Use a friendly and professional tone in your response.\n        If you cannot find the answer in the provided context, do not pretend to know it.\n        Instead, respond with \"I don't know\".\n\n        Context:\\n\"\"\"\n\n    prompt_end = f\"\\n\\nQuestion: {query}\\nAnswer:\"\n\n    prompt = prompt_start + \"\\n\\n---\\n\\n\".join(contexts) + prompt_end\n\n    res = openai_client.completions.create(\n        model=\"gpt-3.5-turbo-instruct\",\n        prompt=prompt,\n        temperature=0,\n        max_tokens=636,\n        top_p=1,\n        frequency_penalty=0,\n        presence_penalty=0,\n        stop=None,\n    )\n\n    return (contexts, res.choices[0].text)\n\n\nqdrant_qna_dataset = load_dataset(\"atitaarora\/qdrant_doc_qna\", split=\"train\")\n\n\ndef create_deepeval_dataset(dataset, eval_size, retrieval_window_size):\n    test_cases = []\n    for i in range(eval_size):\n        entry = dataset[i]\n        question = entry[\"question\"]\n        answer = entry[\"answer\"]\n        context, rag_response = query_with_context(\n            question, retrieval_window_size\n        )\n        test_case = deepeval.test_case.LLMTestCase(\n            input=question,\n            actual_output=rag_response,\n            expected_output=answer,\n            retrieval_context=context,\n        )\n        test_cases.append(test_case)\n    return test_cases\n\n\ntest_cases = create_deepeval_dataset(\n    qdrant_qna_dataset, EVAL_SIZE, RETRIEVAL_SIZE\n)\n\ndeepeval.login_with_confident_api_key(CONFIDENT_AI_API_KEY)\n\ndeepeval.evaluate(\n    test_cases=test_cases,\n    metrics=[\n        deepeval.metrics.AnswerRelevancyMetric(),\n        deepeval.metrics.FaithfulnessMetric(),\n        deepeval.metrics.ContextualPrecisionMetric(),\n        deepeval.metrics.ContextualRecallMetric(),\n        deepeval.metrics.ContextualRelevancyMetric(),\n    ],\n)<\/qdrant_api_key><\/qdrant_url><\/confident_ai_api_key><\/openai_api_key><\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-9-open-llms-access\">9. Open LLMs Access<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"847\" height=\"477\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-16-4.webp\" alt=\"Open LLM Access\" class=\"wp-image-229547\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-16-4.webp 847w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-16-4-300x169.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-16-4-768x433.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-16-4-150x84.webp 150w\" sizes=\"auto, (max-width: 847px) 100vw, 847px\"\/><\/figure>\n<\/div>\n<p>Platforms enabling local or API-based access to open LLMs include:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Ollama<\/strong>: Allows running open LLMs locally.<\/li>\n<li><strong>Groq, Hugging Face, Together AI<\/strong>: Provide API integrations for open LLMs.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-example-of-open-llms-access-for-rag-building\">Example of Open LLMs Access for RAG Building<\/h3>\n<p><strong>Download Ollama:<\/strong>\u00a0<a href=\"https:\/\/ollama.com\/download\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Click here to download<\/a><\/p>\n<pre class=\"wp-block-code\"><code>curl -fsSL https:\/\/ollama.com\/install.sh | sh<\/code><\/pre>\n<p>After this, pull the DeepSeek R1:1.5b using:<\/p>\n<pre class=\"wp-block-code\"><code>ollama pull deepseek-r1:1.5b<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-install-the-required-libraries\">Install the required libraries<\/h4>\n<pre class=\"wp-block-code\"><code>!pip install langchain==0.3.11\n!pip install langchain-openai==0.2.12\n!pip install langchain-community==0.3.11\n!pip install langchain-chroma==0.1.4<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-open-ai-embedding-models\">Open AI Embedding Models<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_openai import OpenAIEmbeddings\nopenai_embed_model = OpenAIEmbeddings(model=\"text-embedding-3-small\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-create-a-vector-db-and-persist-on-the-disk\">Create a Vector DB and persist on the disk<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_community.document_loaders import PyPDFLoader\nloader = PyPDFLoader('AgenticAI.pdf')\npages = loader.load_and_split()\ntexts = [doc.page_content for doc in pages]\n\nfrom langchain_chroma import Chroma\nchroma_db = Chroma.from_texts(\n\u00a0\u00a0\u00a0\u00a0texts=texts,\n\u00a0\u00a0\u00a0\u00a0collection_name=\"db_docs\",\n\u00a0\u00a0\u00a0\u00a0collection_metadata={\"hnsw:space\": \"cosine\"},\u00a0 # Set distance function to cosine\nembedding=openai_embed_model\n)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-build-a-rag-chain\">Build a RAG Chain<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_core.prompts import ChatPromptTemplate\nprompt = \"\"\"You are an assistant for question-answering tasks.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Use the following pieces of retrieved context to answer the question.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0If no context is present or if you don't know the answer, just say that you don't know.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Do not make up the answer unless it is there in the provided context.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Keep the answer concise and to the point with regard to the question.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Question:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{question}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Context:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{context}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Answer:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"\"\"\nprompt_template = ChatPromptTemplate.from_template(prompt)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-load-connection-to-llm\">Load Connection to LLM<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_community.llms import Ollama\ndeepseek = Ollama(model=\"deepseek-r1:1.5b\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-langchain-syntax-for-rag-chain\">LangChain Syntax for RAG Chain<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain.chains import Retrieval\nrag_chain = Retrieval.from_chain_type(llm=deepseek,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0chain_type=\"stuff\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0retriever=similarity_threshold_retriever,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0chain_type_kwargs={\"prompt\": prompt_template})\nquery = \"Tell the Leaders\u2019 Perspectives on Agentic AI\"\nrag_chain.invoke(query)\n{'query': 'Tell the Leaders\u2019 Perspectives on Agentic AI',<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output-5\">Output<\/h4>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1104\" height=\"556\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-15-3.webp\" alt=\"output\" class=\"wp-image-229548\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-15-3.webp 1104w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-15-3-300x151.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-15-3-768x387.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/image-15-3-150x76.webp 150w\" sizes=\"auto, (max-width: 1104px) 100vw, 1104px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>Building effective RAG applications isn\u2019t just about plugging in a language model\u2014it\u2019s about choosing the right RAG Developer stack across the board, from frameworks and embeddings to vector databases and retrieval tools. When these components are thoughtfully integrated, they enable intelligent, scalable systems that can chat with PDFs, pull relevant facts in real time, and generate context-aware responses. As the ecosystem continues to evolve, staying agile with your tools and grounded in solid architecture will be key to building reliable, future-proof AI solutions.<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/pankaj9786\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_Lb7Lh0T.webp\" width=\"48\" height=\"48\" alt=\"Pankaj Singh\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>                Hi, I am Pankaj Singh Negi &#8211; Senior Content Editor | Passionate about storytelling and crafting compelling narratives that transform ideas into impactful content. I love reading about technology revolutionizing our lifestyle.                 <\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Building a RAG (Retrieval-Augmented Generation) application isn\u2019t just about plugging in a few tools\u2014it\u2019s about choosing the right stack that makes retrieval and generation not just possible but efficient and scalable. Let\u2019s say you\u2019re working on something like \u201cSmart Chat with PDF\u201d\u2014an AI app that lets users interact with PDFs conversationally. It\u2019s not as simple [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":170682,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[10921,19198,2059,32726,4172],"dealstore":[],"offerexpiration":[],"class_list":["post-170681","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-comprehensive","tag-developer","tag-guide","tag-rag","tag-stack"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>A Comprehensive Guide to RAG Developer Stack - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=170681\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"A Comprehensive Guide to RAG Developer Stack - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Building a RAG (Retrieval-Augmented Generation) application isn\u2019t just about plugging in a few tools\u2014it\u2019s about choosing the right stack that makes retrieval and generation not just possible but efficient and scalable. Let\u2019s say you\u2019re working on something like \u201cSmart Chat with PDF\u201d\u2014an AI app that lets users interact with PDFs conversationally. It\u2019s not as simple [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=170681\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-04-04T09:10:23+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png\" \/>\n\t<meta property=\"og:image:width\" content=\"2444\" \/>\n\t<meta property=\"og:image:height\" content=\"2106\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"19 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=170681#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=170681\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"A Comprehensive Guide to RAG Developer Stack\",\"datePublished\":\"2025-04-04T09:10:23+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=170681\"},\"wordCount\":2126,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=170681#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png\",\"keywords\":[\"Comprehensive\",\"Developer\",\"Guide\",\"RAG\",\"Stack\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=170681#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=170681\",\"url\":\"https:\/\/fivemor.com\/?p=170681\",\"name\":\"A Comprehensive Guide to RAG Developer Stack - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=170681#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=170681#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png\",\"datePublished\":\"2025-04-04T09:10:23+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=170681#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=170681\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=170681#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png\",\"width\":2444,\"height\":2106},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=170681#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"A Comprehensive Guide to RAG Developer Stack\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"A Comprehensive Guide to RAG Developer Stack - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=170681","og_locale":"en_US","og_type":"article","og_title":"A Comprehensive Guide to RAG Developer Stack - Som2ny Network","og_description":"Building a RAG (Retrieval-Augmented Generation) application isn\u2019t just about plugging in a few tools\u2014it\u2019s about choosing the right stack that makes retrieval and generation not just possible but efficient and scalable. Let\u2019s say you\u2019re working on something like \u201cSmart Chat with PDF\u201d\u2014an AI app that lets users interact with PDFs conversationally. It\u2019s not as simple [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=170681","og_site_name":"Som2ny Network","article_published_time":"2025-04-04T09:10:23+00:00","og_image":[{"width":2444,"height":2106,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png","type":"image\/png"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"19 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=170681#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=170681"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"A Comprehensive Guide to RAG Developer Stack","datePublished":"2025-04-04T09:10:23+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=170681"},"wordCount":2126,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=170681#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png","keywords":["Comprehensive","Developer","Guide","RAG","Stack"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=170681#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=170681","url":"https:\/\/fivemor.com\/?p=170681","name":"A Comprehensive Guide to RAG Developer Stack - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=170681#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=170681#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png","datePublished":"2025-04-04T09:10:23+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=170681#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=170681"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=170681#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/image-17.png","width":2444,"height":2106},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=170681#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"A Comprehensive Guide to RAG Developer Stack"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/170681","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=170681"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/170681\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/170682"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=170681"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=170681"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=170681"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=170681"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=170681"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}