{"id":189269,"date":"2025-04-17T17:04:21","date_gmt":"2025-04-17T17:04:21","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/how-to-build-agentic-rag-using-gpt-4-1\/"},"modified":"2025-04-17T17:04:21","modified_gmt":"2025-04-17T17:04:21","slug":"how-to-build-agentic-rag-using-gpt-4-1","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=189269","title":{"rendered":"How to Build Agentic RAG Using GPT-4.1?"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/09\/retrieval-augmented-generation-rag-in-ai\/\">Retrieval-Augmented Generation (RAG)<\/a> systems enhance generative AI capabilities by integrating external document retrieval to produce contextually rich responses. With the release of GPT 4.1, characterized by exceptional instruction-following, coding excellence, long-context support (up to 1 million tokens), and notable affordability, building agentic RAG systems becomes more powerful, efficient, and accessible. In this article, we\u2019ll discover what makes GPT-4.1 so powerful and learn how to build an agentic RAG system using GPT-4.1 mini.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-overview-of-gpt-4-1\">Overview of GPT 4.1<\/h2>\n<p>GPT 4.1 significantly improves upon its predecessors, providing substantial gains in:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Coding<\/strong>: Achieves a 55% success rate on SWE-bench Verified, significantly outperforming GPT 4o.<\/li>\n<li><strong>Instruction Following<\/strong>: Enhanced capabilities to handle complex, multi-step, and nuanced instructions effectively.<\/li>\n<li><strong>Long Context<\/strong>: Supports a context window of up to 1 million tokens, suitable for broad data analysis. However, retrieval accuracy slightly decreases with extended contexts.<\/li>\n<li><strong>Cost Efficiency<\/strong>: GPT-4.1 offers 83% lower costs and 50% reduced latency compared to GPT-4o.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-what-s-new-in-gpt-4-1\">What\u2019s New in GPT 4.1?<\/h2>\n<p>OpenAI has rolled out the GPT-4.1 lineup, including three models: GPT-4.1, GPT-4.1 Mini, and GPT-4.1 Nano. Here is what it offers:<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-1m-token-context-think-bigger-prompts\">1. 1M Token Context: Think Bigger Prompts<\/h3>\n<p>One of the headline features is the <strong>1-million-token context window<\/strong> \u2013 a first for OpenAI. You can now feed in <em>massive<\/em> blocks of code, research papers, or entire document sets in one go. That said, while it handles scale impressively, pinpoint accuracy fades as the input grows, so it\u2019s best used for broad context understanding rather than surgical precision.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-coding-upgrades-smarter-multilingual-more-accurate\">2. Coding Upgrades: Smarter, Multilingual, More Accurate<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"787\" height=\"454\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/1-2.webp\" alt=\"SWE BENCH\" class=\"wp-image-231440\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/1-2.webp 787w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/1-2-300x173.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/1-2-768x443.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/1-2-150x87.webp 150w\" sizes=\"(max-width: 787px) 100vw, 787px\"\/><\/figure>\n<\/div>\n<p>When it comes to programming, GPT-4.1 steps up significantly:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Python Benchmarks<\/strong>: It scores 55% on SWE-bench verified, outdoing GPT-4o.<\/li>\n<li><strong>Multilingual Code Tasks<\/strong>: Thanks to the Polyglot benchmark, it can handle multiple languages better than before.<\/li>\n<li>Ideal for auto-generating code, debugging, or even assisting in full-stack builds.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-3-better-instruction-following\">3. Better Instruction Following<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"786\" height=\"463\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/2-2.webp\" alt=\"Hard subset\" class=\"wp-image-231441\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/2-2.webp 786w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/2-2-300x177.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/2-2-768x452.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/2-2-150x88.webp 150w\" sizes=\"auto, (max-width: 786px) 100vw, 786px\"\/><\/figure>\n<\/div>\n<p>GPT-4.1 is now more responsive to <strong>multi-step instructions<\/strong> and nuanced formatting rules. Whether you\u2019re designing workflows or building AI agents, this model is much better at doing what you <em>actually<\/em> ask for.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-4-speed-amp-cost-half-the-latency-fraction-of-the-price\">4. Speed &amp; Cost: Half the Latency, Fraction of the Price<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"785\" height=\"569\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/3-2.webp\" alt=\"intelligence by latency\" class=\"wp-image-231442\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/3-2.webp 785w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/3-2-300x217.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/3-2-768x557.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/3-2-150x109.webp 150w\" sizes=\"auto, (max-width: 785px) 100vw, 785px\"\/><\/figure>\n<\/div>\n<p>This version is optimized for performance and affordability:<\/p>\n<ul class=\"wp-block-list\">\n<li>50% faster response times<\/li>\n<li>83% cheaper than GPT-4o<\/li>\n<li>The Nano variant is particularly geared for high-frequency, budget-sensitive use \u2013 perfect for scaling applications with tight margins.<\/li>\n<li>The GPT-4.1 mini model is designed to balance intelligence, speed, and cost. It offers high intelligence and fast speed, making it suitable for many use cases.\n<ul class=\"wp-block-list\">\n<li><strong>Pricing<\/strong>: $0.4 \u2013 $1.6 per input-output.<\/li>\n<li><strong>Input<\/strong>: Text and image.<\/li>\n<li><strong>Output<\/strong>: Text.<\/li>\n<li><strong>Context Window<\/strong>: 1,047,576 tokens (large capacity for processing).<\/li>\n<li><strong>Max Output<\/strong>: 32,768 tokens.<\/li>\n<li><strong>Knowledge Cutoff<\/strong>: June 1, 2024.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1394\" height=\"676\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Screenshot-from-2025-04-17-13-31-38.webp\" alt=\"cost of gpt 4.1 mini\" class=\"wp-image-231469\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Screenshot-from-2025-04-17-13-31-38.webp 1394w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Screenshot-from-2025-04-17-13-31-38-300x145.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Screenshot-from-2025-04-17-13-31-38-768x372.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Screenshot-from-2025-04-17-13-31-38-150x73.webp 150w\" sizes=\"auto, (max-width: 1394px) 100vw, 1394px\"\/><\/figure>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"883\" height=\"743\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/4-2.webp\" alt=\"Aider polygot benchmark\" class=\"wp-image-231443\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/4-2.webp 883w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/4-2-300x252.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/4-2-768x646.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/4-2-150x126.webp 150w\" sizes=\"auto, (max-width: 883px) 100vw, 883px\"\/><\/figure>\n<\/div>\n<p>Read this article to know more: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/04\/open-ai-gpt-4-1\/\" target=\"_blank\" rel=\"noreferrer noopener\">All About OpenAI\u2019s Latest GPT 4.1 Family<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-building-agentic-rag-using-gpt-4-1-mini\">Building Agentic RAG Using GPT 4.1 mini<\/h2>\n<p>I am building a multi-document, agentic RAG system with GPT 4.1 mini. Here\u2019s the workflow.<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Ingests<\/strong> two long PDFs (on ML and GenAI economics).<\/li>\n<li><strong>Chunks<\/strong> them into overlapping pieces (chunk_size=5000, chunk_overlap=300)\u2014designed to preserve context.<\/li>\n<li><strong>Embeds<\/strong> these chunks using OpenAI\u2019s text-embedding-3-small model.<\/li>\n<li><strong>Stores<\/strong> them in two separate Chroma vector stores for efficient similarity-based retrieval.<\/li>\n<li><strong>Wraps<\/strong> the retrieval+LLM prompt logic into two chains (one per topic).<\/li>\n<li><strong>Exposes<\/strong> these chains as tools to a LangChain Zero-Shot Agent, which routes the query to the right context.<\/li>\n<li>Queries like \u201cWhy is Self-Attention used?\u201d or \u201cHow will marketing change with GenAI?\u201d are answered accurately and contextually\u2014thanks to the big chunks and high-quality retrieval.<\/li>\n<\/ol>\n<h3 class=\"wp-block-heading\" id=\"h-1-setup-and-installation\">1. Setup and Installation<\/h3>\n<pre class=\"wp-block-code\"><code>!pip install langchain==0.3.23\n!pip install -U langchain-openai\n!pip install langchain-community==0.3.11\n!pip install langchain-chroma==0.1.4\n!pip install pypdf<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-install-the-necessary-imports\">Install the necessary imports<\/h4>\n<pre class=\"wp-block-code\"><code>from langchain_community.document_loaders import PyPDFLoader\nfrom langchain.text_splitter import RecursiveCharacterTextSplitter\nfrom langchain_chroma import Chroma\nfrom langchain_openai import ChatOpenAI, OpenAIEmbeddings\nfrom langchain_core.prompts import ChatPromptTemplate\nfrom langchain_core.runnables import RunnablePassthrough\nfrom langchain_core.output_parsers import StrOutputParser\nfrom langchain.agents import AgentType, Tool, initialize_agent<\/code><\/pre>\n<p>I am pinning specific versions of <a href=\"https:\/\/www.langchain.com\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LangChain<\/a> packages and related dependencies for compatibility\u2014smart move.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-openai-api-key\">2. OpenAI API Key<\/h3>\n<pre class=\"wp-block-code\"><code>from getpass import getpass\nOPENAI_KEY = getpass('Enter Open AI API Key: ')\nimport os\nos.environ['OPENAI_API_KEY'] = OPENAI_KEY<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-3-load-pdfs-using-pypdfloader\">3. Load PDFs using PyPDFLoader<\/h3>\n<pre class=\"wp-block-code\"><code>pdf_dir = \"\/content\/document_pdf\"\nmachinelearning_paper = os.path.join(pdf_dir, \"Machinelearningalgorithm.pdf\")\ngenai_paper = os.path.join(pdf_dir, \"the-economic-potential-of-generative-ai-\nthe-next-productivity-frontier.pdf\")\n\n# Load individual PDF documents\nprint(\"Loading ml pdf...\")\nml_loader = PyPDFLoader(machinelearning_paper)\nml_documents = ml_loader.load()\n\nprint(\"Loading genai pdf...\")\ngenai_loader = PyPDFLoader(genai_paper)\ngenai_documents = genai_loader.load()<\/code><\/pre>\n<p>Loads the PDFs into LangChain Document objects. Each page becomes one Document.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-3-chunk-with-recursivecharactertextsplitter\">3. Chunk with RecursiveCharacterTextSplitter<\/h3>\n<pre class=\"wp-block-code\"><code># Split the documents\ntext_splitter = RecursiveCharacterTextSplitter(chunk_size=5000, chunk_overlap=300)\n\nml_splits = text_splitter.split_documents(ml_documents)\ngenai_splits = text_splitter.split_documents(genai_documents)\n\nprint(f\"Created {len(ml_splits)} splits for ml PDF\")\nprint(f\"Created {len(genai_splits)} splits for genai PDF\")<\/code><\/pre>\n<p>This is the heart of your long-context handling. This tool:<\/p>\n<ul class=\"wp-block-list\">\n<li>Keeps chunks under 5000 tokens.<\/li>\n<li>Preserves context with 300 overlap.<\/li>\n<\/ul>\n<p>The recursive splitter tries splitting on paragraphs \u2192 sentences \u2192 characters, preserving as much semantic structure as possible.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-ml-splits-3\">Ml_splits[:3]<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1728\" height=\"493\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-1.webp\" alt=\"Output\" class=\"wp-image-231446\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-1.webp 1728w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-1-300x86.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-1-768x219.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-1-1536x438.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-1-150x43.webp 150w\" sizes=\"auto, (max-width: 1728px) 100vw, 1728px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-genai-splits-5\">genai_splits[:5]<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1736\" height=\"531\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/11-1.webp\" alt=\"Output\" class=\"wp-image-231448\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/11-1.webp 1736w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/11-1-300x92.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/11-1-768x235.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/11-1-1536x470.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/11-1-150x46.webp 150w\" sizes=\"auto, (max-width: 1736px) 100vw, 1736px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-4-embedding-with-openaiembeddings\">4. Embedding with OpenAIEmbeddings<\/h3>\n<pre class=\"wp-block-code\"><code># details here: https:\/\/openai.com\/blog\/new-embedding-models-and-api-updates\nopenai_embed_model = OpenAIEmbeddings(model=\"text-embedding-3-small\")<\/code><\/pre>\n<p>I am using the 2024 text-embedding-3-small model, which is:<\/p>\n<ul class=\"wp-block-list\">\n<li>Smaller, faster, yet more accurate than older models.<\/li>\n<li>Great for cost-effective, high-quality retrieval.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-5-store-in-chroma-vector-stores\">5. Store in Chroma Vector Stores<\/h3>\n<pre class=\"wp-block-code\"><code># Create separate vectorstores\nml_vectorstore = Chroma.from_documents(\n\u00a0\u00a0\u00a0documents=ml_splits,\n\u00a0\u00a0\u00a0embedding=openai_embed_model,\n\u00a0\u00a0\u00a0collection_metadata={\"hnsw:space\": \"cosine\"},\n\u00a0\u00a0\u00a0collection_name=\"ml-knowledge\"\n)\ngenai_vectorstore = Chroma.from_documents(\n\u00a0\u00a0\u00a0documents=genai_splits,\n\u00a0\u00a0\u00a0embedding=openai_embed_model,\n\u00a0\u00a0\u00a0collection_metadata={\"hnsw:space\": \"cosine\"},\n\u00a0\u00a0\u00a0collection_name=\"genai-knowledge\"\n)<\/code><\/pre>\n<p>Here, I am creating two vector stores:<\/p>\n<ul class=\"wp-block-list\">\n<li>One for ML-related chunks<\/li>\n<li>One for GenAI-related chunks<\/li>\n<\/ul>\n<p>Using cosine similarity for retrieval:<\/p>\n<pre class=\"wp-block-code\"><code>ml_retriever = ml_vectorstore.as_retriever(search_type=\"similarity_score_threshold\",search_kwargs={\"k\": 5,\"score_threshold\": 0.3})\n\ngenai_retriever = genai_vectorstore.as_retriever(search_type=\"similarity_score_threshold\",search_kwargs={\"k\": 5,\"score_threshold\": 0.3})<\/code><\/pre>\n<p>Only return the top 5 chunks with enough similarity. Keeps answers tight.<\/p>\n<pre class=\"wp-block-code\"><code>query = \"what are ML algorithms?\"\ntop3_docs = ml_retriever.invoke(query)\ntop3_docs<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1733\" height=\"803\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/12-1.webp\" alt=\"Output\" class=\"wp-image-231449\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/12-1.webp 1733w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/12-1-300x139.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/12-1-768x356.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/12-1-1536x712.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/12-1-150x70.webp 150w\" sizes=\"auto, (max-width: 1733px) 100vw, 1733px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-6-retrieval-and-prompt-creation\">6. Retrieval and Prompt Creation<\/h3>\n<pre class=\"wp-block-code\"><code># Create the prompt templates\nml_prompt = ChatPromptTemplate.from_template(\n\n\u00a0\u00a0\u00a0\"\"\"\n\u00a0\u00a0\u00a0You are an expert in machine learning algorithms with deep technical knowledge of the field.\n\u00a0\u00a0\u00a0Answer the following question based solely on the provided context extracted from relevant machine learning research documents.\n\u00a0\u00a0\u00a0Context:\n\u00a0\u00a0\u00a0{context}\n\u00a0\u00a0\u00a0Question:\n\u00a0\u00a0\u00a0{question}\n\n\u00a0\u00a0\u00a0If the answer cannot be found in the context, please respond with: \"I don't have enough information to answer this question based on the provided context.\"\n\u00a0\u00a0\u00a0\"\"\"\n)\n\ngenai_prompt = ChatPromptTemplate.from_template(\n\n\u00a0\u00a0\u00a0\"\"\"\n\u00a0\u00a0\u00a0You are an expert in the economic impact and potential of generative AI technologies across industries and markets.\n\u00a0\u00a0\u00a0Answer the following question based only on the provided context related to the economic aspects of generative AI.\n\u00a0\u00a0\u00a0Context:\n\u00a0\u00a0\u00a0{context}\n\u00a0\u00a0\u00a0Question:\n\u00a0\u00a0\u00a0{question}\n\n\u00a0\u00a0\u00a0\u00a0If the answer cannot be found in the context, please state \"I don't have enough information to answer this question based on the provided context.\"\n\u00a0\u00a0\u00a0\"\"\"\n\n)<\/code><\/pre>\n<p>Here, I am creating context-specific prompts:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>ML QA system<\/strong>: asks about algorithms, training, etc.<\/li>\n<li><strong>GenAI QA system<\/strong>: focuses on economic impact, cross-industry uses.<\/li>\n<\/ul>\n<p>These prompts also guard against hallucination with:<\/p>\n<p><em>\u201cIf the answer cannot be found in the context\u2026 respond with: \u2018I don\u2019t have enough information\u2026\u2019\u201d<\/em><\/p>\n<p>Perfect for reliability.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-7-lcel-chains\">7. LCEL Chains<\/h3>\n<pre class=\"wp-block-code\"><code>from langchain_openai import ChatOpenAI\nllm = ChatOpenAI(model_name=\"gpt-4.1-mini-2025-04-14\", temperature=0)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>def format_docs(docs):\n\u00a0\u00a0\u00a0return \"\\n\\n\".join(doc.page_content for doc in docs)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code># Create the RAG chains using LCEL\nml_chain = (\n\u00a0\u00a0\u00a0{\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"context\": lambda question: format_docs(ml_retriever.get_relevant_documents(question)),\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"question\": RunnablePassthrough()\n\n\u00a0\u00a0\u00a0}\n\n\u00a0\u00a0\u00a0| ml_prompt\n\u00a0\u00a0\u00a0| llm\n\u00a0\u00a0\u00a0| StrOutputParser()\n\n)\n\ngenai_chain = (\n\n\u00a0\u00a0\u00a0{\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"context\": lambda question: format_docs(genai_retriever.get_relevant_documents(question)),\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"question\": RunnablePassthrough()\n\n\u00a0\u00a0\u00a0}\n\n\u00a0\u00a0\u00a0| genai_prompt\n\u00a0\u00a0\u00a0| llm\n\u00a0\u00a0\u00a0| StrOutputParser()\n\n)<\/code><\/pre>\n<p>This is where LangChain Expression Language (LCEL) shines.<\/p>\n<ol class=\"wp-block-list\">\n<li>Retrieve chunks<\/li>\n<li>Format them as context<\/li>\n<li>Inject into the prompt<\/li>\n<li>Send to gpt-4.1-mini<\/li>\n<li>Parse the response string<\/li>\n<\/ol>\n<p>It\u2019s elegant, reusable, and modular.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-8-define-tools-for-agent\">8. Define Tools for Agent<\/h3>\n<pre class=\"wp-block-code\"><code># Define the tools\n\ntools = [\n\u00a0\u00a0\u00a0Tool(\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0name=\"ML Knowledge QA System\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0func=ml_chain.invoke,\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0description=\"Useful for when you need to answer questions related to machine learning concepts, models, training techniques, evaluation metrics, algorithms and practical implementations. Covers supervised and unsupervised learning, model optimization, bias-variance tradeoff, feature engineering, and algorithm selection. Input should be a fully formed question.\"\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0),\n\n\u00a0\u00a0\u00a0Tool(\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0name=\"GenAI QA System\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0func=genai_chain.invoke,\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0description=\"Useful for when you need to answer questions about the economic impact, market potential, and cross-industry implications of generative AI technologies. Input should be a fully formed question. Responses are based strictly on the provided context related to the economics of generative AI.\"\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0)\n\n]<\/code><\/pre>\n<p>Each chain becomes a Tool in LangChain. Tools are like plug-and-play capabilities for the agent.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-9-initialize-agent\">9. Initialize Agent<\/h3>\n<pre class=\"wp-block-code\"><code># Initialize the agent\nagent = initialize_agent(\n\u00a0\u00a0\u00a0tools,\n\u00a0\u00a0\u00a0llm,\n\u00a0\u00a0\u00a0agent=AgentType.ZERO_SHOT_REACT_DESCRIPTION,\n\u00a0\u00a0\u00a0verbose=True\n)<\/code><\/pre>\n<p>I am using the Zero-Shot ReAct agent, which interprets the query, decides which tool (ML or GenAI) to use, and routes the input accordingly.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-10-query-time\">10. Query Time!<\/h3>\n<pre class=\"wp-block-code\"><code>result = agent.invoke(\"How marketing and sale could be transformed using Generative AI?\")<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1738\" height=\"724\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13.webp\" alt=\"Output\" class=\"wp-image-231450\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13.webp 1738w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-300x125.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-768x320.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-1536x640.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-150x62.webp 150w\" sizes=\"auto, (max-width: 1738px) 100vw, 1738px\"\/><\/figure>\n<\/div>\n<p><strong>Agent:<\/strong><\/p>\n<ol class=\"wp-block-list\">\n<li>Chooses GenAI QA System<\/li>\n<li>Retrieves top context chunks from GenAI vectorstore<\/li>\n<li>Formats prompt<\/li>\n<li>Sends to GPT-4.1<\/li>\n<li>Returns a grounded, non-hallucinated answer<\/li>\n<\/ol>\n<pre class=\"wp-block-code\"><code>result1 = agent.invoke(\"why Self-Attention is used?\")<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1726\" height=\"388\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14.webp\" alt=\"Output\" class=\"wp-image-231451\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14.webp 1726w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-300x67.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-768x173.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-1536x345.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-150x34.webp 150w\" sizes=\"auto, (max-width: 1726px) 100vw, 1726px\"\/><\/figure>\n<\/div>\n<p><strong>Agent:<\/strong><\/p>\n<ol class=\"wp-block-list\">\n<li>Chooses ML QA System<\/li>\n<li>Retrieves top context chunks from ML vectorstore<\/li>\n<li>Formats prompt<\/li>\n<li>Sends to GPT-4.1mini<\/li>\n<li>Returns a grounded, non-hallucinated answer<\/li>\n<\/ol>\n<pre class=\"wp-block-code\"><code>result2 = agent.invoke(\"what are Tree-based algorithms?\")<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1741\" height=\"483\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15.webp\" alt=\"Output\" class=\"wp-image-231452\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15.webp 1741w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15-300x83.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15-768x213.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15-1536x426.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15-150x42.webp 150w\" sizes=\"auto, (max-width: 1741px) 100vw, 1741px\"\/><\/figure>\n<\/div>\n<p>GPT-4.1 proves to be exceptionally effective for working with large documents, thanks to its extended context window of up to 1 million tokens. This enhancement eliminates the long-standing limitations faced with previous models, where documents had to be heavily chunked into small segments, often losing semantic coherence.\u00a0<\/p>\n<p>With the ability to handle large chunks, such as the 5000-token segments used here, GPT-4.1 can ingest and reason over dense, information-rich sections without missing contextual links across paragraphs or pages. This is especially valuable in scenarios involving complex documents like academic papers or industry whitepapers, where understanding often depends on multi-page continuity. The model handles these extended chunks accurately and delivers context-grounded responses without hallucinations, a capability further amplified by well-designed retrieval prompts.\u00a0<\/p>\n<p>Moreover, in a RAG pipeline, the quality of responses is heavily tied to how much useful context the model can consume at once. GPT-4.1 removes the previous ceiling, making it possible to retrieve and reason over complete conceptual units rather than fragmented excerpts. As a result, you can ask deep, nuanced questions about long documents and receive precise, well-informed answers, making GPT-4.1 a game-changer for production-grade document analysis and retrieval-based applications.<\/p>\n<p>Also read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/07\/building-agentic-rag-systems-with-langgraph\/\" target=\"_blank\" rel=\"noreferrer noopener\">A Comprehensive Guide to Building Agentic RAG Systems with LangGraph<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-more-than-just-a-needle-in-a-haystack\">More Than Just a \u201cNeedle in a Haystack\u201d<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"812\" height=\"635\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/5-2.webp\" alt=\"needle in the Haystack\" class=\"wp-image-231444\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/5-2.webp 812w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/5-2-300x235.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/5-2-768x601.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/5-2-150x117.webp 150w\" sizes=\"auto, (max-width: 812px) 100vw, 812px\"\/><\/figure>\n<\/div>\n<p>This is a <a href=\"https:\/\/arize.com\/blog-course\/the-needle-in-a-haystack-test-evaluating-the-performance-of-llm-rag-systems\/\">needle-in-a-haystack benchmark<\/a> evaluating how well different models can retrieve or reason over a relevant piece of information (a \u201cneedle\u201d) buried within a long context (\u201chaystack\u201d).<\/p>\n<p>GPT-4.1 excels at finding specific facts in large documents, but OpenAI pushed things further with the <strong>OpenAI-MRCR benchmark<\/strong>, which tests multi-fact retrieval:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>With 2 key facts (\u201cneedles\u201d)<\/strong>: GPT-4.1 does better than 4.0.<\/li>\n<li><strong>With 4 or more<\/strong>: Larger models like GPT-4.5 still dominate, especially in shorter input scenarios.<\/li>\n<\/ul>\n<p>8-needle scenario \u2013 meaning 8 relevant pieces of information are embedded in a longer sequence of tokens, and the model is tested on its ability to retrieve or reference them accurately.<\/p>\n<p>So, while GPT-4.1 handles basic long-context tasks well, it\u2019s not quite ready for deep, interconnected reasoning yet.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-needle\">2 Needle<\/h3>\n<p>This typically refers to a simpler version of the task, possibly with fewer categories or simpler decision points. The \u201caccuracy\u201d in this case is measured by how well the model performs when distinguishing between two categories or making two distinct decisions.<\/p>\n<figure style=\"text-align: center;\">\n  <img alt=\"OpenAI MRCR accuracy, 2 needle\" loading=\"lazy\" width=\"810\" height=\"790\" decoding=\"async\" data-nimg=\"1\" class=\"mx-auto\" style=\"color:transparent\" sizes=\"auto, (min-width: 1728px) 1728px, 100vw\" srcset=\"https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=640&amp;q=90 640w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=750&amp;q=90 750w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=828&amp;q=90 828w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=1080&amp;q=90 1080w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=1200&amp;q=90 1200w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=1920&amp;q=90 1920w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=2048&amp;q=90 2048w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=3840&amp;q=90 3840w\" src=\"https:\/\/images.ctfassets.net\/kftzwdyauwt9\/2oTJ2p3iGsEPnBrYeNhxbb\/9d14d937dc6004da8a49561af01b6781\/OpenAI-MRCR_accuracy_2needle_Lightmode.svg?w=3840&amp;q=90\"\/><figcaption style=\"text-align: center;\">Source: OpenAI<\/figcaption><\/figure>\n<h3 class=\"wp-block-heading\" id=\"h-4-needle\">4 Needle<\/h3>\n<p>This would involve a more complex task where there are four distinct categories or outcomes to predict. It\u2019s a more challenging task for the model compared to \u201c2 needle,\u201d meaning the model has to make more nuanced distinctions.<\/p>\n<figure style=\"text-align: center;\">\n  <img alt=\"OpenAI MRCR accuracy, 4 needle\" loading=\"lazy\" width=\"810\" height=\"790\" decoding=\"async\" data-nimg=\"1\" class=\"mx-auto\" style=\"color:transparent\" sizes=\"auto, (min-width: 1728px) 1728px, 100vw\" srcset=\"https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=640&amp;q=90 640w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=750&amp;q=90 750w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=828&amp;q=90 828w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=1080&amp;q=90 1080w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=1200&amp;q=90 1200w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=1920&amp;q=90 1920w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=2048&amp;q=90 2048w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=3840&amp;q=90 3840w\" src=\"https:\/\/images.ctfassets.net\/kftzwdyauwt9\/5HRJ1DFBDAvOGhAcQ1x61m\/1b4aaca14d8c8caaaf7f3c70f285d089\/OpenAI-MRCR_accuracy_4needle_Lightmode.svg?w=3840&amp;q=90\"\/><figcaption style=\"text-align: center;\">Source: OpenAI<\/figcaption><\/figure>\n<h3 class=\"wp-block-heading\" id=\"h-8-needle\">8 Needle<\/h3>\n<p>An even more complex scenario, where the model has to correctly predict from eight different categories or outcomes. The higher the \u201cneedle\u201d count, the more challenging the task is, requiring the model to demonstrate a broader range of understanding and accuracy.<\/p>\n<figure style=\"text-align: center;\">\n  <img alt=\"OpenAI MRCR accuracy, 8 needle\" loading=\"lazy\" width=\"810\" height=\"790\" decoding=\"async\" data-nimg=\"1\" class=\"mx-auto\" style=\"color:transparent\" sizes=\"auto, (min-width: 1728px) 1728px, 100vw\" srcset=\"https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=640&amp;q=90 640w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=750&amp;q=90 750w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=828&amp;q=90 828w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=1080&amp;q=90 1080w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=1200&amp;q=90 1200w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=1920&amp;q=90 1920w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=2048&amp;q=90 2048w, https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=3840&amp;q=90 3840w\" src=\"https:\/\/images.ctfassets.net\/kftzwdyauwt9\/nOUIReIO4isSJZA5c98FI\/12e3f5f67eba0e996808dcb5181b3ea8\/OpenAI-MRCR_accuracy_8needle_Lightmode.svg?w=3840&amp;q=90\"\/><figcaption style=\"text-align: center;\">Source: OpenAI<\/figcaption><\/figure>\n<p>Still, depending on your use case (especially if you\u2019re working with under 200K tokens), alternatives like DeepSeek-R1 or Gemini 2.5 might give you more value per dollar.<\/p>\n<p>However, if your needs include cutting-edge reasoning or the most up-to-date knowledge, watch GPT-4.5 or competitors like Gemini.<\/p>\n<p>GPT-4.1 may not be a total game-changer, but it\u2019s a smart evolution, especially for developers. OpenAI focused on practical improvements: better coding support, long context processing, and lower costs to make the models more accessible.<\/p>\n<p>Still, areas like benchmark transparency and knowledge freshness leave space for rivals to leap in. As competition ramps up, GPT-4.1 proves OpenAI is listening\u2014now it\u2019s Google, Anthropic, and the rest\u2019s move.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-why-chunking-works-so-well-5000-300-overlap\">Why Chunking Works So Well (5000 + 300 overlap)?<\/h3>\n<p>The config:<\/p>\n<ul class=\"wp-block-list\">\n<li>chunk_size = 5000<\/li>\n<li>chunk_overlap = 300<\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-why-is-this-effective-with-gpt-4-1\">Why is this effective with GPT-4.1?<\/h4>\n<ul class=\"wp-block-list\">\n<li>GPT-4.1 supports 1M token context. Feeding longer chunks is finally useful now. Smaller chunks would\u2019ve missed the semantic glue between ideas spread across paragraphs.<\/li>\n<li>5000-token chunks ensure minimal semantic splitting, capturing large conceptual units like \u201cTransformer architecture\u201d or \u201ceconomic implications of GenAI\u201d.<\/li>\n<li>300-token overlap helps preserve cross-chunk context, preventing cutoff issues.<\/li>\n<\/ul>\n<p>That\u2019s likely why you\u2019re not seeing misses or hallucinations\u2014you\u2019re giving the LLM exactly the chunked context it needs.<\/p>\n<p>Alright, let\u2019s break this down with a step-by-step guide to building an agentic Retrieval-Augmented Generation (RAG) pipeline using GPT-4.1 and leveraging its 1 million token context window capability by chunking and indexing two large PDFs (50+ pages each) to retrieve accurate answers with zero hallucination.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-key-benefits-and-considerations-of-gpt-4-1\">Key Benefits and Considerations of GPT 4.1<\/h2>\n<ul class=\"wp-block-list\">\n<li><strong>Enhanced Retrieval<\/strong>: Superior performance in single-fact retrieval but slightly lower effectiveness in complex, multi-information synthesis tasks compared to larger models like GPT-4.5.<\/li>\n<li><strong>Cost-effectiveness<\/strong>: Particularly the Nano variant, ideal for budget-sensitive, high-throughput tasks.<\/li>\n<li><strong>Developer-friendly<\/strong>: Ideal for coding applications, legal document analysis, and lengthy context tasks.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>GPT-4.1 Mini emerges as a robust and cost-effective foundation for constructing agentic Retrieval-Augmented Generation (RAG) systems. Its support for a 1 million token context window allows for the ingestion of large, semantically rich document chunks, enhancing the model\u2019s ability to provide contextually grounded and accurate responses.\u200b<\/p>\n<p>GPT-4.1 Mini\u2019s enhanced instruction-following capabilities, long-context handling, and affordability make it an excellent choice for developing sophisticated, production-grade RAG applications. Its design facilitates deep, nuanced interactions with extensive documents, positioning it as a valuable asset in the evolving landscape of AI-driven information retrieval.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1744874102130\"><strong class=\"schema-faq-question\"><strong>Why use 5,000-token chunks instead of smaller ones for documents?<\/strong><\/strong> <\/p>\n<p class=\"schema-faq-answer\">Larger chunks let GPT-4.1 \u201csee\u201d bigger ideas all at once\u2014like explaining a whole recipe instead of just listing ingredients. Smaller chunks might split up connected ideas (like separating \u201cwhy self-attention works\u201d from \u201chow it\u2019s calculated\u201d), making answers less accurate.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1744874149274\"><strong class=\"schema-faq-question\"><strong>Why bother splitting PDFs into separate topics (ML vs. GenAI)?<\/strong><\/strong> <\/p>\n<p class=\"schema-faq-answer\">If you dump everything into one pile, the model might mix up answers about machine learning algorithms with economics reports. Separating them is like giving the AI two specialized brains: one for coding and one for business analysis.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1744874159336\"><strong class=\"schema-faq-question\"><strong>Is GPT-4.1 actually cheaper than older models?<\/strong><\/strong> <\/p>\n<p class=\"schema-faq-answer\">Yep! It\u2019s ~83% cheaper than GPT-4.0 for basic tasks, and the Nano variant is built for apps needing tons of queries on a budget (like chatbots for customer support). But if you\u2019re doing ultra-complex tasks, bigger models like GPT-4.5 might still be worth the cost.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1744874170999\"><strong class=\"schema-faq-question\"><strong>Can I use this setup for legal\/financial documents?<\/strong><\/strong> <\/p>\n<p class=\"schema-faq-answer\">Totally. The 1 M-token context means you can feed it entire contracts or reports without losing the bigger picture. Just tweak the prompts to say, \u201cYou\u2019re a legal expert analyzing clauses\u2026\u201d and it\u2019ll adapt.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1744874184097\"><strong class=\"schema-faq-question\"><strong>How does GPT-4.1 handle non-English content?<\/strong><\/strong> <\/p>\n<p class=\"schema-faq-answer\">It\u2019s way better at multilingual tasks than older versions! For coding, it understands mixed languages (like Python + SQL). For text, it supports common languages like Spanish or French\u2014but for niche dialects, competitors like Gemini 2.5 might still edge it out.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1744874194017\"><strong class=\"schema-faq-question\"><strong>What\u2019s the biggest weakness of this RAG setup?<\/strong><\/strong> <\/p>\n<p class=\"schema-faq-answer\">While it\u2019s great at finding single facts in long docs, asking it to connect 8+ hidden details (like solving a mystery novel) can trip it up. For deep analysis, pair it with a human, or maybe, wait for GPT-4.5!<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/pankaj9786\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_Lb7Lh0T.webp\" width=\"48\" height=\"48\" alt=\"Pankaj Singh\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>                Hi, I am Pankaj Singh Negi &#8211; Senior Content Editor | Passionate about storytelling and crafting compelling narratives that transform ideas into impactful content. I love reading about technology revolutionizing our lifestyle.                 <\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Retrieval-Augmented Generation (RAG) systems enhance generative AI capabilities by integrating external document retrieval to produce contextually rich responses. With the release of GPT 4.1, characterized by exceptional instruction-following, coding excellence, long-context support (up to 1 million tokens), and notable affordability, building agentic RAG systems becomes more powerful, efficient, and accessible. In this article, we\u2019ll discover [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":189270,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[25423,5293,74018,32726],"dealstore":[],"offerexpiration":[],"class_list":["post-189269","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-agentic","tag-build","tag-gpt4-1","tag-rag"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>How to Build Agentic RAG Using GPT-4.1? - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=189269\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"How to Build Agentic RAG Using GPT-4.1? - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Retrieval-Augmented Generation (RAG) systems enhance generative AI capabilities by integrating external document retrieval to produce contextually rich responses. With the release of GPT 4.1, characterized by exceptional instruction-following, coding excellence, long-context support (up to 1 million tokens), and notable affordability, building agentic RAG systems becomes more powerful, efficient, and accessible. In this article, we\u2019ll discover [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=189269\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-04-17T17:04:21+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"787\" \/>\n\t<meta property=\"og:image:height\" content=\"454\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"14 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=189269#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=189269\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"How to Build Agentic RAG Using GPT-4.1?\",\"datePublished\":\"2025-04-17T17:04:21+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=189269\"},\"wordCount\":2067,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=189269#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp\",\"keywords\":[\"agentic\",\"Build\",\"GPT4.1\",\"RAG\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=189269#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=189269\",\"url\":\"https:\/\/fivemor.com\/?p=189269\",\"name\":\"How to Build Agentic RAG Using GPT-4.1? - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=189269#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=189269#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp\",\"datePublished\":\"2025-04-17T17:04:21+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=189269#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=189269\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=189269#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp\",\"width\":787,\"height\":454},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=189269#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"How to Build Agentic RAG Using GPT-4.1?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"How to Build Agentic RAG Using GPT-4.1? - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=189269","og_locale":"en_US","og_type":"article","og_title":"How to Build Agentic RAG Using GPT-4.1? - Som2ny Network","og_description":"Retrieval-Augmented Generation (RAG) systems enhance generative AI capabilities by integrating external document retrieval to produce contextually rich responses. With the release of GPT 4.1, characterized by exceptional instruction-following, coding excellence, long-context support (up to 1 million tokens), and notable affordability, building agentic RAG systems becomes more powerful, efficient, and accessible. In this article, we\u2019ll discover [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=189269","og_site_name":"Som2ny Network","article_published_time":"2025-04-17T17:04:21+00:00","og_image":[{"width":787,"height":454,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"14 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=189269#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=189269"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"How to Build Agentic RAG Using GPT-4.1?","datePublished":"2025-04-17T17:04:21+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=189269"},"wordCount":2067,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=189269#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp","keywords":["agentic","Build","GPT4.1","RAG"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=189269#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=189269","url":"https:\/\/fivemor.com\/?p=189269","name":"How to Build Agentic RAG Using GPT-4.1? - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=189269#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=189269#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp","datePublished":"2025-04-17T17:04:21+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=189269#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=189269"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=189269#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/1-2.webp.webp","width":787,"height":454},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=189269#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"How to Build Agentic RAG Using GPT-4.1?"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/189269","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=189269"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/189269\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/189270"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=189269"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=189269"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=189269"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=189269"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=189269"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}