{"id":67077,"date":"2025-02-04T02:46:29","date_gmt":"2025-02-04T02:46:29","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/8-types-of-chunking-for-rag-systems\/"},"modified":"2025-02-04T02:46:29","modified_gmt":"2025-02-04T02:46:29","slug":"8-types-of-chunking-for-rag-systems","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=67077","title":{"rendered":"8 Types of Chunking for RAG Systems"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>The easiest way to learn anything\u2014whether it\u2019s for academics or personal growth\u2014is by breaking it down into smaller, more manageable chunks. Similarly, when you\u2019re tackling a complex subject, it can feel overwhelming at first. However, by dividing it into bite-sized pieces, understanding becomes much easier. Even if it seems like a small concept already, it\u2019s always possible to split it into even more parts, no matter how simple they are. This chunking method makes it easier for a person to grasp or learn something and forms the foundation for how we process information in everyday life. Surprisingly, machines work similarly. Chunking is not just a method but a cognitive psychology concept that plays a vital role in data processing and AI systems that use RAG. Today, we will be talking about 8 types of Chunking in\u00a0 RAG with some Hands-on!!<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-chunking-for-rag-system\">What is Chunking for RAG System?<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"559\" height=\"174\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-17-67a0e1d30d996.webp\" alt=\"\" class=\"wp-image-219373\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-17-67a0e1d30d996.webp 559w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-17-67a0e1d30d996-300x93.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-17-67a0e1d30d996-150x47.webp 150w\" sizes=\"(max-width: 559px) 100vw, 559px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"957\" height=\"290\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-1-67a0e0c9ca199-1.webp\" alt=\"\" class=\"wp-image-219374\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-1-67a0e0c9ca199-1.webp 957w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-1-67a0e0c9ca199-1-300x91.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-1-67a0e0c9ca199-1-768x233.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-1-67a0e0c9ca199-1-150x45.webp 150w\" sizes=\"auto, (max-width: 957px) 100vw, 957px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>Chunking is the process of breaking down large pieces of text into smaller, more manageable parts. This technique is crucial when working with language models because it ensures that the provided data fits within the model\u2019s context window while maintaining the relevance and quality of the information.<\/p>\n<p>By context window, I meant that every language model operates according to the user\u2019s requirements for providing their own data. However, a limitation restricts the user from passing unlimited data to the model. This is because:<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-the-context-limit\">The Context Limit<\/h4>\n<p>There is always a limit on the number of words or tokens that you can provide to the language model. Here\u2019s the context window of OpenAI models:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"939\" height=\"603\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-2-67a0e0c8d03c4.webp\" alt=\"The Context Limit\" class=\"wp-image-219375\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-2-67a0e0c8d03c4.webp 939w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-2-67a0e0c8d03c4-300x193.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-2-67a0e0c8d03c4-768x493.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-2-67a0e0c8d03c4-150x96.webp 150w\" sizes=\"auto, (max-width: 939px) 100vw, 939px\"\/><figcaption class=\"wp-element-caption\">Source: OpenAI<\/figcaption><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-maximizing-signal-to-noise-ratio\">Maximizing Signal-to-Noise Ratio<\/h4>\n<p>Language models perform better when the <strong>signal-to-noise ratio<\/strong> is high. In other words, reducing irrelevant or distracting information in the model\u2019s context window can significantly enhance performance.\u00a0<\/p>\n<p>So. the primary goal of chunking is not just to split data arbitrarily, but to <strong>optimize<\/strong> the way information is presented to the model. Proper chunking enhances the <strong>retrievability<\/strong> of useful content and improves the overall performance of applications relying on AI models.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-why-is-chunking-important\">Why is Chunking Important?<\/h2>\n<p><em>Anton Troynikov, co-founder of Chroma, points out, that unnecessary data within the context window can <\/em><strong><em>measurably degrade<\/em><\/strong><em> the overall effectiveness of an application. By focusing only on relevant content, we can optimize the model\u2019s output and ensure more accurate, efficient responses.<\/em><\/p>\n<p>Makes sense right? Similarly, Chunking is important because:\u00a0<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Overcoming Context Window Limitations<\/strong><strong><br \/><\/strong>Every language model has a fixed context window, which restricts the amount of data that can be processed at once. By chunking, you ensure that essential information is retained within these limits, preventing important data from being omitted or truncated.<\/li>\n<li><strong>Improving Signal-to-Noise Ratio<\/strong><strong><br \/><\/strong>When text is too large and contains unnecessary information, the model\u2019s performance can degrade. Chunking helps in filtering out irrelevant content, ensuring that only the most relevant data is provided to the model, thereby increasing the <strong>signal-to-noise ratio<\/strong> and boosting accuracy.<\/li>\n<li><strong>Enhancing Retrieval Efficiency<\/strong><strong><br \/><\/strong>Properly chunked data makes it easier to locate and retrieve relevant pieces when needed. This is especially important for retrieval-augmented generation (RAG) systems, where accessing the right information quickly can significantly impact response quality.<\/li>\n<li><strong>Task-Specific Optimization<\/strong><strong><br \/><\/strong>Different tasks may require different chunking strategies. For instance, summarization tasks may benefit from larger chunks to maintain coherence, while question-answering tasks might require finer granularity to provide precise answers. The key is to chunk in a way that aligns with the specific needs of the application.<\/li>\n<\/ol>\n<p>In summary, chunking is a foundational step in preparing text data for language models. It helps in balancing data volume, relevance, and retrievability, making it a critical practice in building efficient AI-powered applications.<\/p>\n<p>Let\u2019s understand this with the RAG architecture:<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-rag-architecture-to-comprehend-chunking\">RAG Architecture to Comprehend Chunking<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><a href=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e.webp\"><img loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"659\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e.webp\" alt=\"RAG Architecture to Understand Chunking\" class=\"wp-image-219376\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e.webp 1600w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e-300x124.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e-768x316.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e-1536x633.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-3-67a0e0c8d1f5e-150x62.webp 150w\" sizes=\"auto, (max-width: 1600px) 100vw, 1600px\"\/><\/a><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>In <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/09\/retrieval-augmented-generation-rag-in-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">Retrieval-Augmented Generation (RAG)<\/a>, chunking involves breaking down raw data sources (such as PDFs, spreadsheets, or other documents) into smaller, manageable pieces called \u201c<strong>chunks of text<\/strong>.\u201d The system then processes these chunks, converts them into vector embeddings, and stores them in a vector database (e.g., Chroma) to enable efficient retrieval when a user asks a question.<\/p>\n<p>In short, Chunking refers to <strong>dividing large text data into smaller, manageable pieces<\/strong> to improve retrieval efficiency and relevance in downstream tasks like search and generation.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-chunking\">1. Chunking<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Raw Data Source:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Input data can come from various sources such as PDFs, databases, and reports.<\/li>\n<li>These raw sources often contain large blocks of information that are difficult to process in their entirety.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Data Processing (Chunking Stage):<\/strong>\n<ul class=\"wp-block-list\">\n<li>The large documents are <strong>split into smaller chunks<\/strong>, ensuring that each chunk represents a meaningful segment of information.<\/li>\n<li>These chunks may follow different strategies, such as:\n<ul class=\"wp-block-list\">\n<li><strong>Fixed-size chunks<\/strong> (e.g., 500 words each)<\/li>\n<li><strong>Semantic chunks<\/strong> (split based on meaning or structure, like paragraphs or sections)<\/li>\n<li><strong>Overlapping chunks<\/strong> (to preserve context between chunks)<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/li>\n<li><strong>Embedding Chunks:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Each chunk is passed through an <strong>embedding model<\/strong>, which converts it into a high-dimensional vector representation.<\/li>\n<li>This process encodes the chunk\u2019s meaning, allowing for efficient similarity searches.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-2-chunk-retrieval-using-vector-database\">2. Chunk Retrieval Using Vector Database<\/h3>\n<p>Once the chunks are embedded:<\/p>\n<ul class=\"wp-block-list\">\n<li>When a <strong>user asks a question<\/strong>, the query is also converted into an embedding vector.<\/li>\n<li>A <strong>vector search<\/strong> is performed to find the most relevant chunks from the database (Chroma in this case).<\/li>\n<li>The retrieved chunks (which are the most similar to the query) are sent to the LLM to provide contextual responses.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-3-generation-using-retrieved-chunks\">3. Generation Using Retrieved Chunks<\/h3>\n<p>After chunk retrieval:<\/p>\n<ul class=\"wp-block-list\">\n<li>The <strong>retrieved chunks<\/strong> are bundled with additional components like:\n<ul class=\"wp-block-list\">\n<li><strong>Instruction:<\/strong> Defines how the model should respond.<\/li>\n<li><strong>Context:<\/strong> The retrieved chunk(s) provide the factual basis.<\/li>\n<li><strong>Query:<\/strong> The original user input.<\/li>\n<\/ul>\n<\/li>\n<li>The generator (LLM) then processes this information and generates a coherent response.<\/li>\n<\/ul>\n<p>Also read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/11\/rag-vs-agentic-rag\/\" target=\"_blank\" rel=\"noreferrer noopener\">RAG vs Agentic RAG: A Comprehensive Guide<\/a><\/p>\n<p>Let\u2019s understand the drawbacks of RAG.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-key-drawbacks-of-rag-retrieval-augmented-generation\">Key Drawbacks of RAG (Retrieval-Augmented Generation)<\/h2>\n<ol class=\"wp-block-list\">\n<li><strong>Retrieval Challenges:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Precision and Recall Issues: The retrieval phase often struggles to identify relevant information, leading to:\n<ul class=\"wp-block-list\">\n<li>Selection of misaligned or irrelevant content chunks.<\/li>\n<li>Missing critical information that is essential for accurate responses.<\/li>\n<\/ul>\n<\/li>\n<li>Inadequate Context: A single retrieval based on the original query may fail to capture sufficient context for complex issues.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Generation Difficulties:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Hallucination: The model may generate content that is not supported by the retrieved context, reducing reliability.<\/li>\n<li>Irrelevance, Toxicity, or Bias: Outputs may suffer from:\n<ul class=\"wp-block-list\">\n<li>Irrelevant or off-topic responses.<\/li>\n<li>Toxic or biased language undermines the quality and trustworthiness of the generated content.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/li>\n<li><strong>Augmentation Hurdles:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Integration Challenges: Combining retrieved information with the task at hand can result in:\n<ul class=\"wp-block-list\">\n<li>Disjointed or incoherent outputs.<\/li>\n<li>Redundancy due to repetitive information from multiple sources.<\/li>\n<\/ul>\n<\/li>\n<li>Stylistic and Tonal Inconsistency: Ensuring a consistent tone and style across the generated content adds complexity.<\/li>\n<li>Over-Reliance on Retrieved Content: The model may simply echo retrieved information without synthesizing or adding insightful analysis, limiting the depth of responses.<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n<p>By implementing the right chunking strategies, the <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/11\/rag-pipeline-for-hindi-documents\/\" target=\"_blank\" rel=\"noreferrer noopener\">RAG pipeline<\/a> can achieve more accurate retrieval, richer contextual grounding, and higher-quality response generation, ultimately enhancing the overall system\u2019s reliability and user satisfaction.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-how-to-choose-the-right-chunking-strategy\">How to Choose the Right Chunking Strategy?<\/h2>\n<p>Choosing the right chunking strategy involves carefully considering the <strong>content type<\/strong>, the <strong>embedding model<\/strong>, and the expected <strong>user queries<\/strong>. Here\u2019s a detailed guide tailored to your example scenario:<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-1-understand-the-nature-of-the-content\">1. Understand the Nature of the Content<\/h4>\n<p>Content characteristics heavily influence chunking strategy. Example Scenario:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Scientific documents (e.g., Nature articles):<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Structured content:<\/strong> Sections like Abstract, Introduction, Methods, etc.<\/li>\n<li><strong>Dense information:<\/strong> Each section may contain multiple key points.<\/li>\n<li>Long paragraphs and citations.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Chunking Strategy for Such Content:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>By logical sections:<\/strong> Treat sections like \u201cAbstract,\u201d \u201cMethods,\u201d etc., as individual chunks.<\/li>\n<li><strong>Smaller sub-chunks:<\/strong> Break long sections (e.g., 500\u2013800 tokens) into subsections by paragraph or semantic boundaries.<\/li>\n<li><strong>Maintain context:<\/strong> Avoid cutting in the middle of a thought or example to preserve semantic meaning.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-2-align-with-the-embedding-model\">2. Align with the Embedding Model<\/h4>\n<p>Different embedding models have varying limitations and strengths. Key Considerations:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Token Limitations:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Many embedding models (like OpenAI\u2019s models) have token limits. Ensure chunks fit well within these limits.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Semantic Encoding:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Embedding models work best when input chunks contain coherent and self-contained ideas.<\/li>\n<\/ul>\n<\/li>\n<li>A good chunk typically includes a full sentence, paragraph, or logically connected set of points.<\/li>\n<\/ul>\n<p><strong>Steps to Optimize<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Calculate Token Sizes:<\/strong> Use tools or scripts to estimate the token count of your content to ensure compatibility with the embedding model.<\/li>\n<li><strong>Pre-process with Overlapping Context: <\/strong>When breaking content into chunks, ensure some overlap between chunks (e.g., 20\u201330% overlap) to prevent loss of semantic connections across boundaries.<\/li>\n<li><strong>Prioritize Structure:<\/strong> Embed well-structured and self-contained chunks for better semantic relevance.<\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-3-anticipate-user-queries\">3. Anticipate User Queries<\/h4>\n<p>Understanding what users are likely to search for helps design the chunking strategy. Example User Queries:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>General topics (e.g., \u201cWhat is the method used in this study?\u201d):<\/strong>\n<ul class=\"wp-block-list\">\n<li>Chunks aligned with document sections allow faster retrieval.<\/li>\n<li>Abstract or Results sections might be frequently accessed.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Specific details (e.g., \u201cWhat\u2019s the p-value for Experiment 1?\u201d):<\/strong>\n<ul class=\"wp-block-list\">\n<li>Finer-grained chunking ensures detail-level retrieval.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p>In the next section, I will discuss different chunking strategies in detail.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-1-character-text-chunking\">1. Character Text Chunking<\/h2>\n<p>This method is one of the simplest approaches to chunking or splitting text. It divides the text into fixed-sized chunks of <strong>N characters<\/strong>, regardless of the content or structure. While it is a basic technique, it serves as an excellent starting point for understanding the fundamentals of text chunking and how it works in practice.<\/p>\n<p>This approach is easy and simple to use; however, it is very rigid and does not take into account the structure of your text.<\/p>\n<pre class=\"wp-block-code\"><code>text = \"Clouds come floating into my life, no longer to carry rain or usher storm, but to add color to my sunset sky.\"\nchunks = []\nchunk_size = 35\nchunk_overlap = 5 # Characters\n# Run through the text with the length of your text and iterate every chunk_size,\n# considering the overlap for the starting position of the next chunk.\nfor i in range(0, len(text) - chunk_size + 1, chunk_size - chunk_overlap):\n\u00a0\u00a0\u00a0chunk = text[i:i + chunk_size]\n\u00a0\u00a0\u00a0chunks.append(chunk)\nchunks<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">['Clouds come floating into my life, ',<br\/>\u00a0'ife, no longer to carry rain or ush',<br\/>\u00a0'r usher storm, but to add color to ']<\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-explanation\">Explanation:<\/h3>\n<ol class=\"wp-block-list\">\n<li><strong>Input Text:<\/strong>\n<ul class=\"wp-block-list\">\n<li>A string variable text contains a sentence.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Chunks List Initialization:<\/strong>\n<ul class=\"wp-block-list\">\n<li>chunks = [] creates an empty list to store text segments.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Chunking Parameters:<\/strong>\n<ul class=\"wp-block-list\">\n<li>chunk_size = 35: Defines the length of each chunk to be 35 characters.<\/li>\n<li>chunk_overlap = 5: Specifies that each chunk will overlap with the previous one by 5 characters.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Chunking Process:<\/strong>\n<ul class=\"wp-block-list\">\n<li>The for loop iterates through the text using a step size of chunk_size \u2013 chunk_overlap, meaning new chunks will start every 30 characters but will include the last 5 characters from the previous chunk.<\/li>\n<li>The loop range is determined by len(text) \u2013 chunk_size + 1, ensuring it doesn\u2019t go beyond the text length.<\/li>\n<li>In each iteration, a substring of length chunk_size is extracted from the text and added to the chunks list.<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n<h3 class=\"wp-block-heading\" id=\"h-explanation-of-the-overlapping-mechanism\">Explanation of the Overlapping Mechanism<\/h3>\n<p><strong>Step Size Calculation:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>The loop iterates with a step of chunk_size \u2013 chunk_overlap, which means:<br \/>35\u22125=30.<\/li>\n<li>This means after processing the first 35 characters, the next chunk starts 30 characters after the first one, causing a 5-character overlap.<\/li>\n<\/ul>\n<p>Let\u2019s analyze how the loop runs with the given values:<\/p>\n<p><strong>First chunk (index 0 to 35):<\/strong><strong><br \/><\/strong>Extracts the substring \u201cClouds come floating into my life, \u201c.<br \/>The loop then moves forward by 30 characters.<\/p>\n<p><strong>Second chunk (index 30 to 65):<\/strong><strong><br \/><\/strong>Extracts the substring \u201cife, no longer to carry rain or ush\u201d.<br \/>Notice how the last 5 characters of the previous chunk (\u201clife,\u201d) overlap in this chunk.<\/p>\n<p><strong>Third chunk (index 60 to 95):<\/strong><strong><br \/><\/strong>Extracts the substring \u201cr usher storm, but to add color to \u201c.<br \/>Again, there\u2019s an overlap with the last few characters from the second chunk.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-now-let-s-do-it-with-langchain-nbsp\"><strong>Now let\u2019s do it with Langchain\u00a0<\/strong><\/h4>\n<pre class=\"wp-block-code\"><code>%pip install -qU langchain-text-splitters<\/code><\/pre>\n<p>This command installs the langchain-text-splitters library, which is used for splitting long pieces of text into smaller chunks.<\/p>\n<p>The -q flag suppresses installation output, and -U ensures that the latest version is installed.<\/p>\n<pre class=\"wp-block-code\"><code># Load an example document\nwith open(\"state_of_the_union.txt\") as f:\n\u00a0\u00a0state_of_the_union = f.read()<\/code><\/pre>\n<ul class=\"wp-block-list\">\n<li>Opens the file state_of_the_union.txt and reads its entire content into the variable state_of_the_union as a string.<\/li>\n<li>This document is presumably the transcript of a U.S. State of the Union address.<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>text_splitter = CharacterTextSplitter(\n\u00a0\u00a0separator=\"\\n\\n\",\n\u00a0\u00a0chunk_size=1000,\n\u00a0\u00a0chunk_overlap=200,\n\u00a0\u00a0length_function=len,\n\u00a0\u00a0is_separator_regex=False,\n)<\/code><\/pre>\n<p>This code sets up a CharacterTextSplitter object with the following parameters:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>separator=\u201d\\n\\n\u201d<\/strong>\n<ul class=\"wp-block-list\">\n<li>The document is split by double newline characters (\\n\\n), which typically indicate paragraph breaks in text files.<\/li>\n<\/ul>\n<\/li>\n<li><strong>chunk_size=1000<\/strong>\n<ul class=\"wp-block-list\">\n<li>Each text chunk will contain approximately 1000 characters.<\/li>\n<\/ul>\n<\/li>\n<li><strong>chunk_overlap=200<\/strong>\n<ul class=\"wp-block-list\">\n<li>There will be a 200-character overlap between consecutive chunks to ensure context continuity when processing the text.<\/li>\n<\/ul>\n<\/li>\n<li><strong>length_function=len<\/strong>\n<ul class=\"wp-block-list\">\n<li>Specifies that the length of each chunk is calculated using Python\u2019s built-in len() function, which measures the number of characters.<\/li>\n<\/ul>\n<\/li>\n<li><strong>is_separator_regex=False<\/strong>\n<ul class=\"wp-block-list\">\n<li>Indicates that the separator provided (\u201c\\n\\n\u201d) is a literal string and not a regular expression.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>texts = text_splitter.create_documents([state_of_the_union])\nprint(texts[0])<\/code><\/pre>\n<p>The create_documents() method takes the list of texts (in this case, a single document) and splits it based on the specified parameters (chunk size, overlap, separator).<\/p>\n<p>The result is a list of chunked document objects, where each chunk contains a portion of the original text.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1500\" height=\"313\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-18-67a0e2f95f769.webp\" alt=\"Output\" class=\"wp-image-219382\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-18-67a0e2f95f769.webp 1500w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-18-67a0e2f95f769-300x63.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-18-67a0e2f95f769-768x160.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-18-67a0e2f95f769-150x31.webp 150w\" sizes=\"auto, (max-width: 1500px) 100vw, 1500px\"\/><\/figure>\n<p><strong>Chunking in Action:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>The content is split into paragraphs based on the double newline (\\n\\n) separator.<\/li>\n<li>This ensures the logical separation of ideas while maintaining readability.<\/li>\n<\/ul>\n<p><strong>Overlap Handling:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>The chunk may contain up to 200 characters from the previous chunk to preserve context.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-2-recursive-character-text-splitting\">2. Recursive Character Text Splitting<\/h2>\n<p>Unlike the first method which doesn\u2019t look for the document structure, this method recursively divides text using a predefined list of separators and intelligently merges the resulting smaller chunks into larger ones. The final chunks are optimized to contain no more than N characters, ensuring efficient text processing and context preservation.<\/p>\n<p>It is parameterized by a list of characters. The default list is:<\/p>\n<ul class=\"wp-block-list\">\n<li>\u201c\\n\\n\u201d \u2013 Double new line, or most commonly paragraph breaks<\/li>\n<li>\u201c\\n\u201d \u2013 New lines<\/li>\n<li>\u201d \u201d \u2013 Spaces<\/li>\n<li>\u201c\u201d \u2013 Characters<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>%pip install -qU langchain-text-splitters\n\ntext = \"\"\"\n\nThe Marvel Universe is a vast and interconnected world filled with superheroes, villains, and epic storytelling that has captivated audiences for decades. Founded by visionaries such as Stan Lee, Jack Kirby, and Steve Ditko, Marvel Comics has introduced some of the most iconic characters in pop culture history. From its early beginnings in 1939 as Timely Publications to its transformation into Marvel Comics in the 1960s, the company has consistently pushed the boundaries of storytelling by creating relatable and dynamic characters. Heroes like Spider-Man, Iron Man, Captain America, and Thor have become household names, each with their own compelling backstories and struggles that resonate with fans across generations. Marvel\u2019s success extends beyond the pages of comic books. The launch of the Marvel Cinematic Universe (MCU) in 2008 with the release of Iron Man revolutionized the film industry, introducing interconnected storylines that culminated in epic crossover events such as The Avengers and Infinity War. The MCU\u2019s success is largely attributed to its ability to blend action, humor, and emotional depth while maintaining the essence of the beloved comic book characters. Audiences have followed the journeys of superheroes as they face powerful foes like Thanos and Loki, all while dealing with their own internal conflicts and responsibilities.\"\"\"\n\nfrom langchain_text_splitters import RecursiveCharacterTextSplitter<\/code><\/pre>\n<p>The RecursiveCharacterTextSplitter is imported from the langchain-text-splitters package.<\/p>\n<p>This class is used to split large text documents into smaller chunks efficiently while preserving context.<\/p>\n<pre class=\"wp-block-code\"><code>text_splitter = RecursiveCharacterTextSplitter(\n\u00a0\u00a0\u00a0# Set a really small chunk size, just to show.\n\u00a0\u00a0\u00a0chunk_size=400,\n\u00a0\u00a0\u00a0chunk_overlap=0,\n\u00a0\u00a0\u00a0length_function=len,\n)\ntext_splitter.create_documents([text])<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">[Document(metadata={}, page_content=\"The Marvel Universe is a vast and<br\/>interconnected world filled with superheroes, villains, and epic<br\/>storytelling that has captivated audiences for decades. Founded by<br\/>visionaries such as Stan Lee, Jack Kirby, and Steve Ditko, Marvel Comics has<br\/>introduced some of the most iconic characters in pop\"),<p>\u00a0Document(metadata={}, page_content=\"culture history. From its early<br\/>beginnings in 1939 as Timely Publications to its transformation into Marvel<br\/>Comics in the 1960s, the company has consistently pushed the boundaries of<br\/>storytelling by creating relatable and dynamic characters. Heroes like<br\/>Spider-Man, Iron Man, Captain America, and\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"Thor have become household names, each<br\/>with their own compelling backstories and struggles that resonate with fans<br\/>across generations. Marvel\u2019s success extends beyond the pages of comic<br\/>books. The launch of the Marvel Cinematic Universe (MCU) in 2008 with the<br\/>release of Iron Man revolutionized the\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"film industry, introducing<br\/>interconnected storylines that culminated in epic crossover events such as<br\/>The Avengers and Infinity War. The MCU\u2019s success is largely attributed to<br\/>its ability to blend action, humor, and emotional depth while maintaining<br\/>the essence of the beloved comic book characters.\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"Audiences have followed the journeys of<br\/>superheroes as they face powerful foes like Thanos and Loki, all while<br\/>dealing with their own internal conflicts and responsibilities.\")]<\/p><\/pre>\n<p>The resulting list of Document objects contains multiple chunks of the text, each with overlapping portions to ensure smooth transitions. Here\u2019s a breakdown of the output:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>First Chunk:<\/strong><strong><br \/><\/strong>\u201cThe Marvel Universe is a vast and interconnected world filled with superheroes, \u2026 iconic characters in pop\u201d<\/li>\n<li><strong>Second Chunk:<\/strong><strong><br \/><\/strong>\u201cculture history. From its early beginnings in 1939 as Timely Publications to its transformation into Marvel Comics in the 1960s, \u2026 Iron Man, Captain America, and\u201d<\/li>\n<li><strong>Third Chunk:<\/strong><strong><br \/><\/strong>\u201cThor have become household names, each with their own compelling backstories and struggles that resonate \u2026 Iron Man revolutionized the\u201d<\/li>\n<li><strong>Fourth Chunk:<\/strong><strong><br \/><\/strong>\u201cfilm industry, introducing interconnected storylines that culminated in epic crossover events such as The Avengers \u2026 comic book characters.\u201d<\/li>\n<li><strong>Fifth Chunk:<\/strong><strong><br \/><\/strong>\u201cAudiences have followed the journeys of superheroes as they face powerful foes like Thanos and Loki, \u2026 responsibilities.\u201d<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-3-document-specific-chunking-using-langchain-html-python-json-or-more\">3. Document Specific Chunking Using LangChain( HTML, Python, JSON or more)<\/h2>\n<p>Document-specific chunking is a strategy designed to tailor text-splitting methods to fit different data formats such as images, PDFs, or code snippets. Unlike generic chunking methods, which may not work effectively across various content types, document-specific chunking takes into account the unique structure and characteristics of each format to ensure meaningful segmentation.<\/p>\n<p>For instance, when dealing with Markdown, Python, or JavaScript files, chunking methods are adapted to use format-specific separators, such as headers in Markdown, function definitions in Python, or code blocks in JavaScript. This approach allows for more accurate and context-aware chunking, ensuring that key elements of the content remain intact and understandable.<\/p>\n<p>By adopting document-specific chunking, organizations and developers can efficiently process diverse data types while maintaining logical segmentation, and improving downstream tasks such as search, summarization, and analysis.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-python\">1. Python<\/h3>\n<pre class=\"wp-block-code\"><code>%pip install -qU langchain-text-splitters\nfrom langchain_text_splitters import (Language,RecursiveCharacterTextSplitter,)\nPYTHON_CODE = \"\"\"\ndef hello_world():\n\u00a0\u00a0\u00a0print(\"Hello, World!\")\n# Call the function\nhello_world()\n\"\"\"\npython_splitter = RecursiveCharacterTextSplitter.from_language(\n\u00a0\u00a0\u00a0language=Language.PYTHON, chunk_size=50, chunk_overlap=0\n)\npython_docs = python_splitter.create_documents([PYTHON_CODE])<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>python_docs<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">[Document(metadata={}, page_content=\"def hello_world():\\n\u00a0 \u00a0 print(\"Hello,<br\/>World!\")\"),<br\/>\u00a0Document(metadata={}, page_content=\"# Call the function\\nhello_world()\")]<\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-2-markdown\">2. Markdown<\/h3>\n<pre class=\"wp-block-code\"><code>%pip install -qU langchain-text-splitters\nfrom langchain_text_splitters import(Language,RecursiveCharacterTextSplitter)\n\nmarkdown_text = \"\"\"# \ud83e\udd9c\ufe0f\ud83d\udd17 LangChain\n\u26a1 Building applications with LLMs through composability \u26a1\n## What is LangChain?\n# Hopefully this code block isn't split\nLangChain is a framework for...\nAs an open-source project in a rapidly developing field, we are extremely open to contributions.\n\"\"\"\n\nmd_splitter = RecursiveCharacterTextSplitter.from_language(\n   language=Language.MARKDOWN, chunk_size=60, chunk_overlap=0\n)\nmd_docs = md_splitter.create_documents([markdown_text])<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>md_docs<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">[Document(metadata={}, page_content=\"# \ud83e\udd9c\ufe0f\ud83d\udd17 LangChain\"),<p>\u00a0Document(metadata={}, page_content=\"\u26a1 Building applications with LLMs through composability \u26a1\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"## What is LangChain?\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"# Hopefully this code block isn't split\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"LangChain is a framework for...\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"As an open-source project in a rapidly developing field, we\"),<\/p><p>\u00a0Document(metadata={}, page_content=\"are extremely open to contributions.\")]<\/p><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-4-semantic-chunking\">4. Semantic Chunking<\/h2>\n<p>Semantic chunking is an advanced text-splitting technique that focuses on dividing a document into meaningful chunks based on the actual content and context rather than arbitrary size-based methods such as token count or delimiters. The primary goal of semantic chunking is to ensure that each chunk contains a <strong>single, concise meaning<\/strong>, optimizing it for downstream tasks like embedding into vector representations for machine learning applications.<\/p>\n<p>Traditional chunking methods, such as splitting text by a fixed number of tokens or characters, often result in chunks that contain multiple, unrelated meanings. This can dilute the representation when encoding text into vector embeddings, leading to suboptimal retrieval and processing results. By contrast, semantic chunking works by identifying natural meaning boundaries within the text and segmenting it accordingly to ensure each chunk preserves a coherent and unified concept.<\/p>\n<p>For example, in a newspaper article, different paragraphs may cover various aspects of a single story. A naive chunking approach may group unrelated sections together, leading to mixed embeddings that fail to represent any of the topics accurately. Semantic chunking, however, isolates sections with distinct meanings, ensuring that each vector embedding captures the core essence of that portion.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-implementing-semantic-chunking\">Implementing Semantic Chunking<\/h4>\n<p>In practice, semantic chunking can be implemented using natural language processing (NLP) techniques such as semantic similarity analysis, topic modeling, or machine learning-based segmentation. These methods analyze the underlying meaning of the text and intelligently determine appropriate chunk boundaries.<\/p>\n<p>By adopting semantic chunking, text processing systems can achieve higher accuracy in tasks such as information retrieval, summarization, and AI-driven insights, ensuring that each chunk represents a concise and meaningful unit of information.<\/p>\n<pre class=\"wp-block-code\"><code>!pip install --quiet langchain_experimental langchain_openai<\/code><\/pre>\n<p>This command installs the required packages:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>langchain_experimental<\/strong>: Provides experimental text-splitting techniques, including semantic chunking.<\/li>\n<li><strong>langchain_openai<\/strong>: Provides access to OpenAI\u2019s embedding models for semantic processing.<\/li>\n<\/ul>\n<p>The \u2013quiet flag suppresses unnecessary output during installation.<\/p>\n<pre class=\"wp-block-code\"><code># This is a long document we can split up.\nwith open(\"state_of_the_union.txt\") as f:\n\u00a0\u00a0\u00a0state_of_the_union = f.read()<\/code><\/pre>\n<p>The state_of_the_union.txt file is read into a string variable state_of_the_union.<\/p>\n<p>This text will later be split into meaningful chunks based on semantic differences.<\/p>\n<pre class=\"wp-block-code\"><code>from langchain_experimental.text_splitter import SemanticChunker\nimport os\nfrom langchain_experimental.text_splitter import SemanticChunker\nfrom langchain_openai.embeddings import OpenAIEmbeddings\nfrom getpass import getpass<\/code><\/pre>\n<ul class=\"wp-block-list\">\n<li><strong>os<\/strong>: Used to manage environment variables such as the API key.<\/li>\n<li><strong>SemanticChunker<\/strong>: The class that performs the semantic chunking process.<\/li>\n<li><strong>OpenAIEmbeddings<\/strong>: Provides access to OpenAI\u2019s embedding models to measure sentence similarity.<\/li>\n<li><strong>getpass<\/strong>: Securely prompts the user for their OpenAI API key.<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>os.environ[\"OPENAI_API_KEY\"] = getpass(\"API\")\ntext_splitter = SemanticChunker(\n\u00a0\u00a0\u00a0OpenAIEmbeddings(), breakpoint_threshold_type=\"percentile\"\n)<\/code><\/pre>\n<p>Initializes the SemanticChunker using OpenAI\u2019s embeddings model.<\/p>\n<p>It will automatically calculate the semantic similarity between sentences to determine where to split the text.<\/p>\n<p>Specifies breakpoint_threshold_type=\u201dpercentile\u201d, which means the chunking decision is based on the percentile method for determining split points.<\/p>\n<pre class=\"wp-block-code\"><code>docs = text_splitter.create_documents([state_of_the_union])\nprint(docs[0].page_content)<\/code><\/pre>\n<ul class=\"wp-block-list\">\n<li>This method processes the input text and splits it into meaningful segments using the chosen semantic chunking strategy.<\/li>\n<li>The result is a list of Document objects, each containing a chunk of text.<\/li>\n<\/ul>\n<p>Semantic chunking works by determining where to split text based on differences in <strong>sentence embeddings<\/strong>, which capture the meaning of sentences numerically. The algorithm calculates the difference in meaning between consecutive sentences and splits them when a certain threshold is exceeded.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-methods-to-determine-breakpoints-threshold-types\">Methods to Determine Breakpoints (Threshold Types)<\/h4>\n<p>The chunking behaviour is controlled using the breakpoint_threshold_type parameter, which supports the following methods:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Percentile (Default Method)<\/strong>\n<ul class=\"wp-block-list\">\n<li>Measures the differences between sentence embeddings and splits the text at the top X percentile.<\/li>\n<li>The default percentile is 95.0, adjustable via breakpoint_threshold_amount.<\/li>\n<li>Example: If the differences between sentences follow a distribution, the method splits the largest 5% of differences.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Standard Deviation<\/strong>\n<ul class=\"wp-block-list\">\n<li>Splits chunks when the difference exceeds X standard deviations from the mean.<\/li>\n<li>The default value for X is 3.0.<\/li>\n<li>This method is useful when text has uniform patterns with occasional significant changes.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Interquartile Range (IQR)<\/strong>\n<ul class=\"wp-block-list\">\n<li>Uses statistical quartiles to determine split points by identifying outliers in semantic changes.<\/li>\n<li>The default scaling factor is 1.5, adjustable via breakpoint_threshold_amount.<\/li>\n<li>Effective for texts with moderate variation in meaning.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Gradient-Based Splitting<\/strong>\n<ul class=\"wp-block-list\">\n<li>Uses the gradient of embedding distance to identify split points, applying anomaly detection techniques.<\/li>\n<li>Suitable for domain-specific texts (e.g., legal or medical documents) where topic shifts are subtle.<\/li>\n<li>Works similarly to the percentile method but adapts to highly correlated data.<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-5-agentic-chunking\">5. Agentic Chunking<\/h2>\n<p>Agentic chunking is an advanced method of segmenting documents into smaller, meaningful sections by leveraging a large language model (LLM) to identify natural breakpoints in the text. Unlike traditional chunking methods that rely on fixed character counts, agentic chunking analyzes the content to detect semantically relevant boundaries such as paragraph breaks and topic transitions.<\/p>\n<p>By using AI to determine logical divisions within the text, agentic chunking ensures that each chunk retains contextual integrity and meaning, improving the AI\u2019s ability to process, summarize, and respond effectively. This approach enhances information retrieval, content organization, and decision-making processes by creating well-structured, purpose-driven text segments.<\/p>\n<p>Agentic chunking is particularly useful in applications such as knowledge retrieval, automated summarization, and AI-driven insights, where maintaining coherence and relevance is crucial for optimal performance.<\/p>\n<p>Note: Most people refer to it as Agentic Chunking, but it\u2019s primarily based on LLM-driven chunking.<\/p>\n<p>Talking about the LLM-based Chunking \u2013 It is essentially the process of using a <strong>large language model (LLM)<\/strong>\u2014like GPT-4\u2014to <strong>break down or segment text into more manageable, structured pieces<\/strong>. Instead of using rigid rules (like splitting strictly on sentence boundaries or punctuation), LLM-based chunking leverages the <strong>model\u2019s understanding of language and context<\/strong> to produce chunks in a way that\u2019s more <strong>meaningful and coherent<\/strong>.<\/p>\n<pre class=\"wp-block-code\"><code>!pip install agno openai<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>from typing import List, Optional\nfrom agno.document.base import Document\nfrom agno.document.chunking.strategy import ChunkingStrategy\nfrom agno.models.base import Model\nfrom agno.models.defaults import DEFAULT_OPENAI_MODEL_ID\nfrom agno.models.message import Message\nfrom agno.models.openai import OpenAIChat<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>import os\nos.environ[\"OPENAI_API_KEY\"] = \"your_api_key\"\n\nclass AgenticChunking(ChunkingStrategy):\n   \"\"\"Chunking strategy that uses an LLM to determine natural breakpoints in the text\"\"\"\n\n\n   def __init__(self, model: Optional[Model] = None, max_chunk_size: int = 5000):\n       if \"OPENAI_API_KEY\" not in os.environ:\n           raise ValueError(\"OPENAI_API_KEY environment variable not set.\")\n       self.model = model or OpenAIChat(DEFAULT_OPENAI_MODEL_ID)\n       self.max_chunk_size = max_chunk_size\n\n\n   def chunk(self, document: Document) -&gt; List[Document]:\n       \"\"\"Split text into chunks using LLM to determine natural breakpoints based on context\"\"\"\n       if len(document.content) <\/code><\/pre>\n<pre class=\"wp-block-code\"><code># Example usage\ndocument = Document(\n   id=\"doc1\",\n   content=\"\"\"Recursive chunking divides the input text into smaller chunks in a hierarchical and iterative manner using a set of separators. If the initial attempt at splitting the text doesn\u2019t produce chunks of the desired size or structure, the method recursively calls itself on the resulting chunks with a different separator or criterion until the desired chunk size or structure is achieved. This means that while the chunks aren\u2019t going to be exactly the same size, they\u2019ll still \u201caspire\u201d to be of a similar size.\"\"\",\n   meta_data={\"author\": \"Pankaj\"}\n)\n\nchunker = AgenticChunking(max_chunk_size=200)\nchunks = chunker.chunk(document)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code># Print all chunks\nfor i, chunk in enumerate(chunks, 1):\n   print(f\"Chunk {i} (ID: {chunk.id}, Size: {len(chunk.content)})\")\n   print(chunk.content)\n   print(\"-\" * 50 + \"\\n\")<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">Chunk 1 (ID: doc1_1, Size: 179)<br\/>Recursive chunking divides the input text into smaller chunks in a<br\/>hierarchical and iterative manner using a set of separators. If the initial<br\/>attempt at splitting the text doesn\u2019<br\/>--------------------------------------------------<p>Chunk 2 (ID: doc1_2, Size: 132)<br\/>t produce chunks of the desired size or structure, the method recursively<br\/>calls itself on the resulting chunks with a different sepa<br\/>--------------------------------------------------<\/p><p>Chunk 3 (ID: doc1_3, Size: 104)<br\/>rator or criterion until the desired chunk size or structure is achieved.<br\/>This means that while the chun<br\/>--------------------------------------------------<\/p><p>Chunk 4 (ID: doc1_4, Size: 66)<br\/>ks aren\u2019t going to be exactly the same size, they\u2019ll still \u201caspire<br\/>--------------------------------------------------<\/p><p>Chunk 5 (ID: doc1_5, Size: 26)<br\/>\u201d to be of a similar size.<br\/>--------------------------------------------------<\/p><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-llm-based-chunking-using-openai-library\">LLM-Based Chunking Using OpenAI Library<\/h3>\n<pre class=\"wp-block-code\"><code>from openai import OpenAI<\/code><\/pre>\n<p>Imports the OpenAI library, required to interact with the GPT API.<\/p>\n<pre class=\"wp-block-code\"><code>content = \"An outlier is a data point that significantly deviates from the rest of the data. It can be either much higher or much lower than the other data points, and its pr types of outliers: There are two main types of outliers: Global outliers: Global outliers are isolated data points that are far away from the main body of the data\"<\/code><\/pre>\n<p>This is the input text that will be chunked.<\/p>\n<pre class=\"wp-block-code\"><code># Initialize client with your API key\nclient = OpenAI(api_key=\"API_KEY\")<\/code><\/pre>\n<p>Initializes the OpenAI client using an API key (replace \u201cAPI_KEY\u201d with an actual key to run the code).<\/p>\n<pre class=\"wp-block-code\"><code>response = client.chat.completions.create(\n\u00a0\u00a0\u00a0model=\"gpt-4o\",\n\u00a0\u00a0\u00a0messages=[\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"role\": \"system\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"role\": \"system\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"content\": \"\"\"You are a agentic chunker. Decompose the content into clear and simple propositions:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a01. Split compound sentences into simple sentences\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a02. Separate named entities with descriptions\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a03. Replace pronouns with specific references\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a04. Output as JSON list of strings\"\"\"\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0},\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"role\": \"user\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"content\": f\"Here is the content: {content}\"\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0}\n\u00a0\u00a0\u00a0],\n\u00a0\u00a0\u00a0temperature=0.3\n)<\/code><\/pre>\n<p>Model: Uses gpt-4o for processing.<\/p>\n<p>Messages: The system message defines GPT\u2019s behavior: breaking down text into simple propositions, separating named entities, avoiding pronouns, and outputting as a JSON list.<\/p>\n<p>The user message provides the actual content for chunking.<br \/>Temperature: 0.3 keeps responses deterministic, reducing randomness for more consistent outputs.<\/p>\n<pre class=\"wp-block-code\"><code>print(response.choices[0].message.content)<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">\"An outlier is a data point that significantly deviates from the rest of the data.\",<p>\u00a0\u00a0\"An outlier can be much higher than the other data points.\",<\/p><p>\u00a0\u00a0\"An outlier can be much lower than the other data points.\",<\/p><p>\u00a0\u00a0\"There are two main types of outliers.\",<\/p><p>\u00a0\u00a0\"Global outliers are isolated data points.\",<\/p><p>\u00a0\u00a0\"Global outliers are far away from the main body of the data.\"<\/p><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-6-section-based-chunking\">6. Section Based Chunking<\/h2>\n<p>Section-based chunking is a technique used to divide large texts into meaningful \u201cchunks\u201d or segments based on structural elements like headings, subheadings, paragraphs, or predefined section markers. Unlike topic modeling (which relies on statistical patterns to group content), section-based chunking leverages the document\u2019s inherent structure to create logical divisions.<\/p>\n<p><strong>Structure-Driven:<\/strong><strong><br \/><\/strong>Relies on document formatting like:<\/p>\n<ul class=\"wp-block-list\">\n<li>Headings (e.g., Introduction, Methods, Conclusion)<\/li>\n<li>Numbered sections (e.g., 1.1, 2.3.4)<\/li>\n<li>Bullet points, line breaks, or custom markers.<\/li>\n<\/ul>\n<p><strong>Preserves Context:<\/strong><strong><br \/><\/strong>Keeps related information together, maintaining narrative flow within sections.<\/p>\n<p><strong>Efficient for Structured Documents:<\/strong><strong><br \/><\/strong>Works well with academic papers, reports, PDFs, legal documents, etc.<\/p>\n<pre class=\"wp-block-code\"><code>from sklearn.feature_extraction.text import CountVectorizer\nfrom sklearn.decomposition import LatentDirichletAllocation\nimport numpy as np\nimport fitz\u00a0 # PyMuPDF<\/code><\/pre>\n<p>Function to extract text from a PDF file<\/p>\n<pre class=\"wp-block-code\"><code>def extract_text_from_pdf(pdf_path):\n\u00a0\u00a0\u00a0pdf_document = fitz.open(pdf_path)\n\u00a0\u00a0\u00a0text = \"\"\n\u00a0\u00a0\u00a0for page in pdf_document:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0text += page.get_text()\n\u00a0\u00a0\u00a0return text<\/code><\/pre>\n<p>Topic-based chunking function<\/p>\n<pre class=\"wp-block-code\"><code>def topic_based_chunk(text, num_topics=3):\n\u00a0\u00a0\u00a0sentences = text.split('. ')\n\u00a0\u00a0\u00a0vectorizer = CountVectorizer()\n\u00a0\u00a0\u00a0sentence_vectors = vectorizer.fit_transform(sentences)\n\u00a0\u00a0\u00a0lda = LatentDirichletAllocation(n_components=num_topics, random_state=42)\n\u00a0\u00a0\u00a0lda.fit(sentence_vectors)\n\u00a0\u00a0\u00a0topic_word = lda.components_\n\u00a0\u00a0\u00a0vocabulary = vectorizer.get_feature_names_out()\n\u00a0\u00a0\u00a0topics = []\n\u00a0\u00a0\u00a0for topic_idx, topic in enumerate(topic_word):\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0top_words_idx = topic.argsort()[:-6:-1]\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0topic_keywords = [vocabulary[i] for i in top_words_idx]\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0topics.append(f\"Topic {topic_idx + 1}: {', '.join(topic_keywords)}\")\n\u00a0\u00a0\u00a0chunks_with_topics = []\n\u00a0\u00a0\u00a0for i, sentence in enumerate(sentences):\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0topic_assignments = lda.transform(vectorizer.transform([sentence]))\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0assigned_topic = np.argmax(topic_assignments)\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0chunks_with_topics.append((topics[assigned_topic], sentence))\n\u00a0\u00a0\u00a0return chunks_with_topics<\/code><\/pre>\n<p>Replace \u2018your_file.pdf\u2019 with your actual PDF file path<\/p>\n<pre class=\"wp-block-code\"><code>pdf_path=\"\/content\/1738082270933.pdf\"\npdf_text = extract_text_from_pdf(pdf_path)<\/code><\/pre>\n<p>Get topic-based chunks<\/p>\n<pre class=\"wp-block-code\"><code>topic_chunks = topic_based_chunk(pdf_text, num_topics=3)<\/code><\/pre>\n<p>Display results<\/p>\n<pre class=\"wp-block-code\"><code>for topic, chunk in topic_chunks:\n\u00a0\u00a0\u00a0print(f\"{topic}: {chunk}\\n\")<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">Topic 3: reasoning, r1, deepseek, the, of:\u00a0<p>DeepSeek-R1 is a reasoning-focused large language model (LLM) developed to<br\/>enhance reasoning capabilities in Generative AI systems through advanced<br\/>reinforcement learning techniques.<\/p><\/pre>\n<p><em><strong>Explanation: Topic 3<\/strong> is characterized by keywords like <strong>\u201creasoning,\u201d \u201cR1,\u201d \u201cDeepSeek\u201d<\/strong>, which frequently appear in sentences about the DeepSeek model.<\/em><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-7-contextual-chunking\">7. Contextual Chunking<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"900\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-13-67a0e0c32086c.webp\" alt=\"7. Contextual Chunking\" class=\"wp-image-219390\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-13-67a0e0c32086c.webp 1600w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-13-67a0e0c32086c-300x169.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-13-67a0e0c32086c-768x432.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-13-67a0e0c32086c-1536x864.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/unnamed-13-67a0e0c32086c-150x84.webp 150w\" sizes=\"auto, (max-width: 1600px) 100vw, 1600px\"\/><figcaption class=\"wp-element-caption\">Source: Anthropic<\/figcaption><\/figure>\n<\/div>\n<p>Contextual Chunking in Retrieval-Augmented Generation (RAG) refers to the strategy of segmenting documents or data into meaningful \u201cchunks\u201d that preserve the semantic context. This technique enhances the retrieval and generation performance of RAG models by ensuring that the model has access to coherent, context-rich pieces of information, rather than arbitrary or fragmented text segments.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-why-is-it-important\">Why Is It Important?<\/h4>\n<p>In RAG systems, the process involves two main steps:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Retrieval<\/strong>: Finding relevant chunks from a large knowledge base.<\/li>\n<li><strong>Generation<\/strong>: Using the retrieved chunks to produce a coherent response.<\/li>\n<\/ol>\n<p>If the chunks are poorly segmented, the retrieval process might fetch incomplete or contextually weak information, leading to subpar generation quality. <strong>Contextual chunking<\/strong> helps mitigate this by ensuring that each chunk contains enough semantic information to be useful on its own.<\/p>\n<p>Here\u2019s how you set the chunk process prompt for contextual chunking:\u00a0<\/p>\n<pre class=\"wp-block-code\"><code># create chunk context generation chain\nfrom langchain.prompts import ChatPromptTemplate\nfrom langchain.schema import StrOutputParser\nfrom langchain_openai import ChatOpenAI\nchatgpt = ChatOpenAI(model_name=\"gpt-4o-mini\", temperature=0)\ndef generate_chunk_context(document, chunk):\n\u00a0\u00a0\u00a0chunk_process_prompt = \"\"\"You are an AI assistant specializing in research\u00a0\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0paper analysis. Your task is to provide brief,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0relevant context for a chunk of text based on the\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0following research paper.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Here is the research paper:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<paper>\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{paper}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/paper>\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Here is the chunk we want to situate within the whole\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0document:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<chunk>\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{chunk}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/chunk>\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Provide a concise context (3-4 sentences max) for this\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0chunk, considering the following guidelines:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0- Give a short succinct context to situate this chunk\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0within the overall document for the purposes of\u00a0\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0improving search retrieval of the chunk.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0- Answer only with the succinct context and nothing\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0else.\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0- Context should be mentioned like 'Focuses on ....'\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0do not mention 'this chunk or section focuses on...'\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0Context:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"\"\"\n\u00a0\u00a0\u00a0prompt_template = ChatPromptTemplate.from_template(chunk_process_prompt)\n\u00a0\u00a0\u00a0agentic_chunk_chain = (prompt_template\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0|\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0chatgpt\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0|\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0StrOutputParser())\n\u00a0\u00a0\u00a0context = agentic_chunk_chain.invoke({'paper': document, 'chunk': chunk})\n\u00a0\u00a0\u00a0return context<\/code><\/pre>\n<p>For more information, refer to this article \u2013 <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/12\/contextual-rag-systems-with-hybrid-search-and-reranking\/\" target=\"_blank\" rel=\"noreferrer noopener\">A Comprehensive Guide to Building Contextual RAG Systems with Hybrid Search and Reranking<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-8-late-chunking\">8. Late Chunking<\/h2>\n<p>Late Chunking addresses the challenges of maintaining contextual coherence when processing long documents for retrieval applications. Unlike traditional chunking approaches that segment text early in the pipeline, potentially disrupting long-distance contextual dependencies, Late Chunking leverages long-context embedding models to generate contextual chunk embeddings. This ensures that references spread across multiple text segments (like pronouns or entity mentions) are preserved within their broader context, leading to higher-quality vector representations and more effective retrieval performance. This method mitigates the shortcomings of conventional RAG pipelines, particularly in handling anaphoric references and fragmented information.<\/p>\n<p>To see how Jina Embeddings works explore this: <a href=\"https:\/\/jina.ai\/api-dashboard\/embedding\">Jina Embeddings<\/a>.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-how-late-chunking-works\">How Late Chunking Works?<\/h4>\n<p>When breaking down a Wikipedia article into smaller chunks, phrases like \u201cits\u201d or \u201cthe city\u201d often refer back to something mentioned earlier, such as \u201cBerlin\u201d in the first sentence. However, splitting the text disconnects these references from the original entity, making it difficult for embedding models to correctly associate them with \u201cBerlin.\u201d This results in less accurate vector representations and weaker performance in retrieval-augmented generation (RAG) systems.<\/p>\n<p>Late Chunking addresses this issue by processing the entire text\u2014or as much of it as possible\u2014through the transformer layer of the embedding model before splitting it into chunks. This approach generates token-level vector representations that capture the full context of the text. Afterward, the system applies mean pooling to each chunk to create embeddings, ensuring they retain important contextual information since the full text was initially considered.<\/p>\n<p>Unlike basic chunking methods that process each chunk in isolation, Late Chunking allows every chunk to retain influence from the broader document context. As a result, references like \u201cits\u201d and \u201cthe city\u201d remain correctly associated with \u201cBerlin,\u201d even when appearing in different chunks. This improves RAG systems\u2019 accuracy, making them more context-aware and capable of delivering better, more coherent answers.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-implementation-and-performance-gains\">Implementation and Performance Gains<\/h4>\n<pre class=\"wp-block-code\"><code>!pip install transformers==4.43.4<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>from transformers import AutoModel\nfrom transformers import AutoTokenizer<\/code><\/pre>\n<p># load model and tokenizer<\/p>\n<pre class=\"wp-block-code\"><code>tokenizer = AutoTokenizer.from_pretrained('jinaai\/jina-embeddings-v2-base-en', trust_remote_code=True)\nmodel = AutoModel.from_pretrained('jinaai\/jina-embeddings-v2-base-en', trust_remote_code=True)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>def chunk_by_sentences(input_text: str, tokenizer: callable):\n\u00a0\u00a0\u00a0\"\"\"\n\u00a0\u00a0\u00a0Split the input text into sentences using the tokenizer\n\u00a0\u00a0\u00a0:param input_text: The text snippet to split into sentences\n\u00a0\u00a0\u00a0:param tokenizer: The tokenizer to use\n\u00a0\u00a0\u00a0:return: A tuple containing the list of text chunks and their corresponding token spans\n\u00a0\u00a0\u00a0\"\"\"\n\u00a0\u00a0\u00a0inputs = tokenizer(input_text, return_tensors=\"pt\", return_offsets_mapping=True)\n\u00a0\u00a0\u00a0punctuation_mark_id = tokenizer.convert_tokens_to_ids('.')\n\u00a0\u00a0\u00a0sep_id = tokenizer.convert_tokens_to_ids('[SEP]')\n\u00a0\u00a0\u00a0token_offsets = inputs['offset_mapping'][0]\n\u00a0\u00a0\u00a0token_ids = inputs['input_ids'][0]\n\u00a0\u00a0\u00a0chunk_positions = [\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0(i, int(start + 1))\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0for i, (token_id, (start, end)) in enumerate(zip(token_ids, token_offsets))\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0if token_id == punctuation_mark_id\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0and (\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0token_offsets[i + 1][0] - token_offsets[i][1] &gt; 0\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0or token_ids[i + 1] == sep_id\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0)\n\u00a0\u00a0\u00a0]\n\u00a0\u00a0\u00a0chunks = [\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0input_text[x[1] : y[1]]\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0for x, y in zip([(1, 0)] + chunk_positions[:-1], chunk_positions)\n\u00a0\u00a0\u00a0]\n\u00a0\u00a0\u00a0span_annotations = [\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0(x[0], y[0]) for (x, y) in zip([(1, 0)] + chunk_positions[:-1], chunk_positions)\n\u00a0\u00a0\u00a0]\n\u00a0\u00a0\u00a0return chunks, span_annotations<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>import requests\ndef chunk_by_tokenizer_api(input_text: str, tokenizer: callable):\n\u00a0\u00a0\u00a0# Define the API endpoint and payload\n\u00a0\u00a0\u00a0url=\"https:\/\/tokenize.jina.ai\/\"\n\u00a0\u00a0\u00a0payload = {\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"content\": input_text,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"return_chunks\": \"true\",\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"max_chunk_length\": \"1000\"\n\u00a0\u00a0\u00a0}\n\u00a0\u00a0\u00a0# Make the API request\n\u00a0\u00a0\u00a0response = requests.post(url, json=payload)\n\u00a0\u00a0\u00a0response_data = response.json()\n\u00a0\u00a0\u00a0# Extract chunks and positions from the response\n\u00a0\u00a0\u00a0chunks = response_data.get(\"chunks\", [])\n\u00a0\u00a0\u00a0chunk_positions = response_data.get(\"chunk_positions\", [])\n\u00a0\u00a0\u00a0# Adjust chunk positions to match the input format\n\u00a0\u00a0\u00a0span_annotations = [(start, end) for start, end in chunk_positions]\n\u00a0\u00a0\u00a0return chunks, span_annotations<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>nput_text = \"Berlin is the capital and largest city of Germany, both by area and by population. Its more than 3.85 million inhabitants make it the European Union's most populous city, as measured by population within city limits. The city is also one of the states of Germany, and is the third smallest state in the country in terms of area.\"\n\n# determine chunks\nchunks, span_annotations = chunk_by_sentences(input_text, tokenizer)\nprint('Chunks:\\n- \"' + '\"\\n- \"'.join(chunks) + '\"')<\/code><\/pre>\n<pre class=\"wp-block-preformatted\">Chunks:<p>- \"Berlin is the capital and largest city of Germany, both by area and by<br\/>population.\"<\/p><p>- \" Its more than 3.85 million inhabitants make it the European Union's most<br\/>populous city, as measured by population within city limits.\"<\/p><p>- \" The city is also one of the states of Germany, and is the third smallest<br\/>state in the country in terms of area.\"<\/p><\/pre>\n<pre class=\"wp-block-code\"><code>def late_chunking(\n\u00a0\u00a0\u00a0model_output: 'BatchEncoding', span_annotation: list, max_length=None\n):\n\u00a0\u00a0\u00a0token_embeddings = model_output[0]\n\u00a0\u00a0\u00a0outputs = []\n\u00a0\u00a0\u00a0for embeddings, annotations in zip(token_embeddings, span_annotation):\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0if (\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0max_length is not None\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0):\u00a0 # remove annotations which go beyond the max-length of the model\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0annotations = [\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0(start, min(end, max_length - 1))\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0for (start, end) in annotations\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0if start = 1\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0]\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pooled_embeddings = [\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0embedding.detach().cpu().numpy() for embedding in pooled_embeddings\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0]\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0outputs.append(pooled_embeddings)\n\u00a0\u00a0\u00a0return outputs<\/code><\/pre>\n<pre class=\"wp-block-code\"><code># chunk before\nembeddings_traditional_chunking = model.encode(chunks)\n# chunk afterwards (context-sensitive chunked pooling)\ninputs = tokenizer(input_text, return_tensors=\"pt\")\nmodel_output = model(**inputs)\nembeddings = late_chunking(model_output, [span_annotations])[0]<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>import numpy as np\ncos_sim = lambda x, y: np.dot(x, y) \/ (np.linalg.norm(x) * np.linalg.norm(y))\nberlin_embedding = model.encode('Berlin')\nfor chunk, new_embedding, trad_embeddings in zip(chunks, embeddings, embeddings_traditional_chunking):\n\u00a0\u00a0\u00a0print(f'similarity_new(\"Berlin\", \"{chunk}\"):', cos_sim(berlin_embedding, new_embedding))\n\u00a0\u00a0\u00a0print(f'similarity_trad(\"Berlin\", \"{chunk}\"):', cos_sim(berlin_embedding, trad_embeddings))<\/code><\/pre>\n<p><strong>Output<\/strong><\/p>\n<pre class=\"wp-block-preformatted\">similarity_new(\"Berlin\", \"Berlin is the capital and largest city of Germany,<br\/>both by area and by population.\"): 0.849546<p>similarity_trad(\"Berlin\", \"Berlin is the capital and largest city of Germany,<br\/>both by area and by population.\"): 0.8486219<\/p><p>similarity_new(\"Berlin\", \" Its more than 3.85 million inhabitants make it the<br\/>European Union's most populous city, as measured by population within city<br\/>limits.\"): 0.82489026<\/p><p>similarity_trad(\"Berlin\", \" Its more than 3.85 million inhabitants make it<br\/>the European Union's most populous city, as measured by population within<br\/>city limits.\"): 0.70843387<\/p><p>similarity_new(\"Berlin\", \" The city is also one of the states of Germany, and<br\/>is the third smallest state in the country in terms of area.\"): 0.8498009<\/p><p>similarity_trad(\"Berlin\", \" The city is also one of the states of Germany,<br\/>and is the third smallest state in the country in terms of area.\"):\u00a0<\/p><p>0.75345534<\/p><\/pre>\n<p>Here in the output, you can clearly see there is improvement in the semantic similarity.\u00a0<\/p>\n<p><strong>General Performance Improvement:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>Across all examples, the similarity_new scores are consistently higher than similarity_trad. This indicates that late chunking more effectively captures semantic relationships.<\/li>\n<li>For example:\n<ul class=\"wp-block-list\">\n<li><strong>\u201cBerlin\u201d vs. \u201cThe city is also one of the states of Germany\u2026\u201d<\/strong>\n<ul class=\"wp-block-list\">\n<li>similarity_new: <strong>0.8498<\/strong><\/li>\n<li>similarity_trad: <strong>0.7535<\/strong><\/li>\n<li>The <strong>0.0963<\/strong> improvement highlights better contextual linkage between \u201cthe city\u201d and \u201cBerlin.\u201d<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p><strong>Notable Improvements in Ambiguous References:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>The most significant improvement occurs when dealing with indirect references like <strong>\u201cthe city\u201d<\/strong> instead of explicitly repeating \u201cBerlin.\u201d<\/li>\n<li>In:\n<ul class=\"wp-block-list\">\n<li><strong>\u201cBerlin\u201d vs. \u201cIts more than 3.85 million inhabitants\u2026\u201d<\/strong>\n<ul class=\"wp-block-list\">\n<li>similarity_new: <strong>0.8249<\/strong><\/li>\n<li>similarity_trad: <strong>0.7084<\/strong><\/li>\n<li>The <strong>0.1165<\/strong> difference suggests that late chunking strengthens connections even when the entity isn\u2019t explicitly named.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p><strong>Consistency Across Examples:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>While the traditional method maintains decent performance with direct mentions of \u201cBerlin,\u201d it struggles more with pronouns or indirect references.<\/li>\n<li>The new method sustains high similarity scores even when contextual clues are sparse, reflecting improved semantic memory over longer passages.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>Chunking for RAG systems to manage and optimise data processing is crucial to creating a reliable application. Various chunking strategies\u2014ranging from simple character-based splits to advanced methods like semantic, agentic, and late chunking\u2014help improve data retrievability, contextual relevance, and model performance. Selecting the right chunking approach depends on content type, task requirements, and desired output quality, making it an essential practice for efficient AI-powered applications.<\/p>\n<p>If you find this article helpful then, comment below!<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/pankaj9786\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_Lb7Lh0T.webp\" width=\"48\" height=\"48\" alt=\"Pankaj Singh\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>                Hi, I am Pankaj Singh Negi &#8211; Senior Content Editor | Passionate about storytelling and crafting compelling narratives that transform ideas into impactful content. I love reading about technology revolutionizing our lifestyle.                 <\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>The easiest way to learn anything\u2014whether it\u2019s for academics or personal growth\u2014is by breaking it down into smaller, more manageable chunks. Similarly, when you\u2019re tackling a complex subject, it can feel overwhelming at first. However, by dividing it into bite-sized pieces, understanding becomes much easier. Even if it seems like a small concept already, it\u2019s [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":67078,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[35939,32726,11355,14149],"dealstore":[],"offerexpiration":[],"class_list":["post-67077","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-chunking","tag-rag","tag-systems","tag-types"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>8 Types of Chunking for RAG Systems - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=67077\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"8 Types of Chunking for RAG Systems - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"The easiest way to learn anything\u2014whether it\u2019s for academics or personal growth\u2014is by breaking it down into smaller, more manageable chunks. Similarly, when you\u2019re tackling a complex subject, it can feel overwhelming at first. However, by dividing it into bite-sized pieces, understanding becomes much easier. Even if it seems like a small concept already, it\u2019s [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=67077\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-04T02:46:29+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"35 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=67077#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=67077\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"8 Types of Chunking for RAG Systems\",\"datePublished\":\"2025-02-04T02:46:29+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=67077\"},\"wordCount\":4340,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=67077#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp\",\"keywords\":[\"Chunking\",\"RAG\",\"Systems\",\"Types\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=67077#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=67077\",\"url\":\"https:\/\/fivemor.com\/?p=67077\",\"name\":\"8 Types of Chunking for RAG Systems - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=67077#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=67077#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp\",\"datePublished\":\"2025-02-04T02:46:29+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=67077#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=67077\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=67077#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=67077#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"8 Types of Chunking for RAG Systems\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"8 Types of Chunking for RAG Systems - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=67077","og_locale":"en_US","og_type":"article","og_title":"8 Types of Chunking for RAG Systems - Som2ny Network","og_description":"The easiest way to learn anything\u2014whether it\u2019s for academics or personal growth\u2014is by breaking it down into smaller, more manageable chunks. Similarly, when you\u2019re tackling a complex subject, it can feel overwhelming at first. However, by dividing it into bite-sized pieces, understanding becomes much easier. Even if it seems like a small concept already, it\u2019s [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=67077","og_site_name":"Som2ny Network","article_published_time":"2025-02-04T02:46:29+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"35 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=67077#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=67077"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"8 Types of Chunking for RAG Systems","datePublished":"2025-02-04T02:46:29+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=67077"},"wordCount":4340,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=67077#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp","keywords":["Chunking","RAG","Systems","Types"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=67077#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=67077","url":"https:\/\/fivemor.com\/?p=67077","name":"8 Types of Chunking for RAG Systems - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=67077#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=67077#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp","datePublished":"2025-02-04T02:46:29+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=67077#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=67077"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=67077#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/chunking-in-rag.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=67077#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"8 Types of Chunking for RAG Systems"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/67077","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=67077"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/67077\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/67078"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=67077"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=67077"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=67077"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=67077"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=67077"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}