{"id":74960,"date":"2025-02-08T03:24:04","date_gmt":"2025-02-08T03:24:04","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/contextual-retrieval-for-multimodal-rag-on-slide-decks\/"},"modified":"2025-02-08T03:24:04","modified_gmt":"2025-02-08T03:24:04","slug":"contextual-retrieval-for-multimodal-rag-on-slide-decks","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=74960","title":{"rendered":"Contextual Retrieval for Multimodal RAG on Slide Decks"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Imagine a world where finding information in a document is as easy as asking a question\u2014and getting a response that combines both text and images seamlessly. In this guide, we dive into building a Multimodal<a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/09\/retrieval-augmented-generation-rag-in-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\"> Retrieval-Augmented Generation<\/a> pipeline that can do just that. You\u2019ll learn how to parse text and images from a PDF slide deck using tools like LlamaParse, create contextual summaries for enhanced retrieval, and feed this data into advanced models like GPT-4 for query answering. Along the way, we\u2019ll explore how contextual retrieval improves accuracy, optimize costs with prompt caching, and compare results between baseline and enhanced pipelines. Get ready to unlock the potential of RAG with this step-by-step walkthrough! <\/p>\n<h3 class=\"wp-block-heading\" id=\"h-learning-objectives\">Learning Objectives<\/h3>\n<ul class=\"wp-block-list\">\n<li>Understand how to parse PDF slide decks for text and images using LlamaParse.<\/li>\n<li>Learn to add contextual summaries to text chunks for improved retrieval accuracy.<\/li>\n<li>Build a Multimodal RAG pipeline combining text and images with LlamaIndex.<\/li>\n<li>Explore the integration of multimodal data into models like GPT-4.<\/li>\n<li>Compare retrieval performance between baseline and contextual indices.<\/li>\n<\/ul>\n<p><em><strong>This article was published as a part of the\u00a0<\/strong><\/em><a href=\"https:\/\/www.analyticsvidhya.com\/datahack\/blogathon\" target=\"_blank\" rel=\"noreferrer noopener\"><em><strong>Data Science Blogathon.<\/strong><\/em><\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-building-a-contextual-multimodal-rag-pipeline\">Building a Contextual Multimodal RAG Pipeline<\/h2>\n<p>Contextual retrieval was initially introduced in this Anthropic\u00a0<a href=\"https:\/\/www.anthropic.com\/news\/contextual-retrieval\" target=\"_blank\" rel=\"nofollow noopener\">blog post<\/a>. The high-level intuition is that every chunk is given a concise summary of where that chunk fits in with respect to the overall summary of the document. This allows insertion of high-level concepts\/keywords that enable this chunk to be better retrieved for different types of queries.<\/p>\n<p>These <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/03\/an-introduction-to-large-language-models-llms\/\" target=\"_blank\" rel=\"noreferrer noopener\">LLM<\/a> calls are expensive. Contextual retrieval depends on\u00a0prompt caching\u00a0in order to be efficient.<\/p>\n<p>In this notebook, we use <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/06\/claude-3-5-sonnet\/#:~:text=Claude%203.5%20Sonnet%20is%20the,transcribes%20text%20from%20imperfect%20images.\" target=\"_blank\" rel=\"noreferrer noopener\">Claude 3.5-Sonnet <\/a>to generate contextual summaries. We cache the document as text tokens, but generate contextual summaries by feeding in the parsed text chunk.<\/p>\n<p>We feed both the text and image chunks into the final multimodal RAG pipeline to generate the response.<\/p>\n<p>In a Retrieval-Augmented Generation (RAG) pipeline, we typically:<\/p>\n<ul class=\"wp-block-list\">\n<li>Parse our source data (e.g. PDF documents, images, slides).<\/li>\n<li>Embed and index chunks of text for retrieval.<\/li>\n<li>Retrieve relevant chunks for a given query.<\/li>\n<li>Synthesize a response by feeding the retrieved chunks (and, optionally, any relevant images or additional metadata) into a Large Language Model (LLM).<\/li>\n<\/ul>\n<p>Contextual Retrieval is a neat enhancement to standard RAG. Each chunk of text is annotated with a short summary that situates it within the broader document context. This helps the retriever pick the chunk more accurately for queries that might not match the exact words but relate to the overall topic or concept.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-overview-of-the-multimodal-rag-pipeline\">Overview of the Multimodal RAG Pipeline<\/h3>\n<p>We\u2019ll demonstrate how to build a <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/09\/guide-to-building-multimodal-rag-systems\/\" target=\"_blank\" rel=\"noreferrer noopener\">Multimodal RAG <\/a>pipeline over a PDF slide deck, using:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Anthropic<\/b> as our main LLM (Claude 3.5-Sonnet).<\/li>\n<li><b>VoyageAI<\/b> embeddings for chunk embedding.<\/li>\n<li><b>LlamaIndex <\/b>for our retrieval\/indexing abstractions.<\/li>\n<li><b>LlamaParse<\/b> for extracting text and images from the PDF slides.<\/li>\n<li><b>OpenAI GPT-4<\/b> style multimodal model for final query answering (in text+image mode).<\/li>\n<\/ul>\n<p>We will also show how to cache LLM calls to minimize costs, since Contextual Retrieval can generate a lot of prompt calls.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-environment-setup-and-dependencies\">Environment Setup and Dependencies<\/h2>\n<p>You\u2019ll need to install or upgrade a few packages:<\/p>\n<pre class=\"wp-block-code\"><code>!pip install -U llama-index llama-parse\n!pip install -U llama-index-callbacks-arize-phoenix<\/code><\/pre>\n<p><b>Additionally<\/b>:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Anthropic API Key:<\/b> Set os.environ[\u201cANTHROPIC_API_KEY\u201d] = \u201c\u201d.<\/li>\n<li><b>VoyageAI API Key:<\/b> Set os.environ[\u201cVOYAGE_API_KEY\u201d] = \u201c\u201d.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-setup-observability-with-llamatrace-arize-integration\">Setup Observability with LlamaTrace (Arize Integration)<\/h3>\n<p>We setup an integration with LlamaTrace (integration with Arize).<\/p>\n<p>If you haven\u2019t already done so, make sure to create an account here:\u00a0<a href=\"https:\/\/llamatrace.com\/login\" target=\"_blank\" rel=\"nofollow noopener\">https:\/\/llamatrace.com\/login<\/a>. Then create an API key and put it in the\u00a0PHOENIX_API_KEY\u00a0variable below.<\/p>\n<p>Voyage AI utilizes API keys to monitor usage and manage permissions. To obtain your key, please sign in with your Voyage AI account and click the \u201cCreate new API key\u201d button in the\u00a0<a href=\"https:\/\/dash.voyageai.com\/\" target=\"_blank\" rel=\"nofollow noopener\">dashboard<\/a>. Add Payment details as well , but still Your first 200 million tokens are still free for Voyage series 3 models.<\/p>\n<p>Phoenix API key can be obtained by signing up for LlamaTrace <a href=\"https:\/\/llamatrace.com\/login#\/\" target=\"_blank\" rel=\"nofollow noopener\">here<\/a> , then navigate to the bottom left panel and click on \u2018Keys\u2019 where you should find your\u00a0 API key.<\/p>\n<pre class=\"wp-block-code\"><code>import os\nimport nest_asyncio\n\nnest_asyncio.apply()\n\n# Arize Phoenix\nPHOENIX_API_KEY = \"<phoenix_api_key>\"\nos.environ[\"OTEL_EXPORTER_OTLP_HEADERS\"] = f\"api_key={PHOENIX_API_KEY}\"\nimport llama_index.core\nllama_index.core.set_global_handler(\n    \"arize_phoenix\",\n    endpoint=\"https:\/\/llamatrace.com\/v1\/traces\"\n)<\/phoenix_api_key><\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-load-and-parse-the-pdf-slides\">Load and Parse the PDF Slides<\/h2>\n<p>In our example, we\u2019ll parse the ICONIQ 2024 State of AI Report. This PDF is publicly available at the URL below. If you prefer, you can replace it with any PDF you have.<\/p>\n<pre class=\"wp-block-code\"><code>!mkdir data\n!mkdir data_images_iconiq\n!wget \"https:\/\/cdn.prod.website-files.com\/65e1d7fb19a3e64b5c36fb38\/66eb856e019e59758ef73759_ICONIQ%20Analytics%20%2B%20Insights%20-%20State%20of%20AI%20Sep24.pdf\" -O data\/iconiq_report.pdf<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-model-setup\">Model Setup<\/h3>\n<p>Let\u2019s set up the core components required to build and implement our Multimodal RAG pipeline effectively.<\/p>\n<pre class=\"wp-block-code\"><code>import os\nfrom llama_index.llms.anthropic import Anthropic\nfrom llama_index.embeddings.voyageai import VoyageEmbedding\nfrom llama_index.core import Settings\n\n# Replace with your actual keys\nos.environ[\"ANTHROPIC_API_KEY\"] = \"sk-...\"\nos.environ[\"VOYAGE_API_KEY\"] = \"...\"\n\nllm = Anthropic(model=\"claude-3-5-sonnet-20240620\")\nembed_model = VoyageEmbedding(model_name=\"voyage-3\")\n\nSettings.llm = llm\nSettings.embed_model = embed_model<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-parse-text-and-images-with-llamaparse\">Parse Text and Images with LlamaParse<\/h3>\n<p>In this example, use LlamaParse to parse both the text and images from the document.<\/p>\n<p>We parse out the text with LlamaParse premium.<\/p>\n<p>NOTE: The report has 40 pages, and at ~5c per page, this will cost you $2. Just a heads up!<\/p>\n<p>For obtaining the LlamaCloud API key, click on the \u2018Get started\u2019 here\u00a0https:\/\/www.llamaindex.ai\/contact , and login. Once redirected to the <a href=\"https:\/\/cloud.llamaindex.ai\/\" target=\"_blank\" rel=\"nofollow noopener\">LlamaCloud dashboard<\/a>, generate a new API key by navigating to the API pane on the left.<\/p>\n<pre class=\"wp-block-code\"><code>from llama_parse import LlamaParse\n\nparser = LlamaParse(\n    result_type=\"markdown\",\n    premium_mode=True,\n    # invalidate_cache=True  # Uncomment if you want to force a fresh parse\n    api_key = 'LlamaCloud-API-Key'\n)\nprint(\"Parsing text...\")\nmd_json_objs = parser.get_json_result(\"data\/iconiq_report.pdf\")\nmd_json_list = md_json_objs[0][\"pages\"]\n\nimage_dicts = parser.get_images(md_json_objs, download_path=\"data_images_iconiq\")<\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-build-multimodal-nodes\">Build Multimodal Nodes<\/h2>\n<p>Multimodal nodes are the building blocks that allow us to process and integrate diverse data types like text and images. Here, we\u2019ll construct nodes to parse, embed, and index chunks from a PDF slide deck, setting the foundation for a robust retrieval system.<\/p>\n<p>Each PDF page corresponds to one \u201cnode\u201d containing:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Text <\/b>(parsed into Markdown)<\/li>\n<li><b>Image <\/b>(screenshot of that page)<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-split-pages-into-text-nodes\">Split Pages into Text Nodes<\/h3>\n<p>In this step, we\u2019ll split the PDF pages into smaller, manageable text nodes. This ensures efficient embedding and retrieval by breaking down the content into meaningful chunks for precise contextual analysis.<\/p>\n<pre class=\"wp-block-code\"><code>from pathlib import Path\nfrom llama_index.core.schema import TextNode\nfrom typing import Optional\nimport re\n\ndef get_page_number(file_name):\n    match = re.search(r\"-page_(\\d+)\\.jpg$\", str(file_name))\n    if match:\n        return int(match.group(1))\n    return 0\n\ndef _get_sorted_image_files(image_dir):\n    raw_files = [\n        f for f in list(Path(image_dir).iterdir()) if f.is_file() and \"-page\" in str(f)\n    ]\n    return sorted(raw_files, key=get_page_number)\n\ndef get_text_nodes(image_dir, json_dicts):\n    nodes = []\n    image_files = _get_sorted_image_files(image_dir)\n    md_texts = [d[\"md\"] for d in json_dicts]\n\n    for idx, md_text in enumerate(md_texts):\n        chunk_metadata = {\n            \"page_num\": idx + 1,\n            \"image_path\": str(image_files[idx]),\n            \"parsed_text_markdown\": md_text,\n        }\n        node = TextNode(text=\"\", metadata=chunk_metadata)\n        nodes.append(node)\n\n    return nodes\n\ntext_nodes = get_text_nodes(\"data_images_iconiq\", md_json_list)<\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-add-contextual-summaries\">Add Contextual Summaries<\/h2>\n<p>Contextual retrieval attaches a short, high-level summary to each chunk, describing where it fits into the overall document. We\u2019ll use the LLM to generate these short summaries and store them in each node\u2019s metadata[\u201ccontext\u201d].<\/p>\n<pre class=\"wp-block-code\"><code>from copy import deepcopy\nfrom llama_index.core.llms import ChatMessage\nfrom llama_index.core.prompts import ChatPromptTemplate\nimport time\n\n\nwhole_doc_text = \"\"\"\\\nHere is the entire document.\n<document>\n{WHOLE_DOCUMENT}\n<\/document>\"\"\"\n\nchunk_text = \"\"\"\\\nHere is the chunk we want to situate within the whole document\n<chunk>\n{CHUNK_CONTENT}\n<\/chunk>\nPlease give a short succinct context to situate this chunk within the overall document for \\\nthe purposes of improving search retrieval of the chunk. Answer only with the succinct context and nothing else.\"\"\"\n\n\ndef create_contextual_nodes(nodes, llm):\n    \"\"\"Function to create contextual nodes for a list of nodes\"\"\"\n    nodes_modified = []\n\n    # get overall doc_text string\n    doc_text = \"\\n\".join([n.get_content(metadata_mode=\"all\") for n in nodes])\n\n    for idx, node in enumerate(nodes):\n        start_time = time.time()\n        new_node = deepcopy(node)\n\n        # Combine whole_doc_text and chunk_text into a single string\n        user_content = (\n            f\"{whole_doc_text.format(WHOLE_DOCUMENT=doc_text)}\\n\\n\"\n            f\"{chunk_text.format(CHUNK_CONTENT=node.get_content(metadata_mode=\"all\"))}\"\n        )\n\n        messages = [\n            ChatMessage(role=\"system\", content=\"You are a helpful AI Assistant.\"),\n            ChatMessage(role=\"user\", content=user_content),\n        ]\n\n        # Send messages to the LLM and get a response\n        new_response = llm.chat(messages)\n        new_node.metadata[\"context\"] = str(new_response)\n\n        nodes_modified.append(new_node)\n        print(f\"Completed node {idx}, {time.time() - start_time}\")\n\n    return nodes_modified<\/code><\/pre>\n<p><b\/>Tip: We\u2019re passing an extra_headers parameter with a hypothetical prompt-caching date. This is just to illustrate how you might pass custom headers for Anthropic caching. Actual usage can vary.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-build-and-persist-the-index\">Build and Persist the Index<\/h2>\n<p>We\u2019ll now embed these summarized chunks and store them in a vector store for retrieval. LlamaIndex can persist indices locally or integrate with 40+ external vector databases.<\/p>\n<pre class=\"wp-block-code\"><code>import os\nfrom llama_index.core import (\n    StorageContext,\n    VectorStoreIndex,\n    load_index_from_storage,\n)\n\nif not os.path.exists(\"storage_nodes_iconiq\"):\n    index = VectorStoreIndex(new_text_nodes, embed_model=embed_model)\n    index.set_index_id(\"vector_index\")\n    index.storage_context.persist(\".\/storage_nodes_iconiq\")\nelse:\n    storage_context = StorageContext.from_defaults(persist_dir=\"storage_nodes_iconiq\")\n    index = load_index_from_storage(storage_context, index_id=\"vector_index\")\n\nretriever = index.as_retriever()<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-baseline-index-without-summaries\">Baseline Index (Without Summaries)<\/h3>\n<p>We\u2019ll also build a \u201cbaseline\u201d index on the original text nodes (without the contextual summaries) to compare the difference in retrieval quality.<\/p>\n<pre class=\"wp-block-code\"><code>if not os.path.exists(\"storage_nodes_iconiq_base\"):\n    base_index = VectorStoreIndex(text_nodes, embed_model=embed_model)\n    base_index.set_index_id(\"vector_index\")\n    base_index.storage_context.persist(\".\/storage_nodes_iconiq_base\")\nelse:\n    storage_context = StorageContext.from_defaults(\n        persist_dir=\"storage_nodes_iconiq_base\"\n    )\n    base_index = load_index_from_storage(storage_context, index_id=\"vector_index\")<\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-build-a-multimodal-query-engine\">Build a Multimodal Query Engine<\/h2>\n<p>We want a RAG pipeline that:<\/p>\n<ul class=\"wp-block-list\">\n<li>Retrieves relevant chunks of text.<\/li>\n<li>Also loads the page images.<\/li>\n<li>Sends both text chunks and images to a multimodal LLM (here we illustrate using an OpenAI-like GPT-4 multimodal endpoint, labeled gpt-4o).<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>import base64\nimport openai\nimport os\nfrom typing import Optional, List\n\nfrom llama_index.core.query_engine import CustomQueryEngine\nfrom llama_index.core.base.response.schema import Response\nfrom llama_index.core.retrievers import BaseRetriever\nfrom llama_index.core.prompts import PromptTemplate\nfrom llama_index.core.schema import NodeWithScore, MetadataMode\n\nQA_PROMPT_TMPL = \"\"\"\\\nBelow we give parsed text from slides, as well as images.\n\n---------------------\n{context_str}\n---------------------\n\nGiven the context information and no prior knowledge, please answer the query:\n\nQuery: {query_str}\nAnswer:\n\"\"\"\n\nQA_PROMPT = PromptTemplate(QA_PROMPT_TMPL)\n\ndef encode_image(image_path: str) -&gt; str:\n    \"\"\"If you want to inline a local image in base64.\"\"\"\n    with open(image_path, \"rb\") as f:\n        return base64.b64encode(f.read()).decode(\"utf-8\")\n\n\nclass MultimodalQueryEngine(CustomQueryEngine):\n    \"\"\"\n    Custom multimodal Query Engine that retrieves text nodes,\n    then sends them + image(s) to the new Vision-capable API as documented.\n    \"\"\"\n\n    def __init__(\n        self,\n        retriever: BaseRetriever,\n        model_name: str = \"gpt-4o\",\n        qa_prompt: Optional[PromptTemplate] = None,\n    ) -&gt; None:\n        super().__init__(qa_prompt=qa_prompt or QA_PROMPT)\n        self.retriever = retriever\n        self.model_name = model_name\n\n    def custom_query(self, query_str: str) -&gt; Response:\n        # 1) Retrieve text nodes\n        node_with_scores: List[NodeWithScore] = self.retriever.retrieve(query_str)\n\n        # 2) Build context\n        context_str = \"\\n\\n\".join(\n[nws.node.get_content(metadata_mode=MetadataMode.LLM) for nws in node_with_scores]\n\n        )\n\n        # 3) Format the final prompt\n        formatted_prompt_text = self._qa_prompt.format(\n            context_str=context_str,\n            query_str=query_str,\n        )\n\n        # 4) Build the user message with text + images\n        user_message_content = [\n            {\n                \"type\": \"text\",\n                \"text\": formatted_prompt_text,\n            }\n        ]\n\n        for nws in node_with_scores:\n            image_path = nws.node.metadata.get(\"image_path\", \"\")\n            if image_path:\n                base64_data = encode_image(image_path)\n                image_url = f\"data:image\/jpeg;base64,{base64_data}\"\n                user_message_content.append(\n                    {\n                        \"type\": \"image_url\",\n                        \"image_url\": {\n                            \"url\": image_url,\n                            \"detail\": \"auto\"\n                        },\n                    }\n                )\n\n        messages = [\n            {\n                \"role\": \"user\",\n                \"content\": user_message_content,\n            }\n        ]\n\n        # 5) Call your Vision model\n        response = openai.ChatCompletion.create(\n            model=self.model_name,\n            messages=messages,\n            max_tokens=500,\n        )\n\n        # 6) Return a Response object\n        return Response(\n            response=response.choices[0].message.content,\n            source_nodes=node_with_scores,\n            metadata={},\n        )\n        \n        # 2) Create a query engine\nquery_engine = MultimodalQueryEngine(\n    retriever=index.as_retriever(similarity_top_k=3),\n    model_name=\"gpt-4o\",   # or \"gpt-4o-mini\", \"gpt-4-turbo\", etc.\n)\n\nbase_query_engine = MultimodalQueryEngine(\n    retriever=base_index.as_retriever(similarity_top_k=3),\n    model_name=\"gpt-4o\",\n)<\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-trying-out-queries\">Trying Out Queries<\/h2>\n<p>Let\u2019s query our new pipeline about AI usage by department.<\/p>\n<pre class=\"wp-block-code\"><code>response = query_engine.query(\n    \"Which departments use GenAI the most and how are they using it?\"\n)\nprint(str(response))<\/code><\/pre>\n<p>A typical response might look like this:<\/p>\n<pre class=\"wp-block-preformatted\">Based on the parsed markdown text provided, the departments\/teams that use <br\/>generative AI the most are:<p>1. **AI, Machine Learning, and Data Science** with a score of 4.5.<br\/>2. **IT** with a score of 4.0.<br\/>3. **Engineering \/ R&amp;D** with a score of 3.9.<\/p><p>These scores are derived from a survey where respondents rated the level of<br\/>generative AI usage on a scale of 1-5.<\/p><p>In terms of how these departments are using generative AI:<\/p><p>- **AI, Machine Learning, and Data Science**: While specific use cases for this<br\/>department are not detailed in the provided text, it can be inferred that they are<br\/>likely using generative AI for advanced data analysis, model development, and<br\/>enhancing AI capabilities within the organization.<\/p><p>- **IT**: The IT department is using generative AI for several impactful use cases,<br\/>including:<br\/>- Ticket management<br\/>- Chatbots<br\/>- Customer support and troubleshooting<br\/>- Knowledge management<br\/>- Case summarization<\/p><p>The information about the departments and their use cases comes from the parsed<br\/>markdown text. There are no discrepancies between the parsed markdown and the<br\/>context provided, as the markdown text clearly outlines both the departments with<br\/>the highest usage scores and the specific use cases for the IT department.<\/p><\/pre>\n<p>Comparatively, if we run the same query on the baseline index:<\/p>\n<pre class=\"wp-block-code\"><code>base_response = base_query_engine.query(\n    \"Which departments use GenAI the most and how are they using it?\"\n)\nprint(str(base_response))<\/code><\/pre>\n<p>You\u2019ll see the baseline might have fewer details or slightly different retrieval results. Contextual retrieval gives more precise context around the IT usage specifically. The response would look like:<\/p>\n<pre class=\"wp-block-preformatted\">Based on the parsed markdown text provided, the departments that use Generative AI<br\/>(GenAI) the most are:<p>1. **AI, Machine Learning, and Data Science** - This department has the highest <br\/>weighted average score of 4.5 for GenAI usage, indicating significant adoption. The<br\/>specific use cases are not detailed in the parsed text, but given the nature of the<br\/>department, it is likely involved in developing and refining AI models and<br\/>algorithms.<\/p><p>2. **IT** - With a score of 4.0, the IT department is also a leading user of GenAI.<br\/>The use cases for IT include internal productivity enhancements and IT operations,<br\/>as indicated by the 61% adoption rate for internal productivity and 42% ROI mention<br\/>in IT use cases.<\/p><p>3. **Engineering \/ R&amp;D** - This department has a score of 3.9. While specific use<br\/>cases are not detailed in the parsed text, it is reasonable to infer that GenAI is<br\/>used for product development and research purposes, as suggested by the 69% <br\/>adoption rate for core product performance enhancements and 50% for natural language<br\/>interfaces.<\/p><p>The information is derived from the parsed markdown text, which provides a detailed<br\/>breakdown of GenAI usage by department and specific use cases. There are no <br\/>discrepancies between the parsed markdown and the raw text, as the markdown appears<br\/>to be a structured representation of the same data. The image was not provided, so<br\/>it was not used in forming the answer.<\/p><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-observing-the-benefits-of-contextual-retrieval\">Observing the Benefits of Contextual Retrieval<\/h2>\n<p>Here\u2019s another example query,\u00a0In this next question, the same sources are retrieved with and without contextual retrieval, and the answer is correct for both approaches. This is thanks for LlamaParse Premium\u2019s ability to comprehend graphs.<\/p>\n<pre class=\"wp-block-code\"><code>query = \"what are relevant insights from the 'deep dive on infrastructure' section in terms of model preferences, cost, deployment environments?\"\n\nresponse = query_engine.query(query)\nprint(str(response))<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-output\">Output<\/h3>\n<pre class=\"wp-block-preformatted\">The \"Deep Dive on Infrastructure\" section from the ICONIQ Growth report provides<br\/>insights into the infrastructure aspects necessary for deploying AI solutions. <br\/>However, the parsed markdown text does not explicitly mention model preferences or<br\/>costs in this section. Instead, it focuses on infrastructure tooling and deployment<br\/>environments.<p>From the parsed markdown text, we can gather the following insights related to <br\/>deployment environments:<\/p><p>1. **Deployment Environments**: Enterprises are primarily hosting generative AI<br\/>workloads on the cloud or using a hybrid approach. The preferred deployment methods<br\/>are:<br\/>- Cloud: 56%<br\/>- Hybrid: 42%<br\/>- On-prem: 2%<\/p><p>2. **Cloud Service Providers**: The most utilized cloud service providers for<br\/>hosting AI workloads are:<br\/>- Amazon Web Services (AWS): 68%<br\/>- Microsoft Azure: 61%<br\/>- Google Cloud (GCP): 40%<\/p><p>These insights are derived from the parsed markdown text, specifically from the <br\/>sections discussing \"Cloud Deployment Method\" and \"Infrastructure Tooling.\" There is<br\/>no mention of model preferences or cost considerations in the provided text. If<br\/>there were any discrepancies or additional details in the image or raw text, they<br\/>are not available here, so the answer is based solely on the parsed markdown text<br\/>provided.<\/p><\/pre>\n<p>Now, lets try with the baseline approach:<\/p>\n<pre class=\"wp-block-code\"><code>base_response = base_query_engine.query(query)\nprint(str(base_response))<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-output-0\">Output<\/h3>\n<pre class=\"wp-block-code\"><code>The parsed text from the slides does not provide specific insights regarding model preferences, cost, or deployment environments in the 'deep dive on infrastructure' section. The slide titled \"Deep Dive on Infrastructure\" (page 24) only contains the title, the ICONIQ Growth branding, and confidentiality and copyright notices. There is no detailed information or data presented in the parsed text for this section.<\/code><\/pre>\n<p>Therefore, based on the parsed markdown text provided, there are no relevant insights available from the \u2018deep dive on infrastructure\u2019 section regarding model preferences, cost, or deployment environments. If there were any images associated with this section, they were not provided, and thus no additional insights could be derived from them.<\/p>\n<p>This conclusion is drawn from the parsed markdown text, which lacks any specific information on model preferences, cost, or deployment environments in that section. The image confirms this, as it only shows the title and a graphic without additional details.<\/p>\n<p>If you need insights on these topics, you might want to refer to other sections or slides that specifically address model preferences, costs, or deployment environments.<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Contextual Retrieval<\/strong> might fetch the pages that discuss cloud deployment methods, infrastructure tooling, and cost references, leading to a more thorough response.<\/li>\n<li>The <strong>baseline approach<\/strong> might (in some cases) fail to retrieve the correct chunk or provide less detail.<\/li>\n<\/ul>\n<p>Comparing both answers helps demonstrate that those short \u201ccontextual summaries\u201d in your metadata often lead to more relevant retrieval.<\/p>\n<p>A big thanks to Jerry Liu from LlamaIndex for creating this amazing pipeline.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>In this tutorial, we explored the process of parsing a PDF slide deck using LlamaParse to extract both text and images, enriching each text chunk with contextual summaries to enhance retrieval accuracy. We demonstrated how to build a Multimodal RAG pipeline with LlamaIndex, integrating both textual and visual data into a powerful model like GPT-4, showcasing the potential of multimodal LLMs. Finally, we compared results from a baseline index to a contextual index, highlighting the significant improvements in retrieval precision and relevance achieved through the contextual approach. This comprehensive guide equips you with the tools and techniques to build effective multimodal AI solutions.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h3>\n<ul class=\"wp-block-list\">\n<li>Contextual retrieval improves chunk matching for queries that might not have a direct keyword overlap.<\/li>\n<li>Multimodal RAG can incorporate not just text but also images, charts, or diagrams from slides.<\/li>\n<li>Prompt caching is essential when chunk sizes are large and you\u2019re generating a context summary for each chunk\u2014this can reduce cost significantly.<\/li>\n<li>If you have web-based content (like store listings, large sets of HTML pages), you can use ScrapeGraphAI to fetch that data, then feed it into the same pipeline.<\/li>\n<\/ul>\n<p>With these steps, you can adapt the approach to <b>any <\/b>PDF or external data source\u2014whether it\u2019s a huge enterprise knowledge base, marketing materials, or your company\u2019s internal documentation.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1736147367774\"><strong class=\"schema-faq-question\">Q1. What is \u201cContextual Retrieval\u201d and why do I need it?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. Contextual Retrieval is an approach where each chunk of text in your dataset has a concise summary that situates it within the broader document. This helps your retriever better match relevant chunks\u2014especially for queries that rely on thematic or conceptual overlaps rather than exact keyword matches.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1736147387654\"><strong class=\"schema-faq-question\">Q2. How does Multimodal RAG differ from standard RAG?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. In a Multimodal RAG pipeline, you not only retrieve and feed text chunks into the LLM but also related images, audio, or other modalities. This is especially useful when your data sources are slide decks, PDFs with charts, or any materials that mix text with images. It allows the model to reference both textual and visual content for a more comprehensive answer.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1736147410916\"><strong class=\"schema-faq-question\">Q3. Why do I need LlamaParse to parse PDF slides?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. LlamaParse is a parsing utility that can extract both text and images from a PDF. Traditional PDF extractors often only get the text or struggle with embedded charts and diagrams. With LlamaParse, you can create \u201cnodes\u201d that include a reference to each PDF page\u2019s image file\u2014enabling genuine multimodal retrieval.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1736147454652\"><strong class=\"schema-faq-question\">Q4. Is it mandatory to create a baseline index without contextual summaries?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. No, it isn\u2019t mandatory, but it\u2019s a great way to benchmark the difference. Having a baseline index helps you see how retrieval results change when you add contextual summaries.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><em><strong>This article was published as a part of the\u00a0<\/strong><\/em><a href=\"https:\/\/www.analyticsvidhya.com\/datahack\/blogathon\" target=\"_blank\" rel=\"noreferrer noopener\"><em><strong>Data Science Blogathon.<\/strong><\/em><\/a><\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/adarsh2039075\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_tHXFGNS.webp\" width=\"48\" height=\"48\" alt=\"Adarsh Balan\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>    Hi! I&#8217;m Adarsh, a Business Analytics graduate from ISB, currently deep into research and exploring new frontiers. I&#8217;m super passionate about data science, AI, and all the innovative ways they can transform industries. Whether it&#8217;s building models, working on data pipelines, or diving into machine learning, I love experimenting with the latest tech. AI isn&#8217;t just my interest, it&#8217;s where I see the future heading, and I&#8217;m always excited to be a part of that journey!    <\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Imagine a world where finding information in a document is as easy as asking a question\u2014and getting a response that combines both text and images seamlessly. In this guide, we dive into building a Multimodal Retrieval-Augmented Generation pipeline that can do just that. You\u2019ll learn how to parse text and images from a PDF slide [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":74961,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5815,31610,38577,20383,32726,38576,964],"dealstore":[],"offerexpiration":[],"class_list":["post-74960","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-blogathon","tag-contextual","tag-decks","tag-multimodal","tag-rag","tag-retrieval","tag-slide"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Contextual Retrieval for Multimodal RAG on Slide Decks - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=74960\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Contextual Retrieval for Multimodal RAG on Slide Decks - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Imagine a world where finding information in a document is as easy as asking a question\u2014and getting a response that combines both text and images seamlessly. In this guide, we dive into building a Multimodal Retrieval-Augmented Generation pipeline that can do just that. You\u2019ll learn how to parse text and images from a PDF slide [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=74960\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-08T03:24:04+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"19 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=74960#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=74960\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Contextual Retrieval for Multimodal RAG on Slide Decks\",\"datePublished\":\"2025-02-08T03:24:04+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=74960\"},\"wordCount\":1950,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=74960#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp\",\"keywords\":[\"Blogathon\",\"contextual\",\"Decks\",\"Multimodal\",\"RAG\",\"Retrieval\",\"Slide\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=74960#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=74960\",\"url\":\"https:\/\/fivemor.com\/?p=74960\",\"name\":\"Contextual Retrieval for Multimodal RAG on Slide Decks - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=74960#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=74960#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp\",\"datePublished\":\"2025-02-08T03:24:04+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=74960#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=74960\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=74960#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=74960#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Contextual Retrieval for Multimodal RAG on Slide Decks\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Contextual Retrieval for Multimodal RAG on Slide Decks - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=74960","og_locale":"en_US","og_type":"article","og_title":"Contextual Retrieval for Multimodal RAG on Slide Decks - Som2ny Network","og_description":"Imagine a world where finding information in a document is as easy as asking a question\u2014and getting a response that combines both text and images seamlessly. In this guide, we dive into building a Multimodal Retrieval-Augmented Generation pipeline that can do just that. You\u2019ll learn how to parse text and images from a PDF slide [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=74960","og_site_name":"Som2ny Network","article_published_time":"2025-02-08T03:24:04+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"19 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=74960#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=74960"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Contextual Retrieval for Multimodal RAG on Slide Decks","datePublished":"2025-02-08T03:24:04+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=74960"},"wordCount":1950,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=74960#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp","keywords":["Blogathon","contextual","Decks","Multimodal","RAG","Retrieval","Slide"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=74960#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=74960","url":"https:\/\/fivemor.com\/?p=74960","name":"Contextual Retrieval for Multimodal RAG on Slide Decks - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=74960#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=74960#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp","datePublished":"2025-02-08T03:24:04+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=74960#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=74960"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=74960#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Contextual-Retrieval-for-Multimodal-RAG-on-Slide-Decks-with-LlamaIndex.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=74960#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Contextual Retrieval for Multimodal RAG on Slide Decks"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/74960","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=74960"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/74960\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/74961"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=74960"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=74960"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=74960"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=74960"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=74960"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}