{"id":143733,"date":"2025-03-19T13:11:20","date_gmt":"2025-03-19T13:11:20","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/how-to-build-multimodal-rag-using-docling\/"},"modified":"2025-03-19T13:11:20","modified_gmt":"2025-03-19T13:11:20","slug":"how-to-build-multimodal-rag-using-docling","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=143733","title":{"rendered":"How to Build Multimodal RAG Using Docling?"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Multimodal Retrieval-Augmented Generation (RAG) is a transformative innovation in AI, enabling systems to process and integrate diverse data types such as text, images, audio, and video. This capability is crucial in addressing the challenge of unstructured enterprise data, which predominantly consists of multimodal formats. By leveraging multimodal inputs, RAG enhances contextual understanding, improves accuracy, and expands AI\u2019s applicability across industries like healthcare, customer support, and education.\u00a0Docling is an open-source toolkit developed by IBM to streamline document processing for generative AI applications. We will build Multimodal RAG Capabilities Using Docling.<\/p>\n<p>It converts diverse formats like PDFs, DOCX, and images into structured outputs such as JSON and Markdown, enabling seamless integration with AI frameworks like LangChain and LlamaIndex. By facilitating the extraction of unstructured data and supporting advanced layout analysis, Docling empowers <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/02\/multimodal-rag-with-deepseek-janus-pro\/\" target=\"_blank\" rel=\"noreferrer noopener\">multimodal Retrieval-Augmented Generation (RAG)<\/a> by making complex enterprise data machine-readable and accessible for AI-driven insights<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-learning-objectives\">Learning Objectives<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Exploring Docling<\/strong> \u2013 Understanding how it extracts multimodal information from unstructured files. <\/li>\n<li><strong>Docling Pipeline &amp; AI Models<\/strong> \u2013 Examining its architecture and key AI components. <\/li>\n<li><strong>Unique Features<\/strong> \u2013 Highlighting what makes Docling stand out. <\/li>\n<li><strong>Building a Multimodal RAG System<\/strong> \u2013 Implementing a system using Docling for data extraction and retrieval. <\/li>\n<li><strong>End-to-End Process<\/strong> \u2013 Extracting data from a PDF, generating image descriptions, and querying with a vector DB &amp; Phi 4.<\/li>\n<\/ul>\n<p><em><strong>This article was published as a part of the\u00a0<\/strong><\/em><a href=\"https:\/\/www.analyticsvidhya.com\/datahack\/blogathon\" target=\"_blank\" rel=\"noreferrer noopener\"><em><strong>Data Science Blogathon.<\/strong><\/em><\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-docling-for-unstructured-data\">Docling For Unstructured Data<\/h2>\n<p>Docling is an open-source document processing toolkit developed by <a href=\"https:\/\/research.ibm.com\/blog\/docling-generative-AI\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">IBM<\/a>, designed to convert unstructured files like PDFs, DOCX, and images into structured formats such as JSON and Markdown. Powered by advanced AI models like DocLayNet for layout analysis and TableFormer for table recognition, it enables accurate extraction of text, tables, and images while preserving document structure. With seamless integration into generative AI frameworks like LangChain and LlamaIndex, Docling supports applications such as Retrieval-Augmented Generation (RAG) and question-answering systems. Its lightweight architecture allows efficient performance on standard hardware, making it a cost-effective alternative to SaaS-based solutions for enterprises seeking control over data privacy.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-docling-pipeline\">Docling Pipeline<\/h2>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"735\" height=\"279\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_gPlLqbz.webp\" alt=\"docling pipeline\" class=\"wp-image-226851\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_gPlLqbz.webp 735w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_gPlLqbz-300x114.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_gPlLqbz-150x57.webp 150w\" sizes=\"(max-width: 735px) 100vw, 735px\"\/><figcaption class=\"wp-element-caption\">Docling\u2019s default processing pipeline. The inner part of the model pipeline is easily customizable and extensible. <a href=\"https:\/\/arxiv.org\/html\/2408.09869v1#bib.bib13\" target=\"_blank\" rel=\"nofollow noopener\">Image Source<\/a>\u00a0<\/figcaption><\/figure>\n<\/div>\n<p>Docling implements a linear pipeline of operations, which execute sequentially on each given document (as shown in the above Figure). Each document is first parsed by a PDF backend, which retrieves the programmatic text tokens, consisting of string content and its coordinates on the page, and also renders a bitmap image of each page to support downstream operations. Then, the standard model pipeline applies a sequence of AI models independently on every page in the document to extract features and content, such as layout and table structures. Finally, the results from all pages are aggregated and passed through a post-processing stage, which augments metadata, detects the document language, infers reading order and eventually assembles a typed document object which can be serialized to JSON or Markdown.\u00a0<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-key-ai-models-behind-docling\">Key AI Models Behind Docling<\/h2>\n<p>Traditionally, developers have depended on <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2020\/05\/build-your-own-ocr-google-tesseract-opencv\/\" target=\"_blank\" rel=\"noreferrer noopener\">optical character recognition (OCR)<\/a> for converting documents into digital formats. However, this technology can be slow and prone to errors due to the heavy computational power required. Docling avoids OCR whenever possible, instead using computer vision models that are specifically trained to identify and categorize the visual components of a page.<\/p>\n<p>Docling is based on two models developed by IBM researchers.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-layout-analysis-model\">Layout Analysis Model<\/h3>\n<p>The layout analysis model functions as an object detector, predicting the bounding boxes and categories of various elements within an image of a given page. Its design is based on <a href=\"https:\/\/huggingface.co\/docs\/transformers\/en\/model_doc\/rt_detr\" target=\"_blank\" rel=\"nofollow noopener\">RT-DETR<\/a> and has been re-trained using DocLayNet, our well-known human-annotated dataset for document layout analysis, along with other proprietary datasets. <a href=\"https:\/\/developer.ibm.com\/data\/doclaynet\/\" target=\"_blank\" rel=\"nofollow noopener\">DocLayNet<\/a> is a human-annotated document layout segmentation dataset containing 80863 pages from a broad variety of document sources.<\/p>\n<p>This model utilizes object detection techniques to examine the layout of documents, ranging from machine manuals to annual reports. It then identifies and classifies elements such as blocks of text, images, tables, captions, and more. The Docling pipeline processes page images at a resolution of 72 dpi, enabling them to be handled by a single CPU.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-table-former-model\">Table Former Model<\/h3>\n<p>The TableFormer model, initially introduced in 2022 and subsequently enhanced with a custom token structure language, is a vision-transformer model designed for recovering the structure of tables. It can predict the logical organization of rows and columns in a table based on an input image, identifying which cells belong to column headers, row headers, or the main body of the table. Unlike previous methods, TableFormer effectively handles various table complexities, including partial or absent borders, empty cells, missing rows or columns, cell spans, hierarchical structures in both column and row headings, as well as inconsistencies in indentation or alignment.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-some-key-features-of-docling\">Some Key Features of Docling<\/h3>\n<p>Here are the features:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Versatile Format Support:<\/b> Docling can parse a wide range of document formats, including PDFs, DOCX, PPTX, HTML, images, and more. It exports content into structured formats like JSON and Markdown for seamless integration into AI workflows<\/li>\n<li><b>Advanced PDF Processing:<\/b> It includes sophisticated capabilities such as layout analysis, reading order detection, table structure recognition, and OCR for scanned documents. This ensures the accurate extraction of complex document elements like tables and figures. Docling extracts tables using advanced AI-driven methods, primarily leveraging its custom TableFormer model.<\/li>\n<li><b>Unified Document Representation: <\/b>Docling uses a unified and expressive format to represent parsed documents, making it easier to process and analyze them in downstream applications<\/li>\n<li><b>AI-Ready Integration: <\/b>The toolkit integrates seamlessly with popular AI frameworks like LangChain and LlamaIndex, making it ideal for applications like Retrieval-Augmented Generation (RAG) and question-answering systems<\/li>\n<li><b>Local Execution:<\/b> It supports local execution, enabling secure processing of sensitive data in air-gapped environments<\/li>\n<li><b>Efficient Performance: <\/b>Designed to run on commodity hardware with minimal resource requirements, Docling avoids traditional OCR when possible, speeding up processing by up to 30 times while reducing errors.<\/li>\n<li><b>Modular Architecture:<\/b> Its modular design allows easy customization and extension with new features or models, catering to diverse use cases<\/li>\n<li><b>Open-Source Accessibility: <\/b>Unlike proprietary tools like Watson Document Understanding, Docling is open-source under the MIT license, allowing developers to freely use, customize, and integrate it into their workflows without vendor lock-in or additional costs<\/li>\n<\/ul>\n<p>Docling provides optional support for OCR, for example, to cover scanned PDFs or content in<br \/>bitmap images embedded on a page. Docling relies on EasyOCR, a popular third-party OCR library with support for many languages. These features make Docling a comprehensive solution for document parsing and preparation in generative AI workflows.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-building-a-multimodal-rag-system-using-docling\">Building a Multimodal RAG System using Docling<\/h2>\n<p>In this article, we will first extract all kinds of data \u2013 text, images, and tables from a PDF using Docling. For extracted images, we will use a vision language model to generate the description of the images and save these text descriptions of the images in our VectorDB along with the text data from the original text contents and text from extracted Tables in the PDF. Post this, we will build a RAG system using the vector DB for retrieval along with an LLM (Phi 4) through Ollama for querying from the PDF document.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-hands-on-python-implementation-on-google-colab-using-t4-gpu-free-tier\">Hands-On Python Implementation on Google Colab using T4 GPU (Free Tier)<\/h2>\n<p>You can find the Colab Notebook which has all the steps <a href=\"https:\/\/colab.research.google.com\/drive\/1QH9xnA9O-8x4QNgCCwFJ0LRmNg4QnyqL?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">here<\/a>.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-1-installing-libraries\">Step 1. Installing Libraries<\/h3>\n<p>We first start with installing the necessary libraries<\/p>\n<pre class=\"wp-block-code\"><code>!pip install docling\n\n#Following code added to avoid an error in installation - can be removed if not needed\nimport locale\ndef getpreferredencoding(do_setlocale = True):\n    return \"UTF-8\"\nlocale.getpreferredencoding = getpreferredencoding\n\n!pip install langchain-huggingface<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-2-loading-the-converter-object\">Step 2. Loading the Converter Object<\/h3>\n<p>This code prepares a document converter to process PDF files without OCR but with image generation. It then applies this conversion to a specified PDF file, storing the results in a dictionary.<\/p>\n<p>We use this <a href=\"https:\/\/investor.accenture.com\/~\/media\/Files\/A\/Accenture-IR-V3\/events-and-presentations\/accenture-investor-and-analyst-conference-cfo-slides.pdf\" target=\"_blank\" rel=\"nofollow noopener\">PDF<\/a> (we save it in the current working directory as \u2018accenture.pdf\u2019) which has a lot of charts to test the multimodal retrieval using Docling.<\/p>\n<pre class=\"wp-block-code\"><code>from docling.document_converter import DocumentConverter, PdfFormatOption\nfrom docling.datamodel.base_models import InputFormat\nfrom docling.datamodel.pipeline_options import PdfPipelineOptions\n\npdf_pipeline_options = PdfPipelineOptions(do_ocr=False,generate_picture_images=True,)\nformat_options = {InputFormat.PDF: PdfFormatOption(pipeline_options=pdf_pipeline_options)}\nconverter = DocumentConverter(format_options=format_options)\n\nsources = [ \"\/content\/accenture.pdf\",]\nconversions = {source: converter.convert(source=source).document for source in sources}<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-3-loading-the-model-for-embedding-text\">Step 3. Loading the Model For Embedding Text<\/h3>\n<pre class=\"wp-block-code\"><code>from langchain_huggingface.embeddings import HuggingFaceEmbeddings\nfrom transformers import *\n\nembeddings_model_path = \"ibm-granite\/granite-embedding-30m-english\"\nembeddings_model = HuggingFaceEmbeddings(model_name=embeddings_model_path,)\nembeddings_tokenizer = AutoTokenizer.from_pretrained(embeddings_model_path)<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-4-chunking-the-texts-in-the-document\">Step 4. Chunking the Texts in the Document<\/h3>\n<p>The code below is for the document processing pipeline. It takes converted documents from the previous step and breaks them down into smaller chunks, excluding tables (which is processed separately later). Each chunk is then wrapped into a\u00a0Document\u00a0object with specific metadata. The\u00a0code processes converted documents by splitting them into chunks, skipping tables, and creating new\u00a0Document\u00a0objects with metadata for each chunk.<\/p>\n<pre class=\"wp-block-code\"><code>from docling_core.transforms.chunker.hybrid_chunker import HybridChunker\nfrom docling_core.types.doc.document import TableItem\nfrom langchain_core.documents import Document\n\n\ndoc_id = 0\ntexts: list[Document] = []\n\nfor source, docling_document in conversions.items():\n    for chunk in HybridChunker(tokenizer=embeddings_tokenizer).chunk(docling_document):\n        items = chunk.meta.doc_items\n        if len(items) == 1 and isinstance(items[0], TableItem):\n            continue # we will process tables later\n        refs = \" \".join(map(lambda item: item.get_ref().cref, items))\n        text = chunk.text\n        document = Document(page_content=text,metadata={\"doc_id\": (doc_id:=doc_id+1),\"source\": source,\"ref\": refs,},)\n        texts.append(document)\n\n\n\nprint(f\"{len(texts)} text document chunks created\")<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-5-processing-the-tables-in-the-document\">Step 5. Processing the Tables in the Document<\/h3>\n<p>The code below is designed to process tables from converted documents. It extracts tables, converts them into Markdown format, and wraps each table into a\u00a0Document\u00a0object with specific metadata.<\/p>\n<pre class=\"wp-block-code\"><code>from docling_core.types.doc.labels import DocItemLabel\n\ndoc_id = len(texts)\ntables: list[Document] = []\n\nfor source, docling_document in conversions.items():\n    for table in docling_document.tables:\n        if table.label in [DocItemLabel.TABLE]:\n            ref = table.get_ref().cref\n            text = table.export_to_markdown()\n            document = Document(\n                page_content=text,\n                metadata={\n                    \"doc_id\": (doc_id:=doc_id+1),\n                    \"source\": source,\n                    \"ref\": ref\n                },\n            )\n            tables.append(document)\n\n\nprint(f\"{len(tables)} table documents created\")<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-6-defining-function-for-converting-images-from-pdf-to-base64-form\">Step 6. Defining Function For Converting Images From PDF to base64 form<\/h3>\n<pre class=\"wp-block-code\"><code>import base64\nimport io\nimport PIL.Image\nimport PIL.ImageOps\nfrom IPython.display import display\n\ndef encode_image(image: PIL.Image.Image, format: str = \"png\") -&gt; str:\n    image = PIL.ImageOps.exif_transpose(image) or image\n    image = image.convert(\"RGB\")\n    buffer = io.BytesIO()\n    image.save(buffer, format)\n    encoding = base64.b64encode(buffer.getvalue()).decode(\"utf-8\")\n    return encoding<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-7-pulling-model-from-ollama-for-analysing-images-from-the-pdf\">Step 7. Pulling Model From Ollama For Analysing Images from the PDF<\/h3>\n<p>We will use a vision language model from Ollama to analyse the extracted images from the PDF and generate a description for each of the images. To facilitate the use of Ollama models, we install the following libraries and start up the Ollama server before pulling the model as described below in the code.<\/p>\n<pre class=\"wp-block-code\"><code>!sudo apt update\n!sudo apt install -y pciutils\n!pip install langchain-ollama\n!curl -fsSL https:\/\/ollama.com\/install.sh | sh\n!pip install ollama==0.4.2\n!pip install langchain-community\n\n\n#Enabling threading to start ollama server in a non blocking manner\nimport threading\nimport subprocess\nimport time\n\ndef run_ollama_serve():\n  subprocess.Popen([\"ollama\", \"serve\"])\n\nthread = threading.Thread(target=run_ollama_serve)\nthread.start()\ntime.sleep(5)<\/code><\/pre>\n<p>The code below is designed to process images from converted documents. It extracts images, uses\u00a0a vision model (llama3.2-vision through Ollama) to generate descriptive text for each image, and wraps this text into\u00a0a\u00a0Document\u00a0object with\u00a0specific metadata. Here\u2019s a detailed explanation:<\/p>\n<p>Pulling the \u201cllama3.2-vision\u201d model from Ollama.<\/p>\n<pre class=\"wp-block-code\"><code>!ollama pull llama3.2-vision<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>def encode_image(image: PIL.Image.Image, format: str = \"png\") -&gt; str:\n\n    image = PIL.ImageOps.exif_transpose(image) or image\n    image = image.convert(\"RGB\")\n    buffer = io.BytesIO()\n    image.save(buffer, format)\n    encoding = base64.b64encode(buffer.getvalue()).decode(\"utf-8\")\n    return encoding<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>import ollama\n\npictures: list[Document] = []\ndoc_id = len(texts) + len(tables)\n\nfor source, docling_document in conversions.items():\n    for picture in docling_document.pictures:\n        ref = picture.get_ref().cref\n        image = picture.get_image(docling_document)\n        if image:\n            print(image)\n            response = ollama.chat(\n            model=\"llama3.2-vision\",\n            messages=[{\n              \"role\": \"user\",\n              \"content\": \"Describe this image?\",\n              \"images\": [encode_image(image)]\n            }],\n        )\n            text = response['message']['content'].strip()\n            document = Document(\n                page_content=text,\n                metadata={\n\n                    \"doc_id\": (doc_id:=doc_id+1),\n\n                    \"source\": source,\n\n                    \"ref\": ref,\n\n                },\n\n            )\n\n            pictures.append(document)\n\nprint(f\"{len(pictures)} image descriptions created\")<\/code><\/pre>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"687\" height=\"457\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-67.webp\" alt=\"Output\" class=\"wp-image-226889\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-67.webp 687w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-67-300x200.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-67-150x100.webp 150w\" sizes=\"auto, (max-width: 687px) 100vw, 687px\"\/><\/figure>\n<pre class=\"wp-block-code\"><code>import itertools\nfrom docling_core.types.doc.document import RefItem\n\n# Print all created documents\nfor document in itertools.chain(texts, tables):\n    print(f\"Document ID: {document.metadata['doc_id']}\")\n    print(f\"Source: {document.metadata['source']}\")\n    print(f\"Content:\\n{document.page_content}\")\n    print(\"=\" * 80) # Separator for clarity\n\nfor document in pictures:\n    print(f\"Document ID: {document.metadata['doc_id']}\")\n    source = document.metadata['source']\n    print(f\"Source: {source}\")\n    print(f\"Content:\\n{document.page_content}\")\n    docling_document = conversions[source]\n    ref = document.metadata['ref']\n    picture = RefItem(cref=ref).resolve(docling_document)\n    image = picture.get_image(docling_document)\n    print(\"Image:\")\n    display(image)\n    print(\"=\" * 80) # Separator for clarity<\/code><\/pre>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1893\" height=\"1069\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-66.webp\" alt=\"Output\" class=\"wp-image-226890\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-66.webp 1893w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-66-300x169.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-66-768x434.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-66-1536x867.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-66-150x85.webp 150w\" sizes=\"auto, (max-width: 1893px) 100vw, 1893px\"\/><\/figure>\n<h3 class=\"wp-block-heading\" id=\"h-step-9-storing-in-milvus-vector-db\">Step 9. Storing in Milvus Vector DB<\/h3>\n<p><a href=\"https:\/\/milvus.io\/\" target=\"_blank\" rel=\"nofollow noopener\">Milvus<\/a>\u00a0is a high-performance vector database built for scale. It powers AI applications by efficiently organizing and searching vast amounts of unstructured data, such as text, images, and multi-modal information. We install the langchain-milvus library first and then store the texts, tables and pictures in the vector DB. While defining the vector DB, we also pass the embedding model so that the vector DB converts all the text extracted, including the data from tables and image descriptions, into embeddings before storing them.<\/p>\n<pre class=\"wp-block-code\"><code>!pip install langchain_milvus\n\nimport tempfile\nfrom langchain_core.vectorstores import VectorStore\nfrom langchain_milvus import Milvus\n\n\ndb_file = tempfile.NamedTemporaryFile(prefix=\"vectorstore_\", suffix=\".db\", delete=False).name\nvector_db: VectorStore = Milvus(embedding_function=embeddings_model,connection_args={\"uri\": db_file},auto_id=True,enable_dynamic_field=True,index_params={\"index_type\": \"AUTOINDEX\"},)\n\n\n#add all the LangChain documents for the text, tables and image descriptions to the vector database\nimport itertools\ndocuments = list(itertools.chain(texts, tables, pictures))\nids = vector_db.add_documents(documents)\nprint(f\"{len(ids)} documents added to the vector database\")<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-10-querying-the-model-using-retrieval-augmented-generation-with-phi-4-model\">Step 10. Querying the model using Retrieval Augmented Generation with Phi 4 model<\/h3>\n<p>In the following code, we first pull the \u201cPhi 4\u201d model from Ollama and then use it as the LLM in this RAG system for generating a response post retrieval of the relevant context from the vector DB based on a query.<\/p>\n<pre class=\"wp-block-code\"><code>#Pulling the Ollama model for querying\n!ollama pull phi4\n\n#Querying\nfrom langchain_core.output_parsers import StrOutputParser\nfrom langchain.prompts import ChatPromptTemplate\nfrom langchain_community.chat_models import ChatOllama\nfrom langchain_core.runnables import RunnableLambda, RunnablePassthrough\nretriever = vector_db.as_retriever()\n\n\n# Prompt\ntemplate = \"\"\"Answer the question based only on the following context:\n{context}\n\nQuestion: {question}\n\"\"\"\nprompt = ChatPromptTemplate.from_template(template)\n\n# Local LLM\nollama_llm = \"phi4\"\nmodel_local = ChatOllama(model=ollama_llm)\n\n# Chain\nchain = (\n    {\"context\": retriever, \"question\": RunnablePassthrough()}\n    | prompt\n    | model_local\n    | StrOutputParser()\n)\n\nchain.invoke(\"How much worth in dollars is Strategy &amp; Conslution in Services?\")\n<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-output\"><b>Output<\/b><\/h4>\n<pre class=\"wp-block-preformatted\">According to the context provided, the 'Technology &amp; Strategy\/Consulting'<br\/>section of the company's operations generated a value of $15 billion.<\/pre>\n<p>As seen from the chart below from the document, the response of our multimodal RAG system is correct. With Docling, the information was correctly extracted from the chart and hence the retrieval system was able to provide us with an accurate response.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"736\" height=\"322\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_MAHzhCB.webp\" alt=\"The chart in the Original PDF\" class=\"wp-image-226852\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_MAHzhCB.webp 736w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_MAHzhCB-300x131.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image_MAHzhCB-150x66.webp 150w\" sizes=\"auto, (max-width: 736px) 100vw, 736px\"\/><figcaption class=\"wp-element-caption\">The chart in the Original PDF<\/figcaption><\/figure>\n<\/div>\n<h2 class=\"wp-block-heading\" id=\"h-analyzing-our-rag-system-with-more-queries\">Analyzing Our RAG System with More Queries<\/h2>\n<h4 class=\"wp-block-heading\" id=\"h-what-was-the-revenue-in-germany\">What was the revenue in Germany?<\/h4>\n<pre class=\"wp-block-preformatted\">The revenue in Germany, according to the provided context, is $3 billion.<br\/>This information is listed under the 'Country-Wise Revenue' section of the<br\/>document: \\n\\n. **Germany**: $3 billion\\n\\nIf you need any further details<br\/>or have additional questions, feel free to ask!<\/pre>\n<p>As seen from the chart below from the document, the response of our multimodal RAG system is correct. With Docling, the information was correctly extracted from the chart and hence the retrieval system was able to provide us with an accurate response.<\/p>\n<figure class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\"><img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/__sized__\/article_images\/image_PzkfwwJ-thumbnail_webp-600x300.webp\" alt=\" Chart in Original PDF\"\/><figcaption class=\"wp-element-caption\">Chart in Original PDF<\/figcaption><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-what-was-the-cloud-fy19-revenue\">What was the Cloud FY19 revenue?<\/h4>\n<pre class=\"wp-block-preformatted\">The Cloud FY19 revenue, as provided in the document context, was $11 billion.<br\/>This information is found in the first table under the section titled<br\/>'Cloud' where it states:\\n\\nFY19: $11B\\n\\nThis indicates that the revenue <br\/>from cloud services for fiscal year 2019 was $11 billion.<\/pre>\n<p>As seen from the Table below from the document, the response of our multimodal RAG system is correct. With Docling, the information was correctly extracted from the chart and hence the retrieval system was able to provide us with an accurate response.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter\"><img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/__sized__\/article_images\/image_A4d4cij-thumbnail_webp-600x300.webp\" alt=\" Chart in Original PDF\"\/><figcaption class=\"wp-element-caption\">Chart in Original PDF<\/figcaption><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-what-was-the-industry-x-3-yr-cagr\">What was the Industry X 3 Yr CAGR?<\/h4>\n<pre class=\"wp-block-preformatted\">Based on the provided context from the documents in Accenture\u2019s PDF:\\n\\n-In<br\/>Document with <code>doc_id<\/code> 15 and Document with <code>doc_id<\/code> 3, both mention Industry<br\/>X.\\n-The relevant information is found under a section about revenue growth<br\/>for Industry X:\\n\\n**Document 15** indicates: \"FY19 $10B Industry X FY19 $3B<br\/>FY22 $6.5B 3 Yr. CAGR 2 30%\"\\n\\n**Document 3** reiterates this with similar<br\/>wording: \"Cloud = FY19 $10B Industry X FY19. , Illustrative = . , Cloud =<br\/>$3B. , Illustrative = FY22 $6.5B. , Illustrative = 3 Yr. CAGR 2 30%\"\\n\\nFrom<br\/>these excerpts, the 3-year compound annual growth rate (CAGR) for Industry X<br\/>is **30%.\"**.\\n\\n<\/pre>\n<p>As seen from the previous Table from the document, the response of our multimodal RAG system is correct. With Docling, the information was correctly extracted from the chart and hence the retrieval system was able to provide us with an accurate response<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>In conclusion, Docling stands as a powerful tool for transforming unstructured data into machine-readable formats, making it an essential resource for applications like Multimodal Retrieval-Augmented Generation (RAG). By utilizing advanced AI models and offering seamless integration with popular AI frameworks, Docling enhances the ability to process and query complex documents efficiently. Its open-source nature, combined with versatile format support and modular architecture, makes it an ideal solution for enterprises seeking to leverage generative AI in real-world use cases.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Docling Toolkit<\/strong>: IBM\u2019s open-source tool for extracting structured data (JSON, Markdown) from PDFs, DOCX, and images, enabling seamless AI integration. <\/li>\n<li><strong>Advanced AI Models<\/strong>: Uses Layout Analysis and TableFormer for accurate document processing, reducing reliance on traditional OCR. <\/li>\n<li><strong>AI Framework Integration<\/strong>: Works with LangChain and LlamaIndex, ideal for RAG systems, offering cost-effective AI-driven insights. <\/li>\n<li><strong>Open-Source &amp; Customizable<\/strong>: MIT-licensed, modular, and adaptable for diverse use cases, free from vendor lock-in.<\/li>\n<\/ul>\n<p><strong>The media shown in this article is not owned by Analytics Vidhya and is used at the Author\u2019s discretion.<\/strong><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/krishnaveni140696\/\"\/><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1742299783179\"><strong class=\"schema-faq-question\">Q1. What is Multimodal Retrieval-Augmented Generation (RAG) and how does it work?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. RAG is an AI framework that integrates various data types, such as text, images, audio, and video, to improve contextual understanding and accuracy. By processing multimodal inputs, RAG enables AI systems to generate more accurate insights and extend their applicability across industries like healthcare, education, and customer support.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1742299798692\"><strong class=\"schema-faq-question\">Q2. What is Docling and how does it support AI-driven workflows?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Docling is an open-source document processing toolkit developed by IBM. It converts unstructured documents (e.g., PDFs, DOCX, images) into structured formats such as JSON and Markdown. This conversion enables seamless integration with generative AI frameworks like LangChain and LlamaIndex, facilitating applications like RAG and question-answering systems.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1742299811750\"><strong class=\"schema-faq-question\">Q3. How does Docling handle complex document elements like tables and images?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Docling utilizes advanced AI models like Layout Analysis for detecting document layout elements and TableFormer for recognizing table structures. These models help extract text, tables, and images while preserving the document\u2019s structure, improving accuracy and making complex data machine-readable for AI systems.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1742299827211\"><strong class=\"schema-faq-question\">Q4. Can Docling be used with other AI frameworks and models?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Yes, Docling is designed to integrate seamlessly with popular AI frameworks like LangChain and LlamaIndex. It can be used to power applications like Retrieval-Augmented Generation (RAG) by extracting data from unstructured documents and enabling AI systems to query and retrieve relevant information.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1742299846332\"><strong class=\"schema-faq-question\">Q5. Is Docling a cost-effective solution for enterprises handling sensitive data?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Docling is a cost-effective alternative to SaaS-based document processing tools. It allows local execution, making it ideal for enterprises that need to process sensitive data in air-gapped environments, ensuring data privacy while offering efficient performance on standard hardware. Additionally, Docling is open-source under the MIT license, allowing for easy customization without vendor lock-in.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/mimi6\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_ZkJo4gb.webp\" width=\"48\" height=\"48\" alt=\"Nibedita Dutta\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Nibedita completed her master\u2019s in Chemical Engineering from IIT Kharagpur in 2014 and is currently working as a Senior Data Scientist. In her current capacity, she works on building intelligent ML-based solutions to improve business processes.               <\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Multimodal Retrieval-Augmented Generation (RAG) is a transformative innovation in AI, enabling systems to process and integrate diverse data types such as text, images, audio, and video. This capability is crucial in addressing the challenge of unstructured enterprise data, which predominantly consists of multimodal formats. By leveraging multimodal inputs, RAG enhances contextual understanding, improves accuracy, and [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":143735,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5815,5293,60825,20383,32726],"dealstore":[],"offerexpiration":[],"class_list":["post-143733","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-blogathon","tag-build","tag-docling","tag-multimodal","tag-rag"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>How to Build Multimodal RAG Using Docling? - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=143733\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"How to Build Multimodal RAG Using Docling? - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Multimodal Retrieval-Augmented Generation (RAG) is a transformative innovation in AI, enabling systems to process and integrate diverse data types such as text, images, audio, and video. This capability is crucial in addressing the challenge of unstructured enterprise data, which predominantly consists of multimodal formats. By leveraging multimodal inputs, RAG enhances contextual understanding, improves accuracy, and [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=143733\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-03-19T13:11:20+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"17 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=143733#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=143733\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"How to Build Multimodal RAG Using Docling?\",\"datePublished\":\"2025-03-19T13:11:20+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=143733\"},\"wordCount\":2383,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=143733#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp\",\"keywords\":[\"Blogathon\",\"Build\",\"Docling\",\"Multimodal\",\"RAG\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=143733#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=143733\",\"url\":\"https:\/\/fivemor.com\/?p=143733\",\"name\":\"How to Build Multimodal RAG Using Docling? - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=143733#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=143733#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp\",\"datePublished\":\"2025-03-19T13:11:20+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=143733#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=143733\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=143733#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=143733#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"How to Build Multimodal RAG Using Docling?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"How to Build Multimodal RAG Using Docling? - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=143733","og_locale":"en_US","og_type":"article","og_title":"How to Build Multimodal RAG Using Docling? - Som2ny Network","og_description":"Multimodal Retrieval-Augmented Generation (RAG) is a transformative innovation in AI, enabling systems to process and integrate diverse data types such as text, images, audio, and video. This capability is crucial in addressing the challenge of unstructured enterprise data, which predominantly consists of multimodal formats. By leveraging multimodal inputs, RAG enhances contextual understanding, improves accuracy, and [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=143733","og_site_name":"Som2ny Network","article_published_time":"2025-03-19T13:11:20+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"17 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=143733#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=143733"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"How to Build Multimodal RAG Using Docling?","datePublished":"2025-03-19T13:11:20+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=143733"},"wordCount":2383,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=143733#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp","keywords":["Blogathon","Build","Docling","Multimodal","RAG"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=143733#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=143733","url":"https:\/\/fivemor.com\/?p=143733","name":"How to Build Multimodal RAG Using Docling? - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=143733#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=143733#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp","datePublished":"2025-03-19T13:11:20+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=143733#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=143733"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=143733#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/Multimodal-RAG-using-Docling.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=143733#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"How to Build Multimodal RAG Using Docling?"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/143733","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=143733"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/143733\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/143735"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=143733"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=143733"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=143733"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=143733"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=143733"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}