{"id":163144,"date":"2025-03-29T05:16:29","date_gmt":"2025-03-29T05:16:29","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/how-to-build-multimodal-rag-with-gemma-3-docling\/"},"modified":"2025-03-29T05:16:29","modified_gmt":"2025-03-29T05:16:29","slug":"how-to-build-multimodal-rag-with-gemma-3-docling","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=163144","title":{"rendered":"How to Build Multimodal RAG with Gemma 3 &#038; Docling?"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>In this tutorial, we explore how to set up and execute a sophisticated retrieval-augmented generation (RAG) pipeline in Google Colab. We leverage multiple state-of-the-art tools and libraries-including Gemma 3 for language and vision tasks, Docling for document conversion, LangChain for chain-of-thought orchestration, and Milvus as our vector database build a multimodal system that understands and processes text, tables, and images. Let\u2019s dive into each component and see how they work together.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-multimodal-rag\">What is Multimodal RAG?<\/h2>\n<p>Multimodal RAG (Retrieval-Augmented Generation) extends traditional text-based RAG systems by integrating multiple data modalities, in this case, text, tables, and images. This means that the pipeline not only processes and retrieves text but also leverages vision models to understand and describe image content, making the solution more comprehensive. This multimodal approach is particularly beneficial for documents like annual reports that often contain visual elements, such as charts and diagrams.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-proposed-architecture-of-multimodal-rag-with-gemma-3\">Proposed Architecture of Multimodal RAG with Gemma 3<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><a href=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-122.webp\"><img fetchpriority=\"high\" decoding=\"async\" width=\"1117\" height=\"744\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-122.webp\" alt=\"architecture\" class=\"wp-image-228810\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-122.webp 1117w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-122-300x200.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-122-768x512.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-122-150x100.webp 150w\" sizes=\"(max-width: 1117px) 100vw, 1117px\"\/><\/a><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>The aim of this project is to build a robust <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/enhancing-multimodal-rag-capabilities-using-docling\/\" target=\"_blank\" rel=\"noreferrer noopener\">multimodal RAG<\/a> pipeline that can ingest documents (like PDFs), process text and images, store document embeddings in a vector database, and answer queries by retrieving relevant information. This setup is particularly useful for applications such as analyzing annual reports, extracting financial statements, or summarizing technical papers. By integrating various libraries and tools, we combine the power of language models with document conversion and vector search to create a comprehensive end-to-end solution.<\/p>\n<p>The pipeline uses several key libraries and tools:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Colab-Xterm Extension<\/strong>: This adds terminal support in Colab, allowing us to run shell commands and manage the environment efficiently.<\/li>\n<li><strong>Ollama Models<\/strong>: Provides pre-trained models such as Gemma3, which are used for both language and vision tasks.<\/li>\n<li><strong>Transformers<\/strong>: From Hugging Face, for model loading and tokenization.<\/li>\n<li><strong>LangChain<\/strong>: Orchestrates the chain of processing steps, from prompt creation to document retrieval and generation.<\/li>\n<li><strong>Docling<\/strong>: Converts PDF documents into structured formats, enabling extraction of text, tables, and images.<\/li>\n<li><strong>Milvus<\/strong>: A vector database that stores document embeddings and supports efficient similarity search.<\/li>\n<li><strong>Hugging Face CLI<\/strong>: Used for logging into Hugging Face to access certain models.<\/li>\n<li><strong>Additional Utilities<\/strong>: Such as Pillow for image processing and IPython for display functionalities.<\/li>\n<\/ul>\n<p>Also read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/09\/guide-to-building-multimodal-rag-systems\/\" target=\"_blank\" rel=\"noreferrer noopener\">A Comprehensive Guide to Building Multimodal RAG Systems<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-building-multimodal-rag-with-gemma-3\">Building Multimodal Rag with Gemma 3<\/h2>\n<p>We are building multimodal rag: This approach improves contextual understanding, accuracy, and relevance, especially in fields like healthcare, research, and media analysis. By leveraging cross-modal embeddings, hybrid retrieval strategies, and vision-language models, multimodal RAG systems can provide richer and more insightful responses. The key challenge lies in efficiently integrating and retrieving multimodal data while maintaining coherence and scalability. As AI progresses, developing optimized architectures and retrieval strategies will be crucial for unlocking the full potential of multimodal intelligence. <\/p>\n<h3 class=\"wp-block-heading\" id=\"h-terminal-setup-with-colab-xterm\">Terminal Setup with Colab-Xterm<\/h3>\n<p>First, we install the colab-xterm extension to bring a terminal environment directly into Colab. This allows us to run system commands, install packages, and manage our session more flexibly.<\/p>\n<pre class=\"wp-block-code\"><code>!pip install colab-xterm\u00a0 # Install colab-xterm\n%load_ext colabxterm \u00a0 \u00a0 # Load the xterm extension\n%xterm\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 # Launch an xterm terminal session in Colab<\/code><\/pre>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1708\" height=\"533\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-123.webp\" alt=\"xtream\" class=\"wp-image-228800\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-123.webp 1708w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-123-300x94.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-123-768x240.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-123-1536x479.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-123-150x47.webp 150w\" sizes=\"auto, (max-width: 1708px) 100vw, 1708px\"\/><\/figure>\n<p>This terminal support is especially useful for installing additional dependencies or managing background processes.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-installing-and-managing-ollama-models\">Installing and Managing Ollama Models<\/h3>\n<p>We pull specific <a href=\"https:\/\/ollama.com\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Ollama<\/a> models into our environment using simple shell commands. For example:<\/p>\n<pre class=\"wp-block-code\"><code>!ollama pull gemma3:4b\n!ollama pull llama3.2\n!ollama list<\/code><\/pre>\n<p>These commands ensure that we have the necessary language and vision models available, such as the powerful <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/gemma-3\/\" target=\"_blank\" rel=\"noreferrer noopener\">Gemma 3 model<\/a>, which is central to our multimodal processing.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-installing-essential-python-packages\">Installing Essential Python Packages<\/h3>\n<p>The next step involves installing a host of packages required for our pipeline. This includes libraries for deep learning, text processing, and document handling:<\/p>\n<pre class=\"wp-block-code\"><code>! pip install transformers pillow langchain_community langchain_huggingface langchain_milvus docling langchain_ollama<\/code><\/pre>\n<p>By installing these packages, we prepare the environment for everything from document conversion to retrieval-augmented generation.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-logging-and-hugging-face-authentication\">Logging and Hugging Face Authentication<\/h3>\n<p>Setting up logging is crucial for monitoring pipeline operations:<\/p>\n<pre class=\"wp-block-code\"><code>import logging\nlogging.basicConfig(level=logging.INFO)<\/code><\/pre>\n<p>We also log in to Hugging Face using their CLI to access certain pre-trained models:<\/p>\n<pre class=\"wp-block-code\"><code>!huggingface-cli login<\/code><\/pre>\n<p>This authentication step is necessary for fetching model artifacts and ensuring smooth integration with Hugging Face\u2019s ecosystem.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-configuring-vision-and-language-models-gemma-3\">Configuring Vision and Language Models (Gemma 3)<\/h3>\n<p>The pipeline leverages the Gemma 3 model for both vision and language tasks. For the language side, we set up the model and tokenizer:<\/p>\n<p>This dual setup enables the system to generate textual descriptions from images, making the pipeline truly multimodal.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-document-conversion-with-docling\">Document Conversion with Docling<\/h3>\n<h4 class=\"wp-block-heading\" id=\"h-1-converting-pdfs-to-structured-documents\"><strong>1<\/strong>. Converting PDFs to Structured Documents<\/h4>\n<p>We employ Docling\u2019s DocumentConverter to convert PDFs into structured documents. The conversion process involves extracting text, tables, and images from the source PDFs:<\/p>\n<pre class=\"wp-block-code\"><code>from docling.document_converter import DocumentConverter, PdfFormatOption\nfrom docling.datamodel.base_models import InputFormat\nfrom docling.datamodel.pipeline_options import PdfPipelineOptions\npdf_pipeline_options = PdfPipelineOptions(\n\u00a0\u00a0\u00a0\u00a0do_ocr=False,\n\u00a0\u00a0\u00a0\u00a0generate_picture_images=True,\n)\nformat_options = { InputFormat.PDF: PdfFormatOption(pipeline_options=pdf_pipeline_options) }\nconverter = DocumentConverter(format_options=format_options)\n# Define the sources (URLs) of the documents to be converted.\n# \"https:\/\/arxiv.org\/pdf\/1706.03762\"\nsources = [\n\u00a0\"https:\/\/www.pwc.com\/jm\/en\/research-publications\/pdf\/basic-understanding-of-a-companys-financials.pdf\"\n]\n# Convert the PDF documents from the sources into an internal document format.\nconversions = { source: converter.convert(source=source).document for source in sources }<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-input-file\">Input File<\/h4>\n<p>We\u2019ll be using PwC\u2019s publicly available financial statements. I\u2019ve included the PDF link, and you\u2019re welcome to add your own source links as well!<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-2-extracting-and-chunking-content\">2. Extracting and Chunking Content<\/h4>\n<p>After conversion, we chunk the document into manageable pieces, separating text from tables and images. This segmentation allows each component to be processed independently:<\/p>\n<pre class=\"wp-block-code\"><code>from docling_core.transforms.chunker.hybrid_chunker import HybridChunker\nfrom langchain_core.documents import Document\n# Process text chunks (excluding pure table segments)\ntexts: list[Document] = []\nfor source, docling_document in conversions.items():\n\u00a0\u00a0\u00a0\u00a0for chunk in HybridChunker(tokenizer=embeddings_tokenizer).chunk(docling_document):\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0# Skip table-only chunks; process tables separately\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0if len(chunk.meta.doc_items) == 1:\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0continue\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0document = Document(\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0page_content=chunk.text,\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0metadata={\"source\": source, \"ref\": \"reference details\"}\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0)\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0texts.append(document)<\/code><\/pre>\n<p>This approach not only improves processing efficiency but also facilitates more precise vector storage and retrieval later.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-image-processing-and-encoding\">Image Processing and Encoding<\/h3>\n<p>Images from the documents are processed using Pillow. We convert images into base64-encoded strings that can be embedded directly into prompts:<\/p>\n<pre class=\"wp-block-code\"><code>import base64, io, PIL.Image, PIL.ImageOps\ndef encode_image(image: PIL.Image.Image, format: str = \"png\") -&gt; str:\n\u00a0\u00a0\u00a0\u00a0image = PIL.ImageOps.exif_transpose(image) or image\n\u00a0\u00a0\u00a0\u00a0image = image.convert(\"RGB\")\n\u00a0\u00a0\u00a0\u00a0buffer = io.BytesIO()\n\u00a0\u00a0\u00a0\u00a0image.save(buffer, format)\n\u00a0\u00a0\u00a0\u00a0encoding = base64.b64encode(buffer.getvalue()).decode(\"utf-8\")\n\u00a0\u00a0\u00a0\u00a0return f\"data:image\/{format};base64,{encoding}\"<\/code><\/pre>\n<p>Subsequently, these images are fed into our vision model to generate descriptive text, enhancing the multimodal capabilities of our pipeline.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-creating-a-vector-database-with-milvus\">Creating a Vector Database with Milvus<\/h3>\n<p>To enable fast and accurate retrieval of document embeddings, we set up Milvus as our vector store:<\/p>\n<pre class=\"wp-block-code\"><code>import tempfile\nfrom langchain_core.vectorstores import VectorStore\nfrom langchain_milvus import Milvus\ndb_file = tempfile.NamedTemporaryFile(prefix=\"vectorstore_\", suffix=\".db\", delete=False).name\nvector_db: VectorStore = Milvus(\n\u00a0\u00a0\u00a0\u00a0embedding_function=embeddings_model,\n\u00a0\u00a0\u00a0\u00a0connection_args={\"uri\": db_file},\n\u00a0\u00a0\u00a0\u00a0auto_id=True,\n\u00a0\u00a0\u00a0\u00a0enable_dynamic_field=True,\n\u00a0\u00a0\u00a0\u00a0index_params={\"index_type\": \"AUTOINDEX\"},\n)<\/code><\/pre>\n<p>Documents\u2014whether text, tables, or image descriptions\u2014are then added to the vector database, enabling fast and accurate similarity searches during query execution.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-building-the-retrieval-augmented-generation-rag-chain\">Building the Retrieval-Augmented Generation (RAG) Chain<\/h3>\n<h4 class=\"wp-block-heading\" id=\"h-1-prompt-creation-and-document-wrapping\">1. Prompt Creation and Document Wrapping<\/h4>\n<p>Using LangChain\u2019s prompt templates, we create custom prompts to feed context and queries into our language model:<\/p>\n<pre class=\"wp-block-code\"><code>from langchain.prompts import PromptTemplate\nprompt = \"{input} Given the context: {context}\"\nprompt_template = PromptTemplate.from_template(template=prompt)<\/code><\/pre>\n<p>Each retrieved document is wrapped using a document prompt template, ensuring that the model understands the structure of the input context.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-2-assembling-the-rag-pipeline\">2. Assembling the RAG Pipeline<\/h4>\n<p>We combine the prompt with the vector store to create a retrieval chain that first fetches relevant documents and then uses them to generate a coherent answer:<\/p>\n<pre class=\"wp-block-code\"><code>from langchain.chains.retrieval import create_retrieval_chain\nfrom langchain.chains.combine_documents import create_stuff_documents_chain\ncombine_docs_chain = create_stuff_documents_chain(\n\u00a0\u00a0\u00a0\u00a0llm=model,\n\u00a0\u00a0\u00a0\u00a0prompt=prompt_template,\n\u00a0\u00a0\u00a0\u00a0document_prompt=PromptTemplate.from_template(template=\"\"\"\\\nDocument {doc_id}\n{page_content}\"\"\"),\n\u00a0\u00a0\u00a0\u00a0document_separator=\"\\n\\n\",\n)\nrag_chain = create_retrieval_chain(\n\u00a0\u00a0\u00a0\u00a0retriever=vector_db.as_retriever(),\n\u00a0\u00a0\u00a0\u00a0combine_docs_chain=combine_docs_chain,\n)<\/code><\/pre>\n<p>Queries are then executed against this chain, retrieving context and generating responses based on both the query and the stored document embeddings.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-executing-queries-and-retrieving-information\">Executing Queries and Retrieving Information<\/h3>\n<p>Once the RAG chain is established, you can run queries to retrieve relevant information from your document database. For example:<\/p>\n<pre class=\"wp-block-code\"><code>query = \"Explain Three Key Financial Statements Notes\"\noutputs = rag_chain.invoke({\"input\": query})\nMarkdown(outputs['answer'])<\/code><\/pre>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1165\" height=\"668\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155139.917.webp\" alt=\"Output\" class=\"wp-image-228803\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155139.917.webp 1165w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155139.917-300x172.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155139.917-768x440.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155139.917-150x86.webp 150w\" sizes=\"auto, (max-width: 1165px) 100vw, 1165px\"\/><\/figure>\n<pre class=\"wp-block-code\"><code>query = \"tell me the Contents of an annual report\"\noutputs = rag_chain.invoke({\"input\": query})\nMarkdown(outputs['answer'])<\/code><\/pre>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1151\" height=\"613\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155212.803.webp\" alt=\"Output\" class=\"wp-image-228804\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155212.803.webp 1151w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155212.803-300x160.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155212.803-768x409.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155212.803-150x80.webp 150w\" sizes=\"auto, (max-width: 1151px) 100vw, 1151px\"\/><\/figure>\n<pre class=\"wp-block-code\"><code>query = \"what are the benefits of an annual report?\"\noutputs = rag_chain.invoke({\"input\": query})\nMarkdown(outputs['answer'])<\/code><\/pre>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1142\" height=\"307\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155257.690.webp\" alt=\"output\" class=\"wp-image-228805\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155257.690.webp 1142w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155257.690-300x81.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155257.690-768x206.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/unnamed-2025-03-28T155257.690-150x40.webp 150w\" sizes=\"auto, (max-width: 1142px) 100vw, 1142px\"\/><\/figure>\n<p>The same process can be applied for various queries, such as explaining financial statement notes or summarizing an annual report, thereby demonstrating the versatility of the pipeline.<\/p>\n<p>Here\u2019s the full code: <a href=\"https:\/\/colab.research.google.com\/drive\/1Tnb1_uU9zB8MLsBLQhGQZTG14InE88wX?usp=sharing\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">AV-multimodal-gemma3-rag<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-use-cases\">Use Cases<\/h2>\n<p>This pipeline has numerous applications:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Financial Reporting<\/strong>: Automatically extract and summarize key financial statements, cash flow elements, and annual report details.<\/li>\n<li><strong>Document Analysis<\/strong>: Convert PDFs into structured data for further analysis or machine learning tasks.<\/li>\n<li><strong>Multimodal Search<\/strong>: Enable search and retrieval from mixed media documents, combining textual and visual content.<\/li>\n<li><strong>Business Intelligence<\/strong>: Provide quick insights into complex documents by aggregating and synthesizing information across modalities.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>In this tutorial, we demonstrated how to build a multimodal RAG with Gemma 3 in Google Colab. By integrating tools like Colab-Xterm, Ollama models (Gemma 3), <a href=\"https:\/\/github.com\/docling-project\/docling\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Docling<\/a>, <a href=\"https:\/\/www.langchain.com\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LangChain<\/a>, and <a href=\"https:\/\/milvus.io\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Milvus<\/a>, we created a system capable of processing text, tables, and images. This powerful setup not only enables effective document retrieval but also supports complex query answering and analysis in diverse applications. Whether you\u2019re dealing with financial reports, research papers, or business intelligence tasks, this pipeline offers a versatile and scalable solution.<\/p>\n<p>Happy coding, and enjoy exploring the possibilities of multimodal retrieval-augmented generation!<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/shaik8558834\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_An81zCg.webp\" width=\"48\" height=\"48\" alt=\"Shaik Hamzah\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>GenAI Intern @ Analytics Vidhya | Final Year @ VIT Chennai<br \/>Passionate about AI and machine learning, I&#8217;m eager to dive into roles as an AI\/ML Engineer or Data Scientist where I can make a real impact. With a knack for quick learning and a love for teamwork, I&#8217;m excited to bring innovative solutions and cutting-edge advancements to the table. My curiosity drives me to explore AI across various fields and take the initiative to delve into data engineering, ensuring I stay ahead and deliver impactful projects.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>In this tutorial, we explore how to set up and execute a sophisticated retrieval-augmented generation (RAG) pipeline in Google Colab. We leverage multiple state-of-the-art tools and libraries-including Gemma 3 for language and vision tasks, Docling for document conversion, LangChain for chain-of-thought orchestration, and Milvus as our vector database build a multimodal system that understands and [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":163145,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5293,60825,10099,23318,4977,8558,20383,44762,32726],"dealstore":[],"offerexpiration":[],"class_list":["post-163144","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-build","tag-docling","tag-documents","tag-gemma","tag-images","tag-models","tag-multimodal","tag-query","tag-rag"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>How to Build Multimodal RAG with Gemma 3 &amp; Docling? - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=163144\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"How to Build Multimodal RAG with Gemma 3 &amp; Docling? - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"In this tutorial, we explore how to set up and execute a sophisticated retrieval-augmented generation (RAG) pipeline in Google Colab. We leverage multiple state-of-the-art tools and libraries-including Gemma 3 for language and vision tasks, Docling for document conversion, LangChain for chain-of-thought orchestration, and Milvus as our vector database build a multimodal system that understands and [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=163144\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-03-29T05:16:29+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"1\" \/>\n\t<meta property=\"og:image:height\" content=\"1\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"9 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=163144#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=163144\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"How to Build Multimodal RAG with Gemma 3 &#038; Docling?\",\"datePublished\":\"2025-03-29T05:16:29+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=163144\"},\"wordCount\":1306,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=163144#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp\",\"keywords\":[\"Build\",\"Docling\",\"Documents\",\"Gemma\",\"Images\",\"Models\",\"Multimodal\",\"Query\",\"RAG\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=163144#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=163144\",\"url\":\"https:\/\/fivemor.com\/?p=163144\",\"name\":\"How to Build Multimodal RAG with Gemma 3 & Docling? - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=163144#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=163144#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp\",\"datePublished\":\"2025-03-29T05:16:29+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=163144#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=163144\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=163144#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=163144#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"How to Build Multimodal RAG with Gemma 3 &#038; Docling?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"How to Build Multimodal RAG with Gemma 3 & Docling? - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=163144","og_locale":"en_US","og_type":"article","og_title":"How to Build Multimodal RAG with Gemma 3 & Docling? - Som2ny Network","og_description":"In this tutorial, we explore how to set up and execute a sophisticated retrieval-augmented generation (RAG) pipeline in Google Colab. We leverage multiple state-of-the-art tools and libraries-including Gemma 3 for language and vision tasks, Docling for document conversion, LangChain for chain-of-thought orchestration, and Milvus as our vector database build a multimodal system that understands and [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=163144","og_site_name":"Som2ny Network","article_published_time":"2025-03-29T05:16:29+00:00","og_image":[{"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp","width":1,"height":1,"type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"9 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=163144#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=163144"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"How to Build Multimodal RAG with Gemma 3 &#038; Docling?","datePublished":"2025-03-29T05:16:29+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=163144"},"wordCount":1306,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=163144#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp","keywords":["Build","Docling","Documents","Gemma","Images","Models","Multimodal","Query","RAG"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=163144#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=163144","url":"https:\/\/fivemor.com\/?p=163144","name":"How to Build Multimodal RAG with Gemma 3 & Docling? - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=163144#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=163144#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp","datePublished":"2025-03-29T05:16:29+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=163144#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=163144"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=163144#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/1743225391_Building-a-Multimodal-RAG-Pipeline-using-Gemma-3-and-Docling-1.webp.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=163144#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"How to Build Multimodal RAG with Gemma 3 &#038; Docling?"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/163144","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=163144"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/163144\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/163145"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=163144"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=163144"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=163144"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=163144"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=163144"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}