{"id":108129,"date":"2025-02-25T00:49:34","date_gmt":"2025-02-25T00:49:34","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/mastering-multimodal-rag-with-vertex-ai-gemini-for-content\/"},"modified":"2025-02-25T00:49:34","modified_gmt":"2025-02-25T00:49:34","slug":"mastering-multimodal-rag-with-vertex-ai-gemini-for-content","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=108129","title":{"rendered":"Mastering Multimodal RAG with Vertex AI &#038; Gemini for Content"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Retrieval Augmented Generation (<a href=\"https:\/\/www.youtube.com\/watch?v=bwmX21VsaDk\" target=\"_blank\" rel=\"noreferrer noopener\">RAG<\/a>) has revolutionized how large language models access external data, but traditional approaches are limited to text. With the rise of multimodal data, integrating text and visual information is crucial for comprehensive analysis, especially in complex fields like finance and research. Multimodal RAG addresses this by enabling models to process both text and images for better knowledge retrieval and reasoning. This article explores building a multimodal RAG system using Google\u2019s <a href=\"https:\/\/www-analyticsvidhya-com.webpkgcache.com\/doc\/-\/s\/www.analyticsvidhya.com\/blog\/2023\/12\/what-is-google-gemini-features-usage-and-limitations\/\" target=\"_blank\" rel=\"noreferrer noopener\">Gemini<\/a> models, <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/02\/build-deploy-and-manage-ml-models-with-google-vertex-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">Vertex AI<\/a>, and <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/06\/langchain-guide\/\" target=\"_blank\" rel=\"noreferrer noopener\">LangChain<\/a>, guiding you through environment setup, data processing, embedding generation, and constructing a robust document search engine.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-learning-objectives\">Learning Objectives<\/h4>\n<ul class=\"wp-block-list\">\n<li>Understand the concept of Multimodal RAG and its significance in enhancing data retrieval.<\/li>\n<li>Learn how Gemini can be used to process and integrate both text and visual data.<\/li>\n<li>Explore the capabilities of Vertex AI in building scalable AI models for real-time applications.<\/li>\n<li>Gain insight into how LangChain facilitates seamless integration of language models with external data sources.<\/li>\n<li>Learn how to construct shrewd frameworks that use content and visual information for precise, context-aware reactions.<\/li>\n<li>Know how to apply these innovations for utilize cases like substance era, personalized suggestions, and AI associates.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-multimodal-rag-model-an-overview\">Multimodal RAG Model: An Overview<\/h2>\n<p>Multimodal RAG models combine visual and printed information to supply more strong and context-aware yields. Not at all like conventional Cloth models, which exclusively depend on content, multimodal Clothes are outlined to get and consolidate visual substance such as graphs, charts, and pictures. This dual-processing capability is particularly valuable for analyzing complex records where visuals are as enlightening as content, such as money-related reports, logical papers, or client manuals.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"872\" height=\"572\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1_eTqZBKyd1EMXUcDxfFNWNg.webp\" alt=\"multimodal Retrieval Augmented Generation (RAG) system architecture\" class=\"wp-image-222371\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1_eTqZBKyd1EMXUcDxfFNWNg.webp 872w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1_eTqZBKyd1EMXUcDxfFNWNg-300x197.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1_eTqZBKyd1EMXUcDxfFNWNg-768x504.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1_eTqZBKyd1EMXUcDxfFNWNg-150x98.webp 150w\" sizes=\"(max-width: 872px) 100vw, 872px\"\/><figcaption class=\"wp-element-caption\">Source: Author<\/figcaption><\/figure>\n<\/div>\n<p>By preparing content and pictures, the show offers a more profound understanding of the substance, driving to more precise and smart reactions. This integration relieves the chance of producing deceiving or relevantly erroneous data (commonly known as visualization in machine learning), coming about in more dependable yields for decision-making and investigation.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-key-technologies-used\">Key Technologies Used<\/h2>\n<p>Here\u2019s a summary of each key technology:<\/p>\n<ol class=\"wp-block-list\">\n<li><b>Gemini by Google DeepMind:<\/b> A robust generative AI suite designed for multimodal functions, capable of processing and creating text and images seamlessly.<\/li>\n<li><b>Vertex AI:<\/b> A comprehensive platform for developing, deploying, and scaling machine learning models, known for its vector search feature for multimodal data retrieval.<\/li>\n<li><b>LangChain:<\/b> A tool that streamlines the integration of large language models (LLMs) with various tools and data sources, supporting the connection between models, embeddings, and external resources.<\/li>\n<li><b>Retrieval-Augmented Generation (RAG) Framework:<\/b> Combines retrieval-based and generation-based models to enhance response accuracy by pulling context from external sources before generating outputs, ideal for multimodal content handling.<\/li>\n<li><b>OpenAI\u2019s DALL\u00b7E: <\/b>An image-generation model that translates textual prompts into visual content, enhancing multimodal RAG outputs with tailored and contextually relevant imagery.<\/li>\n<li><b>Transformers for Multimodal Processing:<\/b> The backbone architecture for handling mixed input types, enabling models to process and generate responses involving both text and visual data efficiently.<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-model-architecture-explained\">Model Architecture Explained<\/h2>\n<p>The architecture of a multimodal RAG system involves:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Gemini for Multimodal Processing: <\/b>Handles both text and visual inputs, extracting detailed information.<\/li>\n<li><b>Vertex AI Vector Search:<\/b> Provides a vector store for embedding management, enabling seamless data retrieval.<\/li>\n<li><b>LangChain MultiVectorRetriever:<\/b> Acts as a mediator for retrieving relevant data from the vector store based on user queries.<\/li>\n<li><b>RAG Framework Integration:<\/b> Combines retrieved data with generative capabilities to create accurate, context-rich responses.<\/li>\n<li><b>Multimodal Encoder-Decoder:<\/b> Processes and fuses textual and visual content, ensuring both types of data contribute effectively to the output.<\/li>\n<li><b>Transformers for Hybrid Data Handling:<\/b> Uses attention mechanisms to align and integrate information from different modalities.<\/li>\n<li><b>Fine-Tuning Pipelines: <\/b>Customized training routines that adjust the model\u2019s performance based on specific multimodal datasets for enhanced accuracy and context understanding.<\/li>\n<\/ul>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"872\" height=\"242\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1718492204884.webp\" alt=\"building a multimodal Retrieval Augmented Generation (RAG) system with Gemini and LangChain\" class=\"wp-image-222373\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1718492204884.webp 872w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1718492204884-300x83.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1718492204884-768x213.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/1718492204884-150x42.webp 150w\" sizes=\"auto, (max-width: 872px) 100vw, 872px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-building-a-multimodal-rag-system-with-vertex-ai-gemini-and-langchain\">Building a Multimodal RAG System with Vertex AI, Gemini, and LangChain<\/h2>\n<p>Now let\u2019s get into the actual coding part. In this section, I will guide you through the steps of building a multimodal RAG system for content and images, using Google Gemini, Vertex AI, and LangChain.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-1-setting-up-your-development-environment\">Step 1: Setting Up Your Development Environment<\/h3>\n<p>\u00a0Let\u2019s begin by setting up the environment.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-1-install-necessary-packages\">1. Install necessary packages<\/h4>\n<p>The %pip install command installs all the necessary Python libraries, including google-cloud-aiplatform, langchain, and various document-processing libraries like pypdf.<\/p>\n<pre class=\"wp-block-code\"><code>%pip install -U -q google-cloud-aiplatform langchain-core langchain-google-vertexai langchain-text-splitters langchain-community \"unstructured[all-docs]\" pypdf pydantic lxml pillow matplotlib opencv-python tiktoken<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-2-restart-the-runtime-to-make-sure-new-packages-are-accessible\">2. Restart the runtime to make sure new packages are accessible<\/h4>\n<pre class=\"wp-block-code\"><code>import IPython\n\napp = IPython.Application.instance()\napp.kernel.do_shutdown(True)<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-3-authenticate-the-notebook-environment-google-colab-only\">3. Authenticate the notebook environment (Google Colab only)<\/h4>\n<p>Add the code to authenticate and initialize the Vertex AI environment<br \/>The auth.authenticate_user() function is used for authenticating your Google Cloud account in Google Colab.<\/p>\n<pre class=\"wp-block-code\"><code>import sys\n\n# Additional authentication is required for Google Colab\nif \"google.colab\" in sys.modules:\n    # Authenticate user to Google Cloud\n    from google.colab import auth\n\n    auth.authenticate_user()<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-2-define-google-cloud-project-information\">Step 2: Define Google Cloud Project Information<\/h3>\n<ul class=\"wp-block-list\">\n<li>PROJECT_ID and LOCATION: Define your Google Cloud project and location.<\/li>\n<li>Vertex AI SDK Initialization: The aiplatform.init() function initializes the Vertex AI SDK with your project and bucket information.<\/li>\n<\/ul>\n<p>PROJECT_ID = \u201cYOUR_PROJECT_ID\u201d # @param {type:\u201dstring\u201d}<\/p>\n<pre class=\"wp-block-code\"><code>PROJECT_ID = \"YOUR_PROJECT_ID\"  # @param {type:\"string\"}\nLOCATION = \"us-central1\"  # @param {type:\"string\"}\n\n# For Vector Search Staging\nGCS_BUCKET = \"YOUR_BUCKET_NAME\"  # @param {type:\"string\"}\nGCS_BUCKET_URI = f\"gs:\/\/{GCS_BUCKET}\"<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-3-initialize-the-vertex-ai-sdk\">Step 3: Initialize the Vertex AI SDK<\/h3>\n<pre class=\"wp-block-code\"><code>from google.cloud import aiplatform\n\naiplatform.init(project=PROJECT_ID, location=LOCATION, staging_bucket=GCS_BUCKET_URI)<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-4-import-necessary-libraries\">Step 4: Import Necessary Libraries<\/h3>\n<p>Add the code for constructing the document repository and integrating LangChain:<br \/>Imports various libraries like langchain, IPython, pillow, and others needed for the retrieval and processing pipeline.<\/p>\n<pre class=\"wp-block-code\"><code>import base64\nimport os\nimport re\nimport uuid\n\nfrom IPython.display import Image, Markdown, display\nfrom langchain.prompts import PromptTemplate\nfrom langchain.retrievers.multi_vector import MultiVectorRetriever\nfrom langchain.storage import InMemoryStore\nfrom langchain_core.documents import Document\nfrom langchain_core.messages import AIMessage, HumanMessage\nfrom langchain_core.output_parsers import StrOutputParser\nfrom langchain_core.runnables import RunnableLambda, RunnablePassthrough\nfrom langchain_google_vertexai import (\n    ChatVertexAI,\n    VectorSearchVectorStore,\n    VertexAI,\n    VertexAIEmbeddings,\n)\nfrom langchain_text_splitters import CharacterTextSplitter\nfrom unstructured.partition.pdf import partition_pdf\n\n# from langchain_community.vectorstores import Chroma  # Optional<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-5-define-model-information\">Step 5: Define Model Information<\/h3>\n<pre class=\"wp-block-code\"><code>MODEL_NAME = \"gemini-1.5-flash\"\nGEMINI_OUTPUT_TOKEN_LIMIT = 8192\n\nEMBEDDING_MODEL_NAME = \"text-embedding-004\"\nEMBEDDING_TOKEN_LIMIT = 2048\n\nTOKEN_LIMIT = min(GEMINI_OUTPUT_TOKEN_LIMIT, EMBEDDING_TOKEN_LIMIT)<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-6-load-the-data\">Step 6: Load the Data<\/h3>\n<h4 class=\"wp-block-heading\" id=\"h-1-get-documents-and-images-from-gcs\">1. Get documents and images from GCS<\/h4>\n<pre class=\"wp-block-code\"><code># Download documents and images used in this notebook\n!gsutil -m rsync -r gs:\/\/github-repo\/rag\/intro_multimodal_rag\/ .\nprint(\"Download completed\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-2-extract-images-tables-and-chunk-text-from-a-pdf-file\">2. Extract images, tables, and chunk text from a PDF file<\/h4>\n<ul class=\"wp-block-list\">\n<li>Partitions a PDF into tables and text using partition_pdf from unstructured.<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>pdf_folder_path = \"\/content\/data\/\" if \"google.colab\" in sys.modules else \"data\/\"\npdf_file_name = \"google-10k-sample-14pages.pdf\"\n\n# Extract images, tables, and chunk text from a PDF file.\nraw_pdf_elements = partition_pdf(\n    filename=pdf_file_name,\n    extract_images_in_pdf=False,\n    infer_table_structure=True,\n    chunking_strategy=\"by_title\",\n    max_characters=4000,\n    new_after_n_chars=3800,\n    combine_text_under_n_chars=2000,\n    image_output_dir_path=pdf_folder_path,\n)\n\n# Categorize extracted elements from a PDF into tables and texts.\ntables = []\ntexts = []\nfor element in raw_pdf_elements:\n    if \"unstructured.documents.elements.Table\" in str(type(element)):\n        tables.append(str(element))\n    elif \"unstructured.documents.elements.CompositeElement\" in str(type(element)):\n        texts.append(str(element))\n\n# Optional: Enforce a specific token size for texts\ntext_splitter = CharacterTextSplitter.from_tiktoken_encoder(\n    chunk_size=10000, chunk_overlap=0\n)\njoined_texts = \" \".join(texts)\ntexts_4k_token = text_splitter.split_text(joined_texts)<\/code><\/pre>\n<ul class=\"wp-block-list\">\n<li>Generate summaries of text elements<\/li>\n<li>A function generate_text_summaries uses Vertex AI\u2019s model to summarize text and tables extracted from the PDF for later use in retrieval.<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>def generate_text_summaries(\n    texts: list[str], tables: list[str], summarize_texts: bool = False\n) -&gt; tuple[list, list]:\n    \"\"\"\n    Summarize text elements\n    texts: List of str\n    tables: List of str\n    summarize_texts: Bool to summarize texts\n    \"\"\"\n\n    # Prompt\n    prompt_text = \"\"\"You are an assistant tasked with summarizing tables and text for retrieval. \\\n    These summaries will be embedded and used to retrieve the raw text or table elements. \\\n    Give a concise summary of the table or text that is well optimized for retrieval. Table or text: {element} \"\"\"\n    prompt = PromptTemplate.from_template(prompt_text)\n    empty_response = RunnableLambda(\n        lambda x: AIMessage(content=\"Error processing document\")\n    )\n    # Text summary chain\n    model = VertexAI(\n        temperature=0, model_name=MODEL_NAME, max_output_tokens=TOKEN_LIMIT\n    ).with_fallbacks([empty_response])\n    summarize_chain = {\"element\": lambda x: x} | prompt | model | StrOutputParser()\n\n    # Initialize empty summaries\n    text_summaries = []\n    table_summaries = []\n\n    # Apply to text if texts are provided and summarization is requested\n    if texts:\n        if summarize_texts:\n            text_summaries = summarize_chain.batch(texts, {\"max_concurrency\": 1})\n        else:\n            text_summaries = texts\n\n    # Apply to tables if tables are provided\n    if tables:\n        table_summaries = summarize_chain.batch(tables, {\"max_concurrency\": 1})\n\n    return text_summaries, table_summaries\n\n\n# Get text, table summaries\ntext_summaries, table_summaries = generate_text_summaries(\n    texts_4k_token, tables, summarize_texts=True\n)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>def encode_image(image_path: str) -&gt; str:\n    \"\"\"Getting the base64 string\"\"\"\n    with open(image_path, \"rb\") as image_file:\n        return base64.b64encode(image_file.read()).decode(\"utf-8\")\n\n\ndef image_summarize(model: ChatVertexAI, base64_image: str, prompt: str) -&gt; str:\n    \"\"\"Make image summary\"\"\"\n    msg = model.invoke(\n        [\n            HumanMessage(\n                content=[\n                    {\"type\": \"text\", \"text\": prompt},\n                    {\n                        \"type\": \"image_url\",\n                        \"image_url\": {\"url\": f\"data:image\/png;base64,{base64_image}\"},\n                    },\n                ]\n            )\n        ]\n    )\n    return msg.content\n\n\ndef generate_img_summaries(path: str) -&gt; tuple[list[str], list[str]]:\n    \"\"\"\n    Generate summaries and base64 encoded strings for images\n    path: Path to list of .jpg files extracted by Unstructured\n    \"\"\"\n\n    # Store base64 encoded images\n    img_base64_list = []\n\n    # Store image summaries\n    image_summaries = []\n\n    # Prompt\n    prompt = \"\"\"You are an assistant tasked with summarizing images for retrieval. \\\n    These summaries will be embedded and used to retrieve the raw image. \\\n    Give a concise summary of the image that is well optimized for retrieval.\n    If it's a table, extract all elements of the table.\n    If it's a graph, explain the findings in the graph.\n    Do not include any numbers that are not mentioned in the image.\n    \"\"\"\n\n    model = ChatVertexAI(model_name=MODEL_NAME, max_output_tokens=TOKEN_LIMIT)\n\n    # Apply to images\n    for img_file in sorted(os.listdir(path)):\n        if img_file.endswith(\".png\"):\n            base64_image = encode_image(os.path.join(path, img_file))\n            img_base64_list.append(base64_image)\n            image_summaries.append(image_summarize(model, base64_image, prompt))\n\n    return img_base64_list, image_summaries\n\n\n# Image summaries\nimg_base64_list, image_summaries = generate_img_summaries(\".\")<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-7-create-and-deploy-a-vertex-ai-vector-search-index-and-endpoint\">Step 7: Create and Deploy a Vertex AI Vector Search Index and Endpoint<\/h3>\n<pre class=\"wp-block-code\"><code># https:\/\/cloud.google.com\/vertex-ai\/generative-ai\/docs\/model-reference\/text-embeddings\nDIMENSIONS = 768  # Dimensions output from textembedding-gecko\n\nindex = aiplatform.MatchingEngineIndex.create_tree_ah_index(\n    display_name=\"mm_rag_langchain_index\",\n    dimensions=DIMENSIONS,\n    approximate_neighbors_count=150,\n    leaf_node_embedding_count=500,\n    leaf_nodes_to_search_percent=7,\n    description=\"Multimodal RAG LangChain Index\",\n    index_update_method=\"STREAM_UPDATE\",\n)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>DEPLOYED_INDEX_ID = \"mm_rag_langchain_index_endpoint\"\n\nindex_endpoint = aiplatform.MatchingEngineIndexEndpoint.create(\n    display_name=DEPLOYED_INDEX_ID,\n    description=\"Multimodal RAG LangChain Index Endpoint\",\n    public_endpoint_enabled=True,\n)<\/code><\/pre>\n<ul class=\"wp-block-list\">\n<li>Deploy Index to Index Endpoint<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>index_endpoint = index_endpoint.deploy_index(\n    index=index, deployed_index_id=\"mm_rag_langchain_deployed_index\"\n)\nindex_endpoint.deployed_indexes<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-8-create-retriever-and-load-documents\">Step 8: Create Retriever and Load Documents<\/h3>\n<pre class=\"wp-block-code\"><code># The vectorstore to use to index the summaries\nvectorstore = VectorSearchVectorStore.from_components(\n    project_id=PROJECT_ID,\n    region=LOCATION,\n    gcs_bucket_name=GCS_BUCKET,\n    index_id=index.name,\n    endpoint_id=index_endpoint.name,\n    embedding=VertexAIEmbeddings(model_name=EMBEDDING_MODEL_NAME),\n    stream_update=True,\n)<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>docstore = InMemoryStore()\n\nid_key = \"doc_id\"\n# Create the multi-vector retriever\nretriever_multi_vector_img = MultiVectorRetriever(\n    vectorstore=vectorstore,\n    docstore=docstore,\n    id_key=id_key,\n)<\/code><\/pre>\n<p>\u2022 Load data into Document Store and Vector Store<\/p>\n<pre class=\"wp-block-code\"><code># Raw Document Contents\ndoc_contents = texts + tables + img_base64_list\n\ndoc_ids = [str(uuid.uuid4()) for _ in doc_contents]\nsummary_docs = [\n    Document(page_content=s, metadata={id_key: doc_ids[i]})\n    for i, s in enumerate(text_summaries + table_summaries + image_summaries)\n]\n\nretriever_multi_vector_img.docstore.mset(list(zip(doc_ids, doc_contents)))\n\n# If using Vertex AI Vector Search, this will take a while to complete.\n# You can cancel this cell and continue later.\nretriever_multi_vector_img.vectorstore.add_documents(summary_docs)<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-9-create-chain-with-retriever-and-gemini-llm\">Step 9: Create Chain with Retriever and Gemini LLM<\/h3>\n<pre class=\"wp-block-code\"><code>def looks_like_base64(sb):\n    \"\"\"Check if the string looks like base64\"\"\"\n    return re.match(\"^[A-Za-z0-9+\/]+[=]{0,2}$\", sb) is not None\n\n\ndef is_image_data(b64data):\n    \"\"\"\n    Check if the base64 data is an image by looking at the start of the data\n    \"\"\"\n    image_signatures = {\n        b\"\\xFF\\xD8\\xFF\": \"jpg\",\n        b\"\\x89\\x50\\x4E\\x47\\x0D\\x0A\\x1A\\x0A\": \"png\",\n        b\"\\x47\\x49\\x46\\x38\": \"gif\",\n        b\"\\x52\\x49\\x46\\x46\": \"webp\",\n    }\n    try:\n        header = base64.b64decode(b64data)[:8]  # Decode and get the first 8 bytes\n        for sig, format in image_signatures.items():\n            if header.startswith(sig):\n                return True\n        return False\n    except Exception:\n        return False\n\n\ndef split_image_text_types(docs):\n    \"\"\"\n    Split base64-encoded images and texts\n    \"\"\"\n    b64_images = []\n    texts = []\n    for doc in docs:\n        # Check if the document is of type Document and extract page_content if so\n        if isinstance(doc, Document):\n            doc = doc.page_content\n        if looks_like_base64(doc) and is_image_data(doc):\n            b64_images.append(doc)\n        else:\n            texts.append(doc)\n    return {\"images\": b64_images, \"texts\": texts}\n\n\ndef img_prompt_func(data_dict):\n    \"\"\"\n    Join the context into a single string\n    \"\"\"\n    formatted_texts = \"\\n\".join(data_dict[\"context\"][\"texts\"])\n    messages = [\n        {\n            \"type\": \"text\",\n            \"text\": (\n                \"You are financial analyst tasking with providing investment advice.\\n\"\n                \"You will be given a mix of text, tables, and image(s) usually of charts or graphs.\\n\"\n                \"Use this information to provide investment advice related to the user's question. \\n\"\n                f\"User-provided question: {data_dict['question']}\\n\\n\"\n                \"Text and \/ or tables:\\n\"\n                f\"{formatted_texts}\"\n            ),\n        }\n    ]\n\n    # Adding image(s) to the messages if present\n    if data_dict[\"context\"][\"images\"]:\n        for image in data_dict[\"context\"][\"images\"]:\n            messages.append(\n                {\n                    \"type\": \"image_url\",\n                    \"image_url\": {\"url\": f\"data:image\/jpeg;base64,{image}\"},\n                }\n            )\n    return [HumanMessage(content=messages)]\n\n\n# Create RAG chain\nchain_multimodal_rag = (\n    {\n        \"context\": retriever_multi_vector_img | RunnableLambda(split_image_text_types),\n        \"question\": RunnablePassthrough(),\n    }\n    | RunnableLambda(img_prompt_func)\n    | ChatVertexAI(\n        temperature=0,\n        model_name=MODEL_NAME,\n        max_output_tokens=TOKEN_LIMIT,\n    )  # Multi-modal LLM\n    | StrOutputParser()\n)<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-10-test-the-model\">Step 10: Test the Model<\/h3>\n<h4 class=\"wp-block-heading\" id=\"h-1-process-user-query\">1. Process User Query<\/h4>\n<pre class=\"wp-block-code\"><code>query = \"What are the EV \/ NTM and NTM rev growth for MongoDB, Cloudflare, and Datadog?\n\"<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-2-get-retrieved-documents\">2. Get Retrieved documents<\/h4>\n<pre class=\"wp-block-code\"><code># List of source documents\ndocs = retriever_multi_vector_img.get_relevant_documents(query, limit=1)\n\n# We get relevant docs\nlen(docs)\n\ndocs<\/code><\/pre>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"856\" height=\"164\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/74ecaca749ae459a_856.webp\" alt=\"RAG system with Vertex AI, Google Gemini, and LangChain\" class=\"wp-image-222374\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/74ecaca749ae459a_856.webp 856w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/74ecaca749ae459a_856-300x57.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/74ecaca749ae459a_856-768x147.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/74ecaca749ae459a_856-150x29.webp 150w\" sizes=\"auto, (max-width: 856px) 100vw, 856px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-3-get-generative-response\">3. Get generative response<\/h4>\n<pre class=\"wp-block-code\"><code>plt_img_base64(docs[3])<\/code><\/pre>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"856\" height=\"425\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856.webp\" alt=\"EV \/ NTM revenue multiples\" class=\"wp-image-222375\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856.webp 856w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856-300x149.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856-768x381.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856-150x74.webp 150w\" sizes=\"auto, (max-width: 856px) 100vw, 856px\"\/><\/figure>\n<pre class=\"wp-block-code\"><code>result = chain_multimodal_rag.invoke(query)\n\nfrom IPython.display import Markdown as md\nmd(result)<\/code><\/pre>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"680\" height=\"268\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856_1.webp\" alt=\"Vertex AI, Google Gemini, and LangChain\" class=\"wp-image-222376\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856_1.webp 680w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856_1-300x118.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/989ad388127f5d60_856_1-150x59.webp 150w\" sizes=\"auto, (max-width: 680px) 100vw, 680px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-practical-applications\">Practical Applications<\/h2>\n<ol class=\"wp-block-list\">\n<li><b>Financial Analysis: <\/b>In financial analysis, information from money-related reports such as adjust sheets, salary articulations, and cash stream reports can be extricated to evaluate a company\u2019s execution and make educated choices.<\/li>\n<li><b>Healthcare: <\/b>Cross-referencing restorative records with pictures like X-rays makes a difference specialists to create precise analyze by comparing the patient\u2019s history with visual information.<\/li>\n<li><b>Education:<\/b> In education, providing explanations alongside diagrams aids in visualizing complex concepts, making them easier to understand and enhancing retention for students.<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>Multimodal RAG (Retrieval-Augmented Generation) combines text and visual data to enhance information retrieval, enabling more contextually accurate and comprehensive AI responses. By leveraging tools like Gemini, Vertex AI, and LangChain, developers can build intelligent systems that efficiently process both textual and visual data.<\/p>\n<p>Gemini enables understanding of diverse data types, while Vertex AI supports scalable model deployment for real-time applications. LangChain streamlines integration with external APIs and databases, allowing seamless interaction with multiple data sources. Together, these technologies provide powerful capabilities for creating context-aware, data-rich systems for use in areas like content generation, personalized recommendations, and interactive AI assistants.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h4>\n<ul class=\"wp-block-list\">\n<li>Multimodal RAG combines text and visual data for more accurate, context-aware information retrieval.<\/li>\n<li>Gemini helps process and understand both text and images, enhancing data richness.<\/li>\n<li>Vertex AI offers tools for scalable, efficient AI model deployment, improving real-time performance.<\/li>\n<li>LangChain simplifies the integration of language models with external data sources, enabling seamless data interaction.<\/li>\n<li>These technologies enable the creation of intelligent systems that improve content generation, personalized recommendations, and interactive AI assistants.<\/li>\n<li>The combination of these tools broadens the scope of AI applications, making them more versatile and accurate across diverse use cases.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1739888920232\"><strong class=\"schema-faq-question\">Q1. What is Multimodal RAG, and why is it important?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. Multimodal RAG (Retrieval Augmented Generation) combines text and visual data to improve the accuracy and context of information retrieval, allowing AI systems to provide more comprehensive and relevant responses.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739945727377\"><strong class=\"schema-faq-question\">Q2. How does Gemini contribute to Multimodal RAG?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. Gemini, by Google, is designed to process both text and visual data, enabling AI models to understand and generate insights from mixed data types, enhancing the overall performance of multimodal systems.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739945776657\"><strong class=\"schema-faq-question\">Q3. What is Vertex AI, and how does it support building intelligent systems?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. Vertex AI may be a stage by Google Cloud that provides tools for sending and overseeing AI models at scale. It streamlines the method of building, preparing, and optimizing models, making it simpler for engineers to execute effective multimodal frameworks.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739945810806\"><strong class=\"schema-faq-question\">Q4. What is LangChain, and how does it enhance AI model integration?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. LangChain is a framework that helps integrate large language models with external data sources, APIs, and databases. It enables seamless interaction with different types of data, enhancing the capabilities of multimodal RAG systems.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739945821823\"><strong class=\"schema-faq-question\">Q5. What are some practical applications of Multimodal RAG in real-world scenarios?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. Multimodal RAG can be applied in areas like personalized recommendations, content generation, image-captioning, healthcare (cross-referencing X-rays with medical records), and AI assistants that provide context-aware responses.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/soumyadarshan5263131\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_0GnayIY.webp\" width=\"48\" height=\"48\" alt=\"Soumyadarshan Dash\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Hello there! I&#8217;m Soumyadarshan Dash, a passionate and enthusiastic person when it comes to data science and machine learning. I&#8217;m constantly exploring new topics and techniques in this field, always striving to expand my knowledge and skills. In fact, upskilling myself is not just a hobby, but a way of life for me.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Retrieval Augmented Generation (RAG) has revolutionized how large language models access external data, but traditional approaches are limited to text. With the rise of multimodal data, integrating text and visual information is crucial for comprehensive analysis, especially in complex fields like finance and research. Multimodal RAG addresses this by enabling models to process both text [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":108130,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5815,7222,20726,17656,20383,32726,33667],"dealstore":[],"offerexpiration":[],"class_list":["post-108129","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-blogathon","tag-content","tag-gemini","tag-mastering","tag-multimodal","tag-rag","tag-vertex"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Mastering Multimodal RAG with Vertex AI &amp; Gemini for Content - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=108129\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Mastering Multimodal RAG with Vertex AI &amp; Gemini for Content - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Retrieval Augmented Generation (RAG) has revolutionized how large language models access external data, but traditional approaches are limited to text. With the rise of multimodal data, integrating text and visual information is crucial for comprehensive analysis, especially in complex fields like finance and research. Multimodal RAG addresses this by enabling models to process both text [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=108129\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-25T00:49:34+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"15 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=108129#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=108129\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Mastering Multimodal RAG with Vertex AI &#038; Gemini for Content\",\"datePublished\":\"2025-02-25T00:49:34+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=108129\"},\"wordCount\":1490,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=108129#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg\",\"keywords\":[\"Blogathon\",\"Content\",\"Gemini\",\"Mastering\",\"Multimodal\",\"RAG\",\"Vertex\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=108129#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=108129\",\"url\":\"https:\/\/fivemor.com\/?p=108129\",\"name\":\"Mastering Multimodal RAG with Vertex AI & Gemini for Content - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=108129#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=108129#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg\",\"datePublished\":\"2025-02-25T00:49:34+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=108129#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=108129\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=108129#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=108129#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Mastering Multimodal RAG with Vertex AI &#038; Gemini for Content\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Mastering Multimodal RAG with Vertex AI & Gemini for Content - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=108129","og_locale":"en_US","og_type":"article","og_title":"Mastering Multimodal RAG with Vertex AI & Gemini for Content - Som2ny Network","og_description":"Retrieval Augmented Generation (RAG) has revolutionized how large language models access external data, but traditional approaches are limited to text. With the rise of multimodal data, integrating text and visual information is crucial for comprehensive analysis, especially in complex fields like finance and research. Multimodal RAG addresses this by enabling models to process both text [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=108129","og_site_name":"Som2ny Network","article_published_time":"2025-02-25T00:49:34+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg","type":"image\/jpeg"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"15 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=108129#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=108129"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Mastering Multimodal RAG with Vertex AI &#038; Gemini for Content","datePublished":"2025-02-25T00:49:34+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=108129"},"wordCount":1490,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=108129#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg","keywords":["Blogathon","Content","Gemini","Mastering","Multimodal","RAG","Vertex"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=108129#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=108129","url":"https:\/\/fivemor.com\/?p=108129","name":"Mastering Multimodal RAG with Vertex AI & Gemini for Content - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=108129#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=108129#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg","datePublished":"2025-02-25T00:49:34+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=108129#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=108129"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=108129#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Mastering-Multimodal-RAG-with-Vertex-AI-Gemini-for-Content-and-Images.jpg","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=108129#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Mastering Multimodal RAG with Vertex AI &#038; Gemini for Content"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/108129","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=108129"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/108129\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/108130"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=108129"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=108129"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=108129"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=108129"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=108129"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}