{"id":50167,"date":"2025-01-27T05:04:03","date_gmt":"2025-01-27T05:04:03","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/a-journey-into-multimodal-llms-part-1\/"},"modified":"2025-01-27T05:04:03","modified_gmt":"2025-01-27T05:04:03","slug":"a-journey-into-multimodal-llms-part-1","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=50167","title":{"rendered":"A Journey into Multimodal LLMs Part 1"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>The human mind naturally perceives language, vision, smell, and touch, enabling us to understand our surroundings. We are particularly inclined toward linguistic thought and visual memory. As GenAI models continue to grow, researchers are now working on extending their capabilities by incorporating multimodality. <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/03\/an-introduction-to-large-language-models-llms\/\" target=\"_blank\" rel=\"noreferrer noopener\">Large Language models (LLMs)<\/a> only accept text as input and produce text as output, which means these models do not process or generate data from other modalities such as images, videos, or voice. LLMs have excelled in handling tasks such as question-answering, text summarization, translation, information retrieval, code generation, and reasoning. However, integrating other modalities with LLMs (Multimodal LLMs) enhances the potential of GenAI models. For instance, training a model by combining text and images solves problems such as visual Q&amp;A, image segmentation, and object detection. Likewise, we can add videos in the same model for more advanced media-related analysis.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-introduction-to-multimodal-llms\">Introduction to Multimodal LLMs<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/11\/generative-ai-for-daily-tasks\/\" target=\"_blank\" rel=\"noreferrer noopener\">Generative AI<\/a> is a subsection of machine learning models allowing for new content generation. We can generate new text after feeding input as text to the model known as text-to-text. However, after extending the capabilities of LLMs with other modalities, we can open the solution to a wide range of use cases such as text-to-image, text-to-video, text-to-speech, image-to-image, and image-to-video. We call such models Large multimodal models (Multimodal LLMs). Training such models happens on large datasets containing text and other modalities so that algorithms can learn the relationships among all the input types. Intuitively, these models are not limited to a single input or output type; they can be adapted to handle inputs from any modality and generate output accordingly. In this way, multimodal LLMs can be seen as providing the system with the ability to process and understand different types of sensory inputs.<\/p>\n<p>This blog is split into two sections; in the first part, I will explore the applications of multimodal LLMs and various architectures, while in the second part, I will train a small vision model.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-datasets\">Datasets<\/h2>\n<p>While combining different input types to create multimodal LLMs may appear straightforward, it becomes more complex when processing data from 1D, 2D, and 3D together. It is a multi-step problem that needs to be solved sequentially in a step-by-step manner, and the data must be carefully curated to enhance the problem-solving capabilities of such models.<\/p>\n<p>For now, we will limit our discussion to text and images. Unlike text, images and videos come in varying sizes and resolutions, so a robust pre-processing technique is needed to standardize all inputs into a single framework. Furthermore, inputs like images, videos, prompts, and metadata should be prepared in a way that helps models build coherent thought processes and maintain logical consistency during inference. Models trained with text, image, and video data are called Large Vision-Language Models (LVLMs).<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-application-of-multimodal-llms\">Application of Multimodal LLMs<\/h2>\n<p>The following image is taken from a <a href=\"https:\/\/arxiv.org\/abs\/2409.12191\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Qwen2-VL paper<\/a> where researchers trained a vision model based on <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/09\/pixtral-12b-vs-qwen2-vl-72b\/\" target=\"_blank\" rel=\"noreferrer noopener\">Qwen2 LLM<\/a> that can solve multiple visual use cases.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"540\" height=\"300\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QqczD6H-thumbnail_webp-600x300-1.webp\" alt=\"Qwen2-VL\" class=\"wp-image-217402\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QqczD6H-thumbnail_webp-600x300-1.webp 540w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QqczD6H-thumbnail_webp-600x300-1-300x167.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QqczD6H-thumbnail_webp-600x300-1-150x83.webp 150w\" sizes=\"(max-width: 540px) 100vw, 540px\"\/><figcaption class=\"wp-element-caption\">Source: Qwen2-VL<\/figcaption><\/figure>\n<\/div>\n<p>The figure below demonstrates how a <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/10\/popular-multimodal-models\/\" target=\"_blank\" rel=\"noreferrer noopener\">Multimodal Language Model (MMLM)<\/a> processes different types of input data (image, text, audio, video) to achieve various objectives. The core part of the diagram, the MMLM, integrates all the different modalities (image, text, audio, video) to process them in combination.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"867\" height=\"368\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QRUQHKM.webp\" alt=\"A generic understanding of the Input and output flow of MMLMs.\" class=\"wp-image-217400\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QRUQHKM.webp 867w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QRUQHKM-300x127.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QRUQHKM-768x326.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_QRUQHKM-150x64.webp 150w\" sizes=\"auto, (max-width: 867px) 100vw, 867px\"\/><figcaption class=\"wp-element-caption\">A generic understanding of the Input and output flow of MMLMs.<\/figcaption><\/figure>\n<\/div>\n<p>Let\u2019s proceed further and understand the different applications of vision models. The complete code used in this blog is stored in <a href=\"https:\/\/github.com\/QuGenAI\/multi-modal\/tree\/main\" target=\"_blank\" rel=\"nofollow noopener\">GitHub<\/a>.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-image-captioning\">1. Image captioning<\/h3>\n<p>It is the task of describing the features of images in words. People are using this feature to generate descriptions of images and innovating a range of engaging captions and relevant hashtags for their social media posts to improve visibility.<\/p>\n<pre class=\"wp-block-code\"><code>image_path = \"quantum.jpg\"\nwith open(image_path, 'rb') as image_file:\n    image_data = image_file.read()\n    \nimage_data = base64.b64encode(image_data).decode(\"utf-8\")\n\nprompt=\"\"\"explain this image\"\"\"\nmessage = HumanMessage(\n    content=[\n        {\"type\": \"text\", \"text\": prompt},\n        {\n            \"type\": \"image_url\",\n            \"image_url\": {\"url\": f\"data:image\/jpeg;base64,{image_data}\"},\n        },\n    ],\n)\nresponse = llm.invoke([message])\nprint(response.content)<\/code><\/pre>\n<p>Information extraction is another application for vision models where we expect the model to retrieve features or data points from the images. For example, we can question the model to identify underlying objects\u2019 colour, text, or feature. Contemporary models use function calling or JSON parsing techniques to extract structured data points from the images.<\/p>\n<pre class=\"wp-block-code\"><code>from langchain.output_parsers import PydanticOutputParser\nfrom langchain_core.prompts import ChatPromptTemplate\nfrom pydantic import BaseModel, Field\nimport json\n\nclass Retrieval(BaseModel):\n    Description: str = Field(description=\"Describe the image\")\n    Machine: str = Field(description=\"Explain what is the machine about\")\n    Color: str = Field(description=\"What are the color used in the image\")\n    People: str = Field(description=\"Count how many male and female are standing their\")\n\nparser = PydanticOutputParser(pydantic_object=Retrieval)\n\nprompt = ChatPromptTemplate.from_messages([\n    (\"system\", \"Extract the requested details as per the given details.\\n'{struct_format}'\\n\"),\n    (\"human\", [\n        {\n            \"type\": \"image_url\",\n            \"image_url\": {\"url\": \"data:image\/jpeg;base64,{image_data}\"},\n        },\n    ]),\n])\n\nchain = prompt | llm | parser\n\nimage_path = \"quantum.jpg\"\nwith open(image_path, 'rb') as image_file:\n    image_data = image_file.read()\n    \nimage_data = base64.b64encode(image_data).decode(\"utf-8\")\n\n\nresponse = chain.invoke({\n    \"struct_format\": parser.get_format_instructions(),\n    \"image_data\": image_data\n})\n\ndata = json.loads(response.model_dump_json())\n\nfor k,v in data.items():\n    print(f\"{k}: {v}\")<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-3-visual-interpretation-amp-reasoning\">3. Visual Interpretation &amp; Reasoning<\/h3>\n<p>It is a use case for a vision model to analyze the image and perform reasoning tasks. For example, the model can interpret the underlying information in images, diagrams, and graphical representations, create step-by-step analyses, and conclude.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-4-ocr-ing\">4. OCR\u2019ing<\/h3>\n<p>It is one of the most important use cases in the area of Document AI where models convert and extract text data from images for downstream tasks.<\/p>\n<pre class=\"wp-block-code\"><code>image_path = \"qubits.png\"\nwith open(image_path, 'rb') as image_file:\n    image_data = image_file.read()\n    \nimage_data = base64.b64encode(image_data).decode(\"utf-8\")\n\nprompt=\"\"\"Extract all the text from the image\"\"\"\nmessage = HumanMessage(\n    content=[\n        {\"type\": \"text\", \"text\": prompt},\n        {\n            \"type\": \"image_url\",\n            \"image_url\": {\"url\": f\"data:image\/jpeg;base64,{image_data}\"},\n        },\n    ],\n)\nresponse = llm.invoke([message])\nprint(response.content)<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-5-object-detection-amp-segmentation\">5. Object Detection &amp; Segmentation<\/h3>\n<p>Vision models are capable of identifying objects in the images and classifying them into defined categories. Mainly in the case of object detection models can locate the objects and classify them whereas in the case of segmentation, vision models can divide the images into different regions based on surrounding pixel values.<\/p>\n<pre class=\"wp-block-code\"><code>from langchain.output_parsers import PydanticOutputParser\nfrom langchain_core.prompts import ChatPromptTemplate\nfrom pydantic import BaseModel, Field\nfrom typing import List\n\nimport json\n\nclass Segmentation(BaseModel):\n    Object: List[str] = Field(description=\"Identify the object and give a name\")\n    Bounding_box: List[List[int]] = Field(description=\"Extract the bounding boxes\")\n\nparser = PydanticOutputParser(pydantic_object=Segmentation)\n\nprompt = ChatPromptTemplate.from_messages([\n    (\"system\", \"Extract all the image objects and their bounding boxes. You must always return valid JSON.\\n'{struct_format}'\\n\"),\n    (\"human\", [\n        {\n            \"type\": \"image_url\",\n            \"image_url\": {\"url\": \"data:image\/jpeg;base64,{image_data}\"},\n        },\n    ]),\n])\n\nchain = prompt | llm | parser\n\nimage_path = \"quantum.jpg\"\nwith open(image_path, 'rb') as image_file:\n    image_data = image_file.read()\n    \nimage_data = base64.b64encode(image_data).decode(\"utf-8\")\n\nresponse = chain.invoke({\n    \"struct_format\": parser.get_format_instructions(),\n    \"image_data\": image_data\n})\n\ndata = json.loads(response.model_dump_json())\n\nfor k,v in data.items():\n    print(f\"{k}: {v}\")\n\n## Complete code is available in GitHub\nplot_bounding_boxes(im=img,labels=data['Object'], bounding_boxes=data['Bounding_box'])<\/code><\/pre>\n<p>Vision models have a wide range of use cases across various industries and are increasingly being integrated into different platforms like Canva, Fireflies, Instagram, and YouTube.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-architecture-of-large-vision-language-models-lvlms\">Architecture of Large Vision-Language Models (LVLMs)<\/h2>\n<p>The primary purpose of developing vision models is to unify features from images, videos, and text. Researchers are exploring different architectures to pretrain Large Vision-Language Models (LVLMs).<br \/>Typically, encoders are employed to extract image features, while text data can be processed using an encoder, a decoder, or a combination of both. Modality projectors, sometimes called connectors, are dense neural networks used to align image features with text representations.<\/p>\n<p>Below is the general overview of common network designs.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-two-tower-vlm\">1. Two-Tower VLM<\/h3>\n<p>The figure below represents the simplest architecture where images and text are encoded separately and trained under a common objective. Here\u2019s a breakdown of the components:<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"457\" height=\"322\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_C5wn84H-1.webp\" alt=\"Two-Tower VLM\" class=\"wp-image-217398\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_C5wn84H-1.webp 457w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_C5wn84H-1-300x211.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_C5wn84H-1-150x106.webp 150w\" sizes=\"auto, (max-width: 457px) 100vw, 457px\"\/><figcaption class=\"wp-element-caption\">Two-Tower VLM<\/figcaption><\/figure>\n<\/div>\n<ul class=\"wp-block-list\">\n<li><b>Image Encoder:<\/b> On the left side, there is an encoder that processes image data. This encoder extracts meaningful features from the image for further processing.<\/li>\n<li><b>Text Encoder:<\/b> On the right side, a similar encoder that encodes text data. It transforms the textual data into a format suitable for the shared objective.<\/li>\n<li><b>Objective<\/b>: Representation of the image and text encoders feed into a shared objective. Here the goal is to align the information from both modalities (image and text).<\/li>\n<\/ul>\n<p>This setup is common in models that aim to learn relationships between images and text. These models also work as the base for multiple downstream tasks like image captioning or visual question answering.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-two-leg-vlm\">2. Two-Leg VLM<\/h3>\n<p>The architecture described below resembles the two-tower approach, but it incorporates a fusion layer (a dense neural network) to merge the features from images and text. Let\u2019s go through each step in detail.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"468\" height=\"375\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_xqjvSSU.webp\" alt=\"Two-Leg VLM\" class=\"wp-image-217396\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_xqjvSSU.webp 468w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_xqjvSSU-300x240.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_xqjvSSU-150x120.webp 150w\" sizes=\"auto, (max-width: 468px) 100vw, 468px\"\/><figcaption class=\"wp-element-caption\">Two-Leg VLM<\/figcaption><\/figure>\n<\/div>\n<ul class=\"wp-block-list\">\n<li><b>Image Encoder:<\/b> This component processes input images. It extracts important features and representations from the image data.<\/li>\n<li><b>Text Encoder:<\/b> The right side component processes textual data. It transforms the text data into meaningful representations.<\/li>\n<li><b>Fusion Layer:<\/b> The key addition in this image is the fusion layer. After the image and text data are encoded separately, their representations are combined or fused in this layer. This is critical for learning relationships between the two modalities (images and text).<\/li>\n<li><b>Objective:<\/b> Ultimately, the fused data is utilized for a shared objective, which could be a downstream task such as classification, caption generation, or question answering.<\/li>\n<\/ul>\n<p>In summary, the image describes a multimodal system where image and text data are encoded separately and then combined at the fusion layer to achieve a unified goal. The fusion layer is crucial for leveraging the information from both data types in a coordinated way.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-3-vlm-with-image-encoder-text-encoder-amp-decoder\">3. VLM with Image Encoder \u2013 Text Encoder &amp; Decoder<\/h3>\n<p>The next architecture we can think of is an encoder for images and splitting the encoder and decoder for textual data. We divided the text into two parts where one part will pass through the encoder, and<br \/>the remaining text data will feed into the decoder and learn further relations during cross-attention. This can be one use case of question-answering from images and their long description combined. Therefore, the image will pass through the encoder, the image description will go through the text encoder, and question-answers will feed into the decoder.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1092\" height=\"417\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_6hpSGO3.webp\" alt=\"VLM with Image Encoder - Text Encoder &amp; Decoder\" class=\"wp-image-217395\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_6hpSGO3.webp 1092w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_6hpSGO3-300x115.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_6hpSGO3-768x293.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_6hpSGO3-150x57.webp 150w\" sizes=\"auto, (max-width: 1092px) 100vw, 1092px\"\/><figcaption class=\"wp-element-caption\">VLM with Image Encoder \u2013 Text Encoder &amp; Decoder<\/figcaption><\/figure>\n<\/div>\n<p>Here is an explanation of the different components:<\/p>\n<ol class=\"wp-block-list\">\n<li><b>Conv Stage<\/b>: This step processes images through a convolutional layer to extract features from the image data.<\/li>\n<li><b>Text Embedding<\/b>: Text data (such as image descriptions) is embedded into a high-dimensional vector representation.<\/li>\n<li><b>Concatenate<\/b>: Both the processed image features and the embedded text features are combined into a unified representation.<\/li>\n<li><b>Encoder<\/b>: The concatenated features are passed through an encoder, which transforms the data into a higher-level representation.<\/li>\n<li><b>Projector<\/b>: After encoding, the features are projected into a space where they can be more easily integrated with features from the decoder.<\/li>\n<li><b>Cross Attention<\/b>: This block enables interaction between the features from the projector and the decoder. In this case, the system learns which parts of the image and text data are most relevant to each other.<\/li>\n<li><b>Concatenate <\/b>Features: Instead of using cross-attention, we can stack features from the projector and decoder together.<\/li>\n<li><b>Decoder<\/b>: The combined features are passed to a decoder, which processes the integrated information and generates output.<\/li>\n<li><b>Objective<\/b>. The objective could be the same as given above.<\/li>\n<\/ol>\n<p>Overall, this diagram represents a system where images and text are processed together. Their features are concatenated or cross-attended, and finally decoded to achieve a specific objective in a multimodal task.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-4-vlm-with-encoder-decoder\">4. VLM with Encoder-Decoder<\/h3>\n<p>Our final architecture talks about an approach where all the images will be passed to encoders whereas text data will go to the decoder. During combined representation learning, we can use either<br \/>cross-attention or simply concatenate the features from both modalities.<\/p>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"682\" height=\"388\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_PatqRF6.webp\" alt=\"VLM with Image Encoder - Text Encoder - Decoder\" class=\"wp-image-217394\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_PatqRF6.webp 682w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_PatqRF6-300x171.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/01\/image_PatqRF6-150x85.webp 150w\" sizes=\"auto, (max-width: 682px) 100vw, 682px\"\/><figcaption class=\"wp-element-caption\">VLM with Image Encoder \u2013 Text Encoder \u2013 Decoder<\/figcaption><\/figure>\n<\/div>\n<p>Following is a step-by-step explanation:<\/p>\n<ul class=\"wp-block-list\">\n<li>Image Encoder: It extracts visual features from the image, transforming it into a numerical representation that the model can understand.<\/li>\n<li>Projector: The projector takes the output from the Image Encoder and projects it into a vector space compatible with the text data.<\/li>\n<li>Cross Attention: This is where the core interaction between the image and text happens. It helps the model align the visual information with the relevant textual context.<\/li>\n<li>Concatenate Features: At the place of using cross attention, we can merely stack the features of both modalities for better comprehensive context contextual learning.<\/li>\n<li>Text Decoder: It takes the concatenated features as input and uses them to predict the next word in the sequence.<\/li>\n<\/ul>\n<p>The model learns to \u201cview\u201d the images, \u201ccomprehend\u201d the text, and then generate a coherent and informative output by aligning the visual and textual information.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>Multimodal LLMs, or <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/07\/vision-language-models\/\" target=\"_blank\" rel=\"noreferrer noopener\">Vision-Language Models (VLMs)<\/a> as discussed in this blog, are trained on image-text datasets to facilitate efficient communication across different data modalities. These models excel at recognizing pixels and addressing visual tasks such as object detection and semantic segmentation. However, it is important to highlight that achieving competitive performance with VLMs demands large datasets and significant computational resources. For instance, Qwen2-VL was trained on 1.4 trillion image and text tokens.<\/p>\n<p>While VLMs can handle various visual tasks, they still show limitations in use cases such as reasoning, image interpretation, and extracting complex data.<\/p>\n<p>I will conclude the first part here, hoping it has provided a clear overview of how vision models are generally trained. It is important to note that developing these models requires a strong understanding of matrix operations, model parallelism, flash attention, and hyperparameter tuning. In the next part, we will explore training our VLMs for a small use case.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-references\">References<\/h4>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/ram-ram00706158077\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_efZ0zC0.webp\" width=\"48\" height=\"48\" alt=\"Ram Singh\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>I am Ram, a data scientist. I work as an Associate Director of Machine Learning at Cleareye.AI. Throughout my career, I have worked on various AI projects, ranging from traditional algorithms to cutting-edge technologies. I have extensive experience with LLMs and Graph Neural Networks. I am always eager to learn, and my next pursuit involves exploring Quantum computing.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>The human mind naturally perceives language, vision, smell, and touch, enabling us to understand our surroundings. We are particularly inclined toward linguistic thought and visual memory. As GenAI models continue to grow, researchers are now working on extending their capabilities by incorporating multimodality. Large Language models (LLMs) only accept text as input and produce text [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":50168,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[6967,18306,20383,3341],"dealstore":[],"offerexpiration":[],"class_list":["post-50167","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-journey","tag-llms","tag-multimodal","tag-part"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>A Journey into Multimodal LLMs Part 1 - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=50167\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"A Journey into Multimodal LLMs Part 1 - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"The human mind naturally perceives language, vision, smell, and touch, enabling us to understand our surroundings. We are particularly inclined toward linguistic thought and visual memory. As GenAI models continue to grow, researchers are now working on extending their capabilities by incorporating multimodality. Large Language models (LLMs) only accept text as input and produce text [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=50167\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-01-27T05:04:03+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"2560\" \/>\n\t<meta property=\"og:image:height\" content=\"1439\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=50167#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=50167\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"A Journey into Multimodal LLMs Part 1\",\"datePublished\":\"2025-01-27T05:04:03+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=50167\"},\"wordCount\":1959,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=50167#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp\",\"keywords\":[\"Journey\",\"LLMs\",\"Multimodal\",\"PART\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=50167#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=50167\",\"url\":\"https:\/\/fivemor.com\/?p=50167\",\"name\":\"A Journey into Multimodal LLMs Part 1 - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=50167#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=50167#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp\",\"datePublished\":\"2025-01-27T05:04:03+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=50167#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=50167\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=50167#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp\",\"width\":2560,\"height\":1439},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=50167#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"A Journey into Multimodal LLMs Part 1\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"A Journey into Multimodal LLMs Part 1 - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=50167","og_locale":"en_US","og_type":"article","og_title":"A Journey into Multimodal LLMs Part 1 - Som2ny Network","og_description":"The human mind naturally perceives language, vision, smell, and touch, enabling us to understand our surroundings. We are particularly inclined toward linguistic thought and visual memory. As GenAI models continue to grow, researchers are now working on extending their capabilities by incorporating multimodality. Large Language models (LLMs) only accept text as input and produce text [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=50167","og_site_name":"Som2ny Network","article_published_time":"2025-01-27T05:04:03+00:00","og_image":[{"width":2560,"height":1439,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=50167#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=50167"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"A Journey into Multimodal LLMs Part 1","datePublished":"2025-01-27T05:04:03+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=50167"},"wordCount":1959,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=50167#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp","keywords":["Journey","LLMs","Multimodal","PART"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=50167#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=50167","url":"https:\/\/fivemor.com\/?p=50167","name":"A Journey into Multimodal LLMs Part 1 - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=50167#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=50167#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp","datePublished":"2025-01-27T05:04:03+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=50167#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=50167"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=50167#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/What-Future-Awaits-with-Multimodal-AI_-scaled-1-2.webp.webp","width":2560,"height":1439},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=50167#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"A Journey into Multimodal LLMs Part 1"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/50167","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=50167"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/50167\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/50168"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=50167"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=50167"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=50167"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=50167"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=50167"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}