{"id":299618,"date":"2025-06-17T23:21:09","date_gmt":"2025-06-17T23:21:09","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/how-does-a-multimodal-llm-work-the-vision-story\/"},"modified":"2025-06-17T23:21:09","modified_gmt":"2025-06-17T23:21:09","slug":"how-does-a-multimodal-llm-work-the-vision-story","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=299618","title":{"rendered":"How Does A Multimodal LLM Work? The Vision Story"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p><span style=\"font-weight: 400;\">Multimodal Large Language Models (MLLMs) have lately become the talk of the AI universe. It is dynamically reshaping how AI systems understand and interact with our complex, multi-sensory world. These multi-sensory inputs that we get can also be coined as our different modalities (images, audio, etc.). From Google\u2019s latest Veo 3, generating state-of-the-art videos to ElevenLabs creating incredibly realistic AI voice overs, these systems are demonstrating capabilities that were once considered to be science fiction.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">This comprehensive guide is the first part of a two-part series exploring the intricate world of multimodal LLMs. The second part of this series will explore how these models generate multimodal content and their practical applications across various industries.<\/span><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-challenges-of-multimodality\"><span style=\"font-weight: 400;\">Challenges of Multimodality<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">Multimodality is definitely one of the greatest capabilities and advancements in AI models. However, when we deal with several modalities, there will be certain challenges that need to be curbed.<\/span> Here are two major challenges we face in this regard:<\/p>\n<ol class=\"wp-block-list\">\n<li><b>How to represent our information?<\/b><span style=\"font-weight: 400;\"><br \/>One of the main challenges of multimodal LLMs is when it comes to representing different types of information. It is how to represent and summarize these multimodal data in a common space which we need to train our multimodal models.<\/span><\/li>\n<li><b>How do we align our different modalities?<\/b><b><br \/><\/b><span style=\"font-weight: 400;\">We have to ensure we identify direct relations between similar elements from different modalities. This is done in two ways:<\/span>\n<ol class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Explicit Alignment:<\/strong> Here, we directly find correspondences between elements of different modalities. For this, we have to train our model across various modalities like audio, text, image, etc. This <\/span>supervised<span style=\"font-weight: 400;\"> or <\/span>rule-based<span style=\"font-weight: 400;\"> alignment<\/span> <span style=\"font-weight: 400;\">is implemented using algorithms like Dynamic Time Warping (DTW), Attention with supervision, or alignment matrices.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Implicit Alignment:<\/strong> This uses internally latent alignment of modalities to better solve different problems. Allowing the model to figure it out itself. Models use techniques like <\/span>self-attention<span style=\"font-weight: 400;\">, <\/span>contrastive learning,<span style=\"font-weight: 400;\"> or <\/span>co-attention mechanisms<span style=\"font-weight: 400;\"> to learn which parts of one modality relate to another.<\/span><\/li>\n<\/ol>\n<\/li>\n<\/ol>\n<figure class=\"wp-block-image size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"1167\" height=\"511\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs1.webp\" alt=\"Multimodal LLMs\" class=\"wp-image-237775\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs1.webp 1167w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs1-300x131.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs1-768x336.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs1-150x66.webp 150w\" sizes=\"(max-width: 1167px) 100vw, 1167px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:1400\/0*GUqbzVcLjQsDDtey\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Medium<\/a><\/figcaption><\/figure>\n<p><span style=\"font-weight: 400;\">Let\u2019s understand this with a small example:<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Since we need to represent the term \u201ccat\u201d whether it\u2019s in the form of text, image, or speech as closely as possible, we should make sure other terms like \u201ddog\u201d are far from the vicinity of the term \u201ccat\u201d. These embeddings from various modalities need to be correctly aligned across the shared dimensional space.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"800\" height=\"483\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-2.webp\" alt=\"How MLLMs Work\" class=\"wp-image-237778\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-2.webp 800w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-2-300x181.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-2-768x464.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-2-200x120.webp 200w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-2-150x91.webp 150w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\"\/><figcaption class=\"wp-element-caption\">Source \u2013 <a href=\"https:\/\/media2.dev.to\/dynamic\/image\/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto\/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxz646fuz6yvep03oc149.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Media2.dev<\/a><\/figcaption><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-representation-learning\"><span style=\"font-weight: 400;\">Representation Learning<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">The solution to our first problem on \u201chow to represent information\u201d can be solved by representation learning. There are 2 types of representations-based learning through which multimodal information could be understood by these multimodal models. These are: Joint Representation and Coordinated Representation<\/span>.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-joint-representation\"><span style=\"font-weight: 400;\">Joint Representation<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">Joint representation could be defined as a single unified representation of different types of information which could be text, image, video, audio, etc. We combine the embeddings of each modality in a single embedding dimension space.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1958\" height=\"498\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-3.webp\" alt=\"Joint interpretation in MLLMs\" class=\"wp-image-237779\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-3.webp 1958w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-3-300x76.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-3-768x195.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-3-1536x391.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-3-150x38.webp 150w\" sizes=\"auto, (max-width: 1958px) 100vw, 1958px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:2000\/format:webp\/1*xIDqBztt65KhSnlADtfFJQ.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Mediu<span style=\"font-weight: 400;\">m<\/span><\/a><\/figcaption><\/figure>\n<p><span style=\"font-weight: 400;\">Here, in this approach, we will pass each modality across its respective encoders. Basically, Text will be passed through a Text Encoder (e.g. BERT) and image across an Image Encoder (e.g. VIT) likewise for other modalities.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1100\" height=\"277\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-4.webp\" alt=\"List of encoders for different types of data\" class=\"wp-image-237781\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-4.webp 1100w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-4-300x76.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-4-768x193.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-4-150x38.webp 150w\" sizes=\"auto, (max-width: 1100px) 100vw, 1100px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:1100\/format:webp\/1*bSS5EucfcsaJmAIPAMsgBQ.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Medium<\/a><\/figcaption><\/figure>\n<p><span style=\"font-weight: 400;\">We get the embeddings for each modality. Later, these embedding representations merge using a concatenation technique. Then, a projection layer or multimodal attention mechanism will assign certain importance to certain features. The resulting joint embedding will contain the complete semantics of all the input modalities.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">This <\/span>entire system is trained. The<span style=\"font-weight: 400;\"> individual modality encoders, the fusion mechanism, and the final task-specific layers are all optimized together using a <\/span>single loss function<span style=\"font-weight: 400;\">. This unified training setup allows the model to learn cross-modal correlations more effectively, especially when the modalities are strongly interdependent (e.g. image and its caption like in the COCO dataset).<\/span><\/p>\n<p><span style=\"font-weight: 400;\">These joint embeddings are particularly useful when the input modalities are closely aligned or when the available training data is limited, as shared representations help in regularizing the learning process and extracting richer, semantically meaningful features from the combined input.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Read more about the <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/04\/evolution-of-embeddings\/\" target=\"_blank\" rel=\"noreferrer noopener\">Evolution of Embeddings.<\/a><\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-coordinated-representation\"><span style=\"font-weight: 400;\">Coordinated Representation<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">Coordinated Representation learning on the other side has a completely different approach. Here, we learn independent representations alone and then coordinate (or align) them together in the fusion stage. In this approach, each modality (text, image, audio, etc.) is handled by its dedicated model, which is trained separately and may also have its loss function and objective.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"2000\" height=\"629\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-5.webp\" alt=\"coordinated representation in MLLMs\" class=\"wp-image-237784\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-5.webp 2000w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-5-300x94.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-5-768x242.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-5-1536x483.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-5-150x47.webp 150w\" sizes=\"auto, (max-width: 2000px) 100vw, 2000px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:2000\/format:webp\/1*nDvetwD21fhOTscIlqrFKg.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">Medium<\/span><\/a><\/figcaption><\/figure>\n<p><span style=\"font-weight: 400;\">Once these models are trained, their individual output embeddings are combined using a <\/span>coordinated fusion mechanism<span style=\"font-weight: 400;\"> like late fusion (simple concatenation), cross-modal attention, or statistical alignment methods such as Canonical Correlation Analysis (CCA). The coordination phase focuses on ensuring that the separate embeddings are semantically aligned with each other so that they can jointly contribute to the final prediction. Unlike joint embeddings, coordinated embeddings allow each modality to <\/span>preserve its own feature structure<span style=\"font-weight: 400;\"> without being forced into a shared representation space prematurely.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">This method is highly effective when modalities are somewhat <\/span>independent or loosely coupled<span style=\"font-weight: 400;\">, when there is <\/span>abundant modality-specific data<span style=\"font-weight: 400;\">, or when computational resources allow for more extensive pre-training. Coordinated embeddings also offer greater flexibility in model architecture and training pipelines, as each modality can be improved independently before coordination.<\/span><\/p>\n<h4 class=\"wp-block-heading\" id=\"h-explicit-vs-implicit-alignment\">Explicit vs Implicit Alignment<\/h4>\n<p><span style=\"font-weight: 400;\">Let\u2019s try to tabulate our understanding here:<\/span><\/p>\n<div class=\"table-responsive mb-3\">\n<table class=\"table table-hover table-bordered\">\n<thead\/>\n<tbody>\n<tr>\n<td><b>Feature<\/b><\/td>\n<td><b>Explicit Alignment<\/b><\/td>\n<td><b>Implicit Alignment<\/b><\/td>\n<\/tr>\n<tr>\n<td><b>Nature<\/b><\/td>\n<td><span style=\"font-weight: 400;\">Supervised \/ Annotated<\/span><\/td>\n<td><span style=\"font-weight: 400;\">Unsupervised \/ Learned during training<\/span><\/td>\n<\/tr>\n<tr>\n<td><b>Need for Labels<\/b><\/td>\n<td><span style=\"font-weight: 400;\">Requires aligned or annotated data<\/span><\/td>\n<td><span style=\"font-weight: 400;\">Does not require explicit alignments<\/span><\/td>\n<\/tr>\n<tr>\n<td><b>Approach<\/b><\/td>\n<td><span style=\"font-weight: 400;\">Manual or rule-based mapping<\/span><\/td>\n<td><span style=\"font-weight: 400;\">Learned via attention or contrastive loss<\/span><\/td>\n<\/tr>\n<tr>\n<td><b>Example Tasks<\/b><\/td>\n<td><span style=\"font-weight: 400;\">Image captioning with bounding boxes<\/span><\/td>\n<td><span style=\"font-weight: 400;\">CLIP, VQA with unsupervised attention<\/span><\/td>\n<\/tr>\n<tr>\n<td><b>Advantages<\/b><\/td>\n<td><span style=\"font-weight: 400;\">High precision, interpretable<\/span><\/td>\n<td><span style=\"font-weight: 400;\">Scalable, flexible, learn fine-grained links<\/span><\/td>\n<\/tr>\n<tr>\n<td><b>Challenges<\/b><\/td>\n<td><span style=\"font-weight: 400;\">Expensive to label, less flexible<\/span><\/td>\n<td><span style=\"font-weight: 400;\">Can be less interpretable, data-hungry<\/span><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><span style=\"font-weight: 400;\">We will now try to understand another important term that we used in the above section named \u201cfusion\u201d next.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">If you want to understand how implicit alignment can be done, read <\/span><a href=\"https:\/\/arxiv.org\/pdf\/1406.5679\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">this<\/span><\/a><span style=\"font-weight: 400;\">. In this research paper, the model embeds fragments of images (objects in the image) and fragments of sentences (typed dependency tree relations) into a common space.<\/span><\/p>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"619\" height=\"293\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-6.webp\" alt=\"how multimodal large language models work\" class=\"wp-image-237793\" style=\"width:840px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-6.webp 619w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-6-300x142.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-6-150x71.webp 150w\" sizes=\"auto, (max-width: 619px) 100vw, 619px\"\/><\/figure>\n<p>Let\u2019s dive a little deeper into this.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-the-concept-of-fusion-in-multimodal-llms\">The Concept of Fusion in Multimodal LLMs<\/h2>\n<p><span style=\"font-weight: 400;\">The cornerstone of multimodal learning lies in understanding how different types of data can be combined effectively. In other words, it serves as a way to accurately align our different modalities across a unified dimensional space. Fusion strategies determine when and how information from different modalities is integrated, fundamentally shaping the model\u2019s ability to understand complex multimodal inputs<\/span>.<\/p>\n<p>Fusion<span style=\"font-weight: 400;\"> refers to the <\/span>integration of information from multiple modalities<span style=\"font-weight: 400;\"> such as text, image, and audio into a <\/span>unified representation<span style=\"font-weight: 400;\">. It plays a critical role in enabling models to <\/span>leverage complementary information<span style=\"font-weight: 400;\"> from each modality. The goal is to combine features in such a way that the model can make more informed predictions.<\/span> <span style=\"font-weight: 400;\">It\u2019s pretty similar to the concept of fusion that we use in Deep Learning.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">There are two widely used strategies for fusion: <\/span><b>Early Fusion<\/b><span style=\"font-weight: 400;\"> and <\/span><b>Late Fusion<\/b><span style=\"font-weight: 400;\">.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1100\" height=\"240\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-7.webp\" alt=\"Early fusion and late fusion\" class=\"wp-image-237795\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-7.webp 1100w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-7-300x65.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-7-768x168.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-7-150x33.webp 150w\" sizes=\"auto, (max-width: 1100px) 100vw, 1100px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:1100\/format:webp\/0*HgiQG5I4cLCwqD0a\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Medium<\/a><\/figcaption><\/figure>\n<p>There also exists a third category \u2013 mid-fusion, about which I will explain in a while.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-early-fusion\"><span style=\"font-weight: 400;\">1. Early Fusion<\/span><\/h3>\n<p>Early Fusion<span style=\"font-weight: 400;\"> represents the simplest approach to multimodal integration, here the raw data from different modalities is combined at the input level itself before any processing occurs. In early fusion systems, data from various sources such as pixel values from images and tokenized text are concatenated or combined through simple operations at the very beginning of the processing pipeline. This approach allows for comprehensive interaction between modalities from the earliest stages of computation, enabling the model to capture subtle correlations and dependencies that might be lost in later-stage fusion approaches.\u00a0<\/span><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Process:<\/strong> Raw modalities -&gt; Feature Extraction (low-level features) -&gt; Concatenation\/Simple Combination -&gt; Joint Processing by a single model.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Pros:<\/strong> It allows the model to learn correlations and interactions between modalities from the earliest stages. It can also be conceptually simpler.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Cons:<\/strong> It can be difficult to implement effectively if modalities have vastly different structures or scales. The combined feature space can become very high-dimensional and unwieldy. It forces a \u201cone-size-fits-all\u201d processing approach early on, which might not be optimal for each modality.<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400;\">Example: Earlier attempts might involve flattening an image and concatenating it with text embeddings before feeding it into a neural network. This is less common in modern sophisticated multimodal LLMs due to their limitations.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-late-fusion\"><span style=\"font-weight: 400;\">2. Late Fusion<\/span><\/h3>\n<p>Late Fusion<span style=\"font-weight: 400;\"> takes the opposite approach, processing each modality independently through specialized networks before combining the results at the decision level. Here separate neural networks process each data type using architectures optimized for that specific modality like convolutional neural networks for images, or transformer architectures for text and VIT for images. The outputs from these specialized processors are then combined using techniques such as weighted averaging, concatenation, or more sophisticated fusion modules.\u00a0<\/span><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Process: <\/strong>Modality A -&gt; Model A -&gt; Output A; Modality B -&gt; Model B -&gt; Output B. Then, Output A and Output B are combined (using averaging, voting, a small neural network, etc.).<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Pros:<\/strong> It allows for optimal, specialized processing of each modality using models best suited for it. It is simpler to implement if you already have strong unimodal models. It\u2019s more robust in missing modalities.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Cons:<\/strong> It fails to capture low-level features between modalities because they are processed in isolation for too long. Also, the fusion happens too late to influence the feature learning within each modality stream.<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400;\">Example: An image classifier identifies objects in an image, and a text classifier analyzes a caption. A separate module then combines\/fuses these classifications to say if the caption accurately describes the image.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-3-mid-fusion\"><span style=\"font-weight: 400;\">3. Mid Fusion<\/span><\/h3>\n<p>Mid Fusion<span style=\"font-weight: 400;\"> or intermediate fusion strikes a balance between early and late approaches by integrating multimodal information at various intermediate layers of the network. This strategy enables the model to capture both low-level cross-modal interactions and high-level semantic relationships. Mid-fusion architectures often employ attention mechanisms or specialized transfer modules that allow information to flow between modality-specific processing streams at multiple points throughout the network. The Multimodal Transfer Module (MMTM) uses this approach by using squeeze and excitation operations to recalculate channel-wise features in each CNN stream based on information from multiple modalities.<\/span><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Process:<\/strong> Modality A -&gt; Partial Processing A -&gt; Features A; Modality B -&gt; Partial Processing B -&gt; Features B. Then, Features A and Features B are combined and fed into a joint multimodal processing network.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Pros: <\/strong>It allows specialized initial processing while still enabling the model to learn rich cross-modal relationships at a deeper feature level. It also offers more flexibility.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Cons:<\/strong> It can be more complex to design and train. Finding the optimal point and method of fusion can be challenging in this case.<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400;\">Example: Most modern vision-language models (like LLaVA) use this. An image encoder processes the image into a set of feature vectors, and a text encoder processes the text into token embeddings. These are then projected and combined in a way that allows a central LLM to attend to both.<\/span><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-core-encoder-architectures\"><span style=\"font-weight: 400;\">Core Encoder Architectures<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">Let\u2019s now try to get an over-the-top understanding of some widely used encoders that are utilized in the VLMS.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">If you would like to learn more about various Large Vision Language model architectures click <\/span><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/01\/multimodal-llms\/\"><span style=\"font-weight: 400;\">here<\/span><\/a><span style=\"font-weight: 400;\">.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-clip-contrastive-language-image-pre-training\"><span style=\"font-weight: 400;\">CLIP: Contrastive Language-Image Pre-training<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">CLIP represents a foundational breakthrough in multimodal learning, introducing a simple yet powerful approach to learning joint representations of images and text through contrastive pre-training. The architecture consists of two separate encoders: a <\/span>vision encoder<span style=\"font-weight: 400;\"> that processes images and a <\/span>text encoder<span style=\"font-weight: 400;\"> that processes natural language descriptions. These encoders are trained jointly using a contrastive objective that encourages the model to associate images with their corresponding textual descriptions while distinguishing them from unrelated text-image pairs.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"655\" height=\"424\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-8.webp\" alt=\"CLIP: Contrastive Language-Image Pre-training\" class=\"wp-image-237801\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-8.webp 655w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-8-300x194.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-8-150x97.webp 150w\" sizes=\"auto, (max-width: 655px) 100vw, 655px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:1310\/1*LEc2qQNO6Vumrv5lqSpuhA.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">Medium<\/span><\/a><\/figcaption><\/figure>\n<p><span style=\"font-weight: 400;\">The training process for CLIP involves presenting the model with batches (for the sake of understanding the above image let\u2019s say n=5) of n image-caption pairs, where each image is paired with its correct textual description. The model computes embeddings for all images and texts in the batch, creating two sets of n-dimensional vectors.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The contrastive loss function encourages high similarity between correct image-text pairs while penalizing high similarity between incorrect pairs. As we can visualize in the above image the diagonal weights will be maximized and the rest will be penalized. Mathematically, this is expressed as a symmetric cross-entropy loss over similarity scores, where the temperature parameter controls the sharpness of the distribution.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">CLIP\u2019s effectiveness came from its ability to learn from naturally occurring image-text pairs found on the internet (400 million scrapped information from the web), eliminating the need for manually annotated datasets. This approach enables the model to learn rich semantic relationships that generalize well to downstream tasks. The learned representations demonstrate remarkable zero-shot capabilities, allowing the model to perform image classification and retrieval tasks on categories it has never seen during training. The success of CLIP has inspired numerous follow-up works and established contrastive pre-training as a dominant methodology in multimodal learning.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Also, do consider reading about <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/06\/vision-transformers-vit-revolutionizing-computer-vision\/\" target=\"_blank\" rel=\"noreferrer noopener\">ViT here<\/a>.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-siglip-sigmoid-loss-for-improved-efficiency\"><span style=\"font-weight: 400;\">SigLIP: Sigmoid Loss for Improved Efficiency<\/span><\/h3>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/10\/googles-siglip\/\" target=\"_blank\" rel=\"noreferrer noopener\"><span style=\"font-weight: 400;\">SigLIP<\/span><\/a><span style=\"font-weight: 400;\"> represents an evolution of the CLIP architecture that addresses some of the computational limitations of the original contrastive approach. While CLIP requires computing similarities between all pairs of images and texts in a batch, SigLIP employs a pairwise sigmoid loss that operates on individual image-text pairs independently. This modification eliminates the need for a global view of all pairwise similarities within a batch, enabling more efficient scaling to larger batch sizes while maintaining or improving performance.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The sigmoid loss function used in SigLIP offers several advantages over the traditional contrastive loss. It provides a more stable training mechanism and better performance with smaller batch sizes, making the approach more accessible with limited computational resources. The pairwise nature of the loss enables more flexible training configurations and better handling of datasets with varying numbers of positive examples per sample.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">SigLIP\u2019s architecture maintains the dual-encoder structure of CLIP but incorporates architectural improvements and training optimizations that enhance both efficiency and effectiveness. The model uses separate image and text encoders to generate representations for both modalities, with the sigmoid loss encouraging similarity between matched pairs and dissimilarity between unmatched pairs. This approach has demonstrated superior performance across various image-text tasks while offering improved computational efficiency compared to traditional contrastive methods.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"702\" height=\"222\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-9.webp\" alt=\"SigLIP: Sigmoid Loss for Improved Efficiency\" class=\"wp-image-237802\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-9.webp 702w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-9-300x95.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-9-150x47.webp 150w\" sizes=\"auto, (max-width: 702px) 100vw, 702px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source: <\/span><a href=\"https:\/\/cdn.hashnode.com\/res\/hashnode\/image\/upload\/v1722583708806\/ca1896d5-9637-4746-a07b-169f42e24fc8.png?auto=compress,format&amp;format=webp\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">cdn.hashnode<\/span><\/a><\/figcaption><\/figure>\n<h3 class=\"wp-block-heading\" id=\"h-rope-rotary-position-embedding\"><span style=\"font-weight: 400;\">RoPE: Rotary Position Embedding<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">Although RoPE can\u2019t be considered as an encoder model, it definitely is an embedding strategy widely used in large language models.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Rotary Position Embedding (RoPE) represents a sophisticated approach to encoding positional information in transformer-based architectures. RoPE encodes the absolute positional information using rotation matrices while naturally including the explicit relative position dependencies in self-attention formulations. This approach provides valuable properties including flexibility to expand to any sequence length, decaying inter-token dependency with increasing relative distances, and the capability to equip linear self-attention with relative position encoding.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The mathematical foundation of RoPE involves applying rotation matrices to embedding vectors based on their positions in the sequence. This rotation-based approach ensures that the dot product between embeddings captures both content similarity and relative positional relationships. The decay property of RoPE means that tokens that are farther apart in the sequence have naturally reduced attention weights, which aligns well with many natural language and multimodal tasks where local context is typically more important than distant context.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1018\" height=\"744\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-10.webp\" alt=\"RoPE embeddings in multimodal LLMs\" class=\"wp-image-237803\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-10.webp 1018w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-10-300x219.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-10-768x561.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-10-150x110.webp 150w\" sizes=\"auto, (max-width: 1018px) 100vw, 1018px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/pbs.twimg.com\/media\/FrqjrsmXoAQhr2R.jpg\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">pbs.twing<\/span><\/a><\/figcaption><\/figure>\n<p><span style=\"font-weight: 400;\">In multimodal applications, RoPE enables models to handle variable-length sequences more effectively, which is crucial when processing multimodal data where different modalities may have different temporal or spatial characteristics. The ability to extrapolate to longer sequences than those seen during training makes RoPE particularly valuable for multimodal models that need to handle diverse input formats and lengths.<\/span><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-case-studies-in-vision-language-models\"><span style=\"font-weight: 400;\">Case Studies in Vision-Language Models<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">Now, let\u2019s see how these concepts and components come together in some open-sourced influential multimodal LLMs, particularly focusing on how they \u201csee.\u201d<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-llava-large-language-and-vision-assistant\"><span style=\"font-weight: 400;\">1. LLaVA (Large Language and Vision Assistant)<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">LLaVA\u2019s core idea is to demonstrate that a remarkably simple architecture can achieve impressive visual reasoning capabilities by efficiently connecting a pre-trained vision encoder (from CLIP) to a pre-trained Large Language Model (Vicuna) using a single, trainable linear projection layer. It leverages the strong existing capabilities of these unimodal models for multimodal understanding.<\/span><\/p>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"562\" height=\"370\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-11.webp\" alt=\"LLaVA (Large Language and Vision Assistant)\" class=\"wp-image-237806\" style=\"width:840px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-11.webp 562w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-11-300x198.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-11-150x99.webp 150w\" sizes=\"auto, (max-width: 562px) 100vw, 562px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-training-process\"><span style=\"font-weight: 400;\">Training Process<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">LLaVA utilizes pre-trained Vicuna LLM and CLIP vision encoder components. The training is a 2-stage instruction-tuning procedure:<\/span><\/p>\n<p><b>Stage 1: Visual Feature Alignment (Pre-training)<\/b><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Goal:<\/strong> Teach the projection layer to map visual features into the LLM\u2019s word embedding space.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Data:<\/strong> A subset of Conceptual Captions (CC3M), containing image-caption pairs.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Method: <\/strong>The image is fed through the (frozen) CLIP-ViT. The output visual features are passed through the (trainable) linear projection layer. These projected visual tokens are prepended to the tokenized caption. The Vicuna LLM (frozen) is then tasked with autoregressively predicting the caption. Only the linear projection layer\u2019s weights are updated.<\/span><\/li>\n<\/ul>\n<p><b>Stage 2: Instruction Fine-tuning (End-to-End)<\/b><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Goal:<\/strong> Improve the model\u2019s ability to follow instructions and engage in complex visual dialogue.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Data:<\/strong> A small, high-quality synthetically generated dataset (LLaVA-Instruct-158K) using GPT-4 to create varied questions about images, detailed descriptions, and complex reasoning tasks. This dataset includes \u2013 Multimodal conversations (58k), Detailed Text Descriptions of images (23k), and Complex reasoning\/complex visual QA (77k)<\/span>.<\/li>\n<li><span style=\"font-weight: 400;\"><strong>Method:<\/strong> Both the projection layer and the LLM weights are fine-tuned on this instruction dataset. The input to the LLM is a combination of projected image features and a textual instruction\/question.<\/span><\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-working\"><span style=\"font-weight: 400;\">Working<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">The LLaVA model processes inputs which can be text, an image, or a combination.<\/span> Here\u2019s how it works:<\/p>\n<ol class=\"wp-block-list\">\n<li><b> Text Input:<\/b><span style=\"font-weight: 400;\"> Vicuna\u2019s native tokenizer and embedding system prepares the provided text (e.g. a question) for the LLM<\/span> <span style=\"font-weight: 400;\">by tokenizing and embeddin<\/span>g it.<\/li>\n<li><b>Image Input:<\/b><span style=\"font-weight: 400;\"> The CLIP vision encoder (specifically, its Vision Transformer, ViT) extracts rich visual features from the image. These features, typically representing image patches, are a sequence of vectors.<\/span><\/li>\n<li><b>Projection<\/b><span style=\"font-weight: 400;\">: These visual feature vectors then pass through the MLP Projection Layer. This layer performs a linear transformation, projecting the visual features into the same dimensionality as Vicuna\u2019s word embeddings. This makes the visual information \u201clook like\u201d word tokens to the LLM.<\/span><\/li>\n<li><b> Combined Input to LLM: <\/b><span style=\"font-weight: 400;\">The model then combines the projected visual tokens with the text token embeddings (e.g., by prepending the visual tokens to the text tokens).<\/span><\/li>\n<li><b> LLM Processing (Fusion &amp; Reasoning):<\/b><span style=\"font-weight: 400;\"> This combined sequence is fed into the Vicuna LLM. The LLM\u2019s attention mechanisms process both types of tokens simultaneously. This is where \u201cFusion\u201d happens, allowing the model to correlate parts of the text with relevant visual tokens. The goal is to achieve Joint embedding (a shared representation space) and Implicit Alignment (connecting visual concepts to textual ones).<\/span><\/li>\n<li><b>Output Generation:<\/b><span style=\"font-weight: 400;\"> Based on the processed combined input, the LLM autoregressively generates a textual response to the query or instruction.<\/span><\/li>\n<\/ol>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"696\" height=\"376\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-12.webp\" alt=\"Multimodal generation\" class=\"wp-image-237810\" style=\"width:840px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-12.webp 696w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-12-300x162.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-12-150x81.webp 150w\" sizes=\"auto, (max-width: 696px) 100vw, 696px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-simplified-version\">Simplified Version<span style=\"font-weight: 400;\"> <\/span><\/h4>\n<p><span style=\"font-weight: 400;\">LLaVA looks at an image and creates captions for the images <span style=\"font-weight: 400;\">using CLIP (vision encoder)<\/span>. A special translator (projection layer) changes these captions into a language the Vicuna LLM understands. The Vicuna brain then reads both the translated captions and any actual text words (like your question). Finally, the Vicuna brain uses all this information to give you an answer in the text.<\/span><\/p>\n<h4 class=\"wp-block-heading\" id=\"h-encoder-decoder-architecture\"><span style=\"font-weight: 400;\">Encoder-Decoder Architecture<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">While not a traditional encoder-decoder in the sequence-to-sequence translation sense, LLaVA uses components that serve these roles:<\/span><\/p>\n<ol class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Vision Encoder:<\/strong> A pre-trained CLIP ViT-L\/14. This model takes an image and outputs visual embeddings (features).<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Language Model (acts as Decoder): <\/strong>Vicuna (an instruction-tuned Llama variant). It takes the visual embeddings (after projection) and text embeddings as input, and autoregressive generates the text output.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Connector\/Projector (The \u201cBridge\u201d): <\/strong>A single linear MLP layer. This is the key new component that translates visual features from the vision encoder\u2019s space to the LLM\u2019s input embedding space.<\/span><\/li>\n<\/ol>\n<h4 class=\"wp-block-heading\" id=\"h-strengths\"><span style=\"font-weight: 400;\">Strengths<\/span><\/h4>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Simplicity &amp; Efficiency: <\/strong>Remarkable performance for its relatively simple architecture and efficient training (especially Stage 1).<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Leverages Pre-trained Models:<\/strong> Effectively utilizes the power of strong, readily available pre-trained vision (CLIP) and language (Vicuna) models.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Cost-Effective Fine-tuning:<\/strong> The initial feature alignment stage only trains a small projection layer, making it computationally cheaper.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Instruction Following: <\/strong>The LLaVA-Instruct-158K dataset was crucial for enabling strong conversational and instruction-following abilities.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Open Source: <\/strong>Contributed significantly to open research in vision-language models.<\/span><\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-limitations\"><span style=\"font-weight: 400;\">Limitations<\/span><\/h4>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Granularity (Early Versions):<\/strong> Original LLaVA often relied on a single global feature vector or a small sequence from the image (e.g., [CLS] token features), which could limit the understanding of very fine-grained details or complex spatial relationships. (Later versions like LLaVA-1.5 improved this by using more patch features and an MLP projector).<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Hallucination:<\/strong> Can sometimes \u201challucinate\u201d objects or details not present in the image, a common issue with LLMs.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Reasoning Depth:<\/strong> While good, reasoning on very complex scenes or abstract visual concepts might be limited compared to larger, more extensively trained models.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Dataset Dependency: <\/strong>Performance is heavily influenced by the quality and nature of the instruction-tuning dataset.<\/span><\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-2-llama-3-vision-llama-3-1-vision-8b-70b\"><span style=\"font-weight: 400;\">2. Llama 3 Vision (Llama 3.1 Vision 8B \/ 70B)<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">Llama 3 Vision aims to build state-of-the-art open-source multimodal models by integrating a powerful vision encoder with the strong base of Llama 3 LLMs. The core idea is to leverage Meta\u2019s advancements in LLMs, vision models, and large-scale training methodologies to create models that can perform complex visual reasoning, understand nuanced visual details, and follow intricate instructions involving images and text.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1400\" height=\"727\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-13.webp\" alt=\"Llama 3 Vision (Llama 3.1 Vision 8B \/ 70B)\" class=\"wp-image-237811\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-13.webp 1400w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-13-300x156.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-13-768x399.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-13-150x78.webp 150w\" sizes=\"auto, (max-width: 1400px) 100vw, 1400px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:1400\/1*KmSRlJXQtWU6fj9SxhYKvw.jpeg\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">Medium<\/span><\/a><\/figcaption><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-training-process-0\"><span style=\"font-weight: 400;\">Training Process<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Llama 3 Vision models leverage pre-trained Llama 3 LLMs and powerful pre-trained vision encoders (e.g., CLIP ViT). The training strategy typically involves:<\/span><\/p>\n<p><b>Stage 1: Large-Scale Multimodal Pre-training<\/b><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Goal:<\/strong> Teach the model fundamental visual concepts and their deep alignment with language at a massive scale.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Data: <\/strong>Billions of image-text pairs from diverse sources (e.g., publicly available web data, licensed datasets). Meta has access to vast (anonymized and privacy-preserving) image-text data.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Method:<\/strong> The vision encoder, a projector module (e.g., a two-layer MLP), and the Llama 3 LLM are trained jointly. The model learns to predict text associated with images or masked portions of text\/images. This stage trains the projector and fine-tunes both the vision encoder and the LLM for multimodal understanding.<\/span><\/li>\n<\/ul>\n<p><b>Stage 2: Instruction Fine-tuning (End-to-End)<\/b><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Goal:<\/strong> Enhance the model\u2019s ability to follow diverse instructions, engage in dialogue, and perform specific multimodal tasks.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Data: <\/strong>A curated mix of high-quality multimodal instruction-following datasets. This includes Visual Question Answering (VQA), image captioning, visual reasoning, object grounding, Optical Character Recognition (OCR) in images, chart\/diagram understanding, etc.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Method: <\/strong>The entire model (or significant parts of it) is fine-tuned on these instruction datasets to improve its helpfulness, safety, and task-specific performance.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Scaling: <\/strong>Meta emphasizes scaling laws, meaning Llama 3 Vision benefits from scaling up the LLM size (e.g., 8B to 70B), the vision encoder size, and the training data volume and quality.<\/span><\/li>\n<\/ul>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"567\" height=\"342\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-14.webp\" alt=\"Llama 3 Vision (Llama 3.1 Vision 8B \/ 70B) - working\" class=\"wp-image-237812\" style=\"width:696px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-14.webp 567w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-14-300x181.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-14-200x120.webp 200w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-14-150x90.webp 150w\" sizes=\"auto, (max-width: 567px) 100vw, 567px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:4800\/format:webp\/0*EeVOSzoREJsKEbSB.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">Medium<\/span><\/a><\/figcaption><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-working-0\"><span style=\"font-weight: 400;\">Working<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Llama 3 Vision processes image and text inputs to generate textual outputs.<\/span><\/p>\n<ol class=\"wp-block-list\">\n<li><b>Text Input:<\/b><span style=\"font-weight: 400;\"> Text (e.g., questions, instructions) is tokenized using Llama 3\u2019s advanced tokenizer (e.g., 128k vocabulary) and converted into token embeddings.<\/span><\/li>\n<li><b> Image Input<\/b><span style=\"font-weight: 400;\">: The input image is preprocessed (e.g., scaled to a resolution like 448\u00d7448 for Llama 3.1 Vision). It\u2019s then fed into a powerful vision encoder (e.g., a CLIP ViT model). The vision encoder processes the image and outputs a sequence of visual embeddings, representing numerous image patches (e.g., Llama 3.1 Vision produces 144 visual tokens from a CLIP ViT-L\/14).<\/span><\/li>\n<li><b>Projection<\/b><span style=\"font-weight: 400;\">: These visual embeddings are passed through a projector module, typically a multi-layer perceptron (e.g., a two-layer MLP in Llama 3.1 Vision). The projector transforms these visual features into embeddings that are compatible with the Llama 3 LLM\u2019s input space.<\/span><\/li>\n<li><b>Combined Input to LLM<\/b><span style=\"font-weight: 400;\">: The projected visual tokens are combined with the text token embeddings. Special image tokens might be used to demarcate visual information within the sequence.<\/span><\/li>\n<li><b>LLM Processing (Fusion &amp; Reasoning)<\/b><span style=\"font-weight: 400;\">: The Llama 3 LLM processes this interleaved sequence of visual and textual tokens. Its sophisticated attention mechanisms (Grouped Query Attention for efficiency with long sequences) allow it to deeply integrate and correlate information from both modalities. This enables Joint embedding and Implicit Alignment at a very fine-grained level.<\/span><\/li>\n<li><b>Output Generation<\/b><span style=\"font-weight: 400;\">: The LLM leverages its vast pre-trained knowledge, detailed visual information, and the textual context to perform reasoning and generate a coherent and relevant textual response.<\/span><\/li>\n<\/ol>\n<h4 class=\"wp-block-heading\" id=\"h-simplified-version-0\">Simplified Version<\/h4>\n<p><span style=\"font-weight: 400;\">Llama 3 Vision uses a very sharp ViT variant model to look at an image, breaking it down into many detailed picture words(patch info). A projector makes these detailed image captions ready for the super-smart Llama 3 LLM. The Llama 3 brain reads these captions along with any text questions you ask it. Because the Llama 3 brain is so big and well-trained, it can understand complex things in the picture and give you very detailed and intelligent answers in the text.<\/span><\/p>\n<h4 class=\"wp-block-heading\" id=\"h-encoder-decoder-architecture-0\"><span style=\"font-weight: 400;\">Encoder-Decoder Architecture<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Similar to LLaVA, it\u2019s a vision encoder + projector + LLM architecture:<\/span><\/p>\n<ol class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Vision Encoder:<\/strong> A powerful, pre-trained Vision Transformer. For Llama 3.1 Vision, this is a CLIP ViT model, potentially a large variant.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Language Model (acts as a Decoder):<\/strong> The Llama 3 model (e.g., Llama 3 8B or Llama 3 70B), which is an autoregressive decoder.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Connector\/Projector:<\/strong> A learnable module, typically an MLP (e.g., a two-layer MLP for Llama 3.1 Vision) to map the sequences of visual features from the ViT output into the LLM\u2019s input embedding space.<\/span><\/li>\n<\/ol>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1037\" height=\"722\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-15.webp\" alt=\"Encoder-Decoder Architecture of Llama 3\" class=\"wp-image-237814\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-15.webp 1037w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-15-300x209.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-15-768x535.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-15-150x104.webp 150w\" sizes=\"auto, (max-width: 1037px) 100vw, 1037px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/miro.medium.com\/v2\/resize:fit:1400\/1*2xqoPM6-ltd-t_O0v0eyXg.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">Medium<\/span><\/a><\/figcaption><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-strengths-0\"><span style=\"font-weight: 400;\">Strengths<\/span><\/h4>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>State-of-the-Art Performance:<\/strong> Aims for top-tier performance on a wide range of vision-language benchmarks due to scale and advanced training.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Scale:<\/strong> Benefits from large base LLMs (Llama 3 8B, 70B), powerful vision encoders, and massive training datasets.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Strong Base LLM:<\/strong> Built upon the highly capable Llama 3 models known for excellent text generation and reasoning.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Improved Reasoning &amp; Reduced Hallucination:<\/strong> Extensive pre-training and fine-tuning on high-quality, diverse data help improve reasoning and reduce the tendency to hallucinate.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Advanced Capabilities:<\/strong> Shows strong performance in areas like OCR, understanding charts\/graphs, and fine-grained visual detail recognition.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Architectural Refinements:<\/strong> Leverages LLM advancements like Grouped Query Attention (GQA) for efficient handling of long sequences (including visual tokens).<\/span><\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-limitations-0\"><span style=\"font-weight: 400;\">Limitations<\/span><\/h4>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\"><strong>Computational Cost:<\/strong> Larger models (like 70B) require significant computational resources for training and inference.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Data Dependency &amp; Bias:<\/strong> Performance and potential biases are still dependent on the vast datasets used for training. Ensuring fairness and mitigating harmful biases is an ongoing challenge.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Hallucination:<\/strong> While reduced, the risk of generating plausible but incorrect information (hallucination) persists, especially for out-of-distribution or highly ambiguous inputs.<\/span><\/li>\n<li><span style=\"font-weight: 400;\"><strong>Complexity: <\/strong>The increased scale and complexity can make debugging, interpretation, and fine-tuning more challenging for end-users compared to simpler models.<\/span><\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-advancements-in-llama-4\"><span style=\"font-weight: 400;\">Advancements in Llama 4<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">While specific, verified details for Llama 4 are still emerging, discussions around its advancements often center on tackling the inherent challenges of large-scale multimodal learning, particularly through architectural innovations like <\/span><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/12\/mixture-of-experts-models\/\" target=\"_blank\" rel=\"noreferrer noopener\">Mixture-of-Experts<\/a> (MoE)<span style=\"font-weight: 400;\">.<\/span><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1920\" height=\"1308\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-16.webp\" alt=\"Llama 4 Multimodal LLM\" class=\"wp-image-237815\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-16.webp 1920w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-16-300x204.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-16-768x523.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-16-1536x1046.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/MLLMs-16-150x102.webp 150w\" sizes=\"auto, (max-width: 1920px) 100vw, 1920px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source \u2013 <\/span><a href=\"https:\/\/scontent.fdel1-4.fna.fbcdn.net\/v\/t39.2365-6\/488655517_650996354186993_1043942188415715102_n.png\" target=\"_blank\" rel=\"noreferrer noopener nofollow\"><span style=\"font-weight: 400;\">scontent<\/span><\/a><\/figcaption><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-1-addressing-computational-complexity-and-scalability-with-moe\">1. Addressing Computational Complexity and Scalability with MoE<\/h4>\n<p><span style=\"font-weight: 400;\">A key conceptual advancement for Llama 4 is the effective implementation of MoE. This architecture significantly mitigates computational costs by activating only a relevant expert. This allows for enhancing model capacity while keeping the computational load for training and inference manageable.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Such efficiency is crucial for handling increasingly large, high-resolution multimodal datasets and long sequence lengths, which would otherwise be bottlenecked by the quadratic scaling of traditional attention mechanisms. This also enables broader scalability solutions, allowing the model to learn from more extensive and diverse data.<\/span><\/p>\n<h4 class=\"wp-block-heading\" id=\"h-2-improved-alignment-of-heterogeneous-data\">2. Improved Alignment of Heterogeneous Data<\/h4>\n<p><span style=\"font-weight: 400;\">With the capacity afforded by MoE and advancements in training strategies, Llama 4 would aim for a more sophisticated alignment of diverse modalities like images and text. This involves developing more robust representations that can capture modality-specific characteristics (e.g., spatial correlations in vision, semantic rules in text) while enabling deeper cross-modal understanding and interaction.<\/span><\/p>\n<p>Llama4 architecture also mentions the use of the <span style=\"font-weight: 400;\">Early Fusion mechanism to align the embeddings into a unified representation space. While not its primary purpose, the increased capacity and specialization within an MoE framework could indirectly aid in better handling statistical and even temporal discrepancies between modalities if trained on appropriate data.<\/span><\/p>\n<h4 class=\"wp-block-heading\" id=\"h-3-enhanced-robustness-and-bias-mitigation\">3. Enhanced Robustness and Bias Mitigation<\/h4>\n<p><span style=\"font-weight: 400;\">Models like Llama 4 are expected to incorporate more advanced strategies to address inherited biases and improve overall robustness. Llama 4 would aim to:<\/span><\/p>\n<ul class=\"wp-block-list\">\n<li><span style=\"font-weight: 400;\">Implement more comprehensive bias mitigation techniques during pre-training and fine-tuning to reduce the amplification of biases through cross-modal interactions.<\/span><\/li>\n<li><span style=\"font-weight: 400;\">Build greater resilience to input quality variations, out-of-distribution data, and adversarial attacks that might exploit cross-modal vulnerabilities. The goal is to achieve more reliable and secure performance across a wider range of real-world scenarios.<\/span><\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\"><span style=\"font-weight: 400;\">Conclusion<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">The evolution of multimodal LLMs represents one of the most significant advances in artificial intelligence, fundamentally changing how machines perceive and interact with the world around us. From the foundational concepts of early and late fusion to the sophisticated architectures of modern systems like Llama 4, we have traced the technical journey that has enabled AI systems to understand and process multiple modalities with human-like sophistication. The technical foundations we explored including contrastive learning principles, joint embedding spaces, and alignment mechanisms provide the theoretical framework that makes multimodal understanding possible.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Our case studies of LLaVA, Llama 3.2 Vision, and Llama 4 illustrate the rapid progression of multimodal capabilities. LLaVA demonstrated that elegant simplicity could achieve remarkable results through visual instruction tuning. Llama 3.2 Vision showed how sophisticated cross-attention mechanisms could enable robust multimodal reasoning. Llama 4 represents the current state-of-the-art, introducing mixture-of-experts architectures and unprecedented context lengths that open entirely new categories of applications. In the second part of this series, we will explore how these Multimodal LLMs are able to understand audio.<\/span><\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/shaik8558834\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_An81zCg.webp\" width=\"48\" height=\"48\" alt=\"Shaik Hamzah\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>GenAI Intern @ Analytics Vidhya | Final Year @ VIT Chennai<br \/>Passionate about AI and machine learning, I&#8217;m eager to dive into roles as an AI\/ML Engineer or Data Scientist where I can make a real impact. With a knack for quick learning and a love for teamwork, I&#8217;m excited to bring innovative solutions and cutting-edge advancements to the table. My curiosity drives me to explore AI across various fields and take the initiative to delve into data engineering, ensuring I stay ahead and deliver impactful projects.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n<p>                    <!-- Free Courses --><\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Multimodal Large Language Models (MLLMs) have lately become the talk of the AI universe. It is dynamically reshaping how AI systems understand and interact with our complex, multi-sensory world. These multi-sensory inputs that we get can also be coined as our different modalities (images, audio, etc.). From Google\u2019s latest Veo 3, generating state-of-the-art videos to [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":299619,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[30122,20383,3972,8767,418],"dealstore":[],"offerexpiration":[],"class_list":["post-299618","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-llm","tag-multimodal","tag-story","tag-vision","tag-work"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>How Does A Multimodal LLM Work? The Vision Story - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=299618\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"How Does A Multimodal LLM Work? The Vision Story - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Multimodal Large Language Models (MLLMs) have lately become the talk of the AI universe. It is dynamically reshaping how AI systems understand and interact with our complex, multi-sensory world. These multi-sensory inputs that we get can also be coined as our different modalities (images, audio, etc.). From Google\u2019s latest Veo 3, generating state-of-the-art videos to [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=299618\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-06-17T23:21:09+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"25 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=299618#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=299618\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"How Does A Multimodal LLM Work? The Vision Story\",\"datePublished\":\"2025-06-17T23:21:09+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=299618\"},\"wordCount\":4957,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=299618#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp\",\"keywords\":[\"LLM\",\"Multimodal\",\"story\",\"vision\",\"Work\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=299618#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=299618\",\"url\":\"https:\/\/fivemor.com\/?p=299618\",\"name\":\"How Does A Multimodal LLM Work? The Vision Story - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=299618#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=299618#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp\",\"datePublished\":\"2025-06-17T23:21:09+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=299618#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=299618\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=299618#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=299618#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"How Does A Multimodal LLM Work? The Vision Story\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"How Does A Multimodal LLM Work? The Vision Story - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=299618","og_locale":"en_US","og_type":"article","og_title":"How Does A Multimodal LLM Work? The Vision Story - Som2ny Network","og_description":"Multimodal Large Language Models (MLLMs) have lately become the talk of the AI universe. It is dynamically reshaping how AI systems understand and interact with our complex, multi-sensory world. These multi-sensory inputs that we get can also be coined as our different modalities (images, audio, etc.). From Google\u2019s latest Veo 3, generating state-of-the-art videos to [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=299618","og_site_name":"Som2ny Network","article_published_time":"2025-06-17T23:21:09+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"25 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=299618#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=299618"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"How Does A Multimodal LLM Work? The Vision Story","datePublished":"2025-06-17T23:21:09+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=299618"},"wordCount":4957,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=299618#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp","keywords":["LLM","Multimodal","story","vision","Work"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=299618#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=299618","url":"https:\/\/fivemor.com\/?p=299618","name":"How Does A Multimodal LLM Work? The Vision Story - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=299618#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=299618#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp","datePublished":"2025-06-17T23:21:09+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=299618#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=299618"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=299618#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/How-Multimodal-LLMs-Work-The-Vision-Story.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=299618#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"How Does A Multimodal LLM Work? The Vision Story"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/299618","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=299618"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/299618\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/299619"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=299618"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=299618"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=299618"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=299618"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=299618"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}