{"id":111732,"date":"2025-02-26T18:35:06","date_gmt":"2025-02-26T18:35:06","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/boosting-image-search-capabilities-using-siglip-2\/"},"modified":"2025-02-26T18:35:06","modified_gmt":"2025-02-26T18:35:06","slug":"boosting-image-search-capabilities-using-siglip-2","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=111732","title":{"rendered":"Boosting Image Search Capabilities Using SigLIP 2"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Boosting image search capabilities has become a critical focus in the realm of digital asset management, e-commerce, and social media platforms. With the ever-increasing volume of visual content generated daily, the need for efficient and accurate image retrieval systems is more pressing than ever. Enter SigLIP 2 (<a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/01\/why-is-sigmoid-function-important-in-artificial-neural-networks\/\" target=\"_blank\" rel=\"noreferrer noopener\">Sigmoid Loss<\/a> for Language-Image Pre-Training), a state-of-the-art multilingual vision-language encoder developed by <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/02\/deepminds-alphageometry2-surpasses-math-olympiad\/\" target=\"_blank\" rel=\"noreferrer noopener\">Google DeepMind<\/a>, which promises to revolutionize how we approach image similarity and search tasks. Its innovative architecture not only improves semantic understanding but also excels in zero-shot classification and image-text retrieval. By utilizing a unified training approach that incorporates self-supervised learning and diverse data curation, SigLIP 2 outperforms previous models in extracting meaningful visual representations.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-learning-objectives\">Learning Objectives<\/h4>\n<ul class=\"wp-block-list\">\n<li>Understand the fundamentals of CLIP models and their role in image retrieval systems.<\/li>\n<li>Identify the limitations of softmax-based loss functions in distinguishing nuanced image differences.<\/li>\n<li>Explore how the SigLIP model overcomes these limitations by utilizing sigmoid loss functions.<\/li>\n<li>Analyze the key advancements and differentiating features of SigLIP 2 over SigLIP.<\/li>\n<li>Implement an image retrieval system based on a user\u2019s image query.<\/li>\n<li>Compare and evaluate the performance of SigLIP 2 against SigLIP in image retrieval tasks.<\/li>\n<\/ul>\n<p><em><strong>This article was published as a part of the\u00a0<\/strong><\/em><a href=\"https:\/\/www.analyticsvidhya.com\/datahack\/blogathon\" target=\"_blank\" rel=\"noreferrer noopener\"><em><strong>Data Science Blogathon.<\/strong><\/em><\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-contrastive-language-image-pre-training-clip\">Contrastive Language-Image Pre-training (CLIP)<\/h2>\n<p>CLIP, which stands for\u00a0Contrastive Language-Image Pre-training, is a groundbreaking multimodal model developed by OpenAI in 2021. It bridges the gap between computer vision and natural language processing by learning a shared representation space for images and text. This innovative approach allows CLIP to understand and correlate both modalities simultaneously, enabling it to perform tasks like zero-shot image classification, image-text retrieval, and captioning.<\/p>\n<p><em>Learn More: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/09\/clip-vit-l14\/\" target=\"_blank\" rel=\"noreferrer noopener\">CLIP VIT-L14: OpenAI\u2019s Multimodal Marvel for Zero-Shot Image Classification<\/a><\/em><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-components-of-clip\">Key Components of CLIP<\/h3>\n<p>The Key components of CLIP consists of a Text Encoder, an Image Encoder with a Contrastive Learning Mechanism. This mechanism aligns the representations of text and images by maximizing the similarity between matching pairs and minimizing it for non-matching pairs.<\/p>\n<figure class=\"wp-block-image size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"807\" height=\"561\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_8dpX9mH.webp\" alt=\"CLIP architecture\" class=\"wp-image-223426\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_8dpX9mH.webp 807w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_8dpX9mH-300x209.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_8dpX9mH-768x534.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_8dpX9mH-150x104.webp 150w\" sizes=\"(max-width: 807px) 100vw, 807px\"\/><figcaption class=\"wp-element-caption\">Source: <a href=\"https:\/\/openai.com\/index\/clip\/\" target=\"_blank\" rel=\"nofollow noopener\">https:\/\/openai.com\/index\/clip\/<\/a><\/figcaption><\/figure>\n<p>CLIP is trained on a large dataset of image-text pairs, typically involving hundreds of millions of examples. The model learns to predict the most relevant text snippet given an image and vice versa.<\/p>\n<p><em>Also Read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/10\/googles-siglip\/\" target=\"_blank\" rel=\"noreferrer noopener\">Google\u2019s SigLIP: A Significant Momentum in CLIP\u2019s Framework<\/a><\/em><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-softmax-function-with-cross-entropy-loss\">Softmax Function with Cross Entropy Loss<\/h3>\n<p>In CLIP, there is an encoder for image and another encoder for text which take the input images and texts to a latent representation. When we have the embeddings (the latent representations) from the encoders, a similarity score (or dot product) is calculated between each image and text pair. The similarity score gives us a measure of how similar the image and the text embeddings are. To train the models to tag the correct text for an image or vice versa, a loss function is utilized whose objective is to maximize the similarity score between the image and text pairs.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"801\" height=\"454\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aNUYlk5.webp\" alt=\"How CLIP works\" class=\"wp-image-223427\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aNUYlk5.webp 801w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aNUYlk5-300x170.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aNUYlk5-768x435.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aNUYlk5-150x85.webp 150w\" sizes=\"auto, (max-width: 801px) 100vw, 801px\"\/><\/figure>\n<p>In CLIP, the <b>softmax function<\/b> is applied to the model\u2019s outputs to obtain a probability distribution like below for every image text pair in a batch.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"655\" height=\"174\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_soQDk0q.webp\" alt=\"softmax function\" class=\"wp-image-223429\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_soQDk0q.webp 655w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_soQDk0q-300x80.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_soQDk0q-150x40.webp 150w\" sizes=\"auto, (max-width: 655px) 100vw, 655px\"\/><\/figure>\n<p>In CLIP, the normalization (as seen in the denominators) is independently performed two times: across images and across texts as shown below in the loss function below \u2013<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"505\" height=\"153\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_GWb6vfr.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223430\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_GWb6vfr.webp 505w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_GWb6vfr-300x91.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_GWb6vfr-150x45.webp 150w\" sizes=\"auto, (max-width: 505px) 100vw, 505px\"\/><\/figure>\n<p>The first term in the above equation finds the best text match for a given query image while the second term finds the best image match for a given query text. \u201cB\u201d is the batch size.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-limitations-of-clip\">Limitations of CLIP<\/h3>\n<ul class=\"wp-block-list\">\n<li><b>Issues in dealing with very Similar Pairs.<\/b> While CLIP leverages the Softmax function to calculate probabilities for image-text pairings, a potential issue arises when using it directly with cosine similarity, as\u00a0<b><i>the Softmax function might not effectively capture the relative distance between image and text embeddings, especially when dealing with very similar pairs, leading to less nuanced comparisons and potentially hindering performance in certain scenarios where fine-grained distinctions are important.\u00a0<\/i><\/b> Softmax tends to push the probabilities of \u201cincorrect\u201d pairings very close to zero, potentially causing the <b>model to miss subtle differences between similar images and text descriptions.<\/b><\/li>\n<li><b>Quadratic Memory Complexity.<\/b> Additionally since in CLIP, the similarity of every positive-pair is normalized by all negative pairs, every GPU has to maintain an\u00a0NxN\u00a0matrix for all pairwise similarities introducing quadratic memory complexity.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-siglip-with-sigmoid-loss-function\">SigLIP with Sigmoid Loss Function<\/h2>\n<p>SigLIP, developed by Google follows a similar framework as CLIP but overcomes CLIP\u2019s above issues by using a<b> sigmoid-based loss <\/b>(in place of softmax based loss)<b>\u00a0<\/b>that operates independently on each image-text pair. Following is the Sigmoid Loss Function used in SigLIP<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"795\" height=\"168\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_Dcu9Sgw.webp\" alt=\"sigmoid loss function\" class=\"wp-image-223431\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_Dcu9Sgw.webp 795w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_Dcu9Sgw-300x63.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_Dcu9Sgw-768x162.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_Dcu9Sgw-150x32.webp 150w\" sizes=\"auto, (max-width: 795px) 100vw, 795px\"\/><figcaption class=\"wp-element-caption\">Source: <a href=\"https:\/\/ahmdtaha.medium.com\/sigmoid-loss-for-language-image-pre-training-2dd5e7d1af84\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">https:\/\/ahmdtaha.medium.com\/sigmoid-loss-for-language-image-pre-training-2dd5e7d1af84<\/a><\/figcaption><\/figure>\n<ul class=\"wp-block-list\">\n<li>Here, <b>\u201cN\u201d <\/b>is the batch size which is present in the denominator so that the loss remains normalized for all batch sizes.<\/li>\n<li><b>\u201c\u03a3(i=1 to N) \u03a3(j=1 to N)\u201d<\/b> is used to sum over the loss for all combinations of image (i) and text (j) pairs.<\/li>\n<li><b>\u201cz_ij\u201d <\/b>is for determining whether the image text pair is positive (1) or negative (-1).<\/li>\n<li><b>\u201ct\u201d <\/b>controls the steepness of the sigmoid curve.<\/li>\n<li><b>\u201cxi \u00b7 yj\u201d<\/b>\u00a0measures how similar the image embeddings and text embeddings are.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-differences-with-respect-to-clip\">Differences with Respect to CLIP<\/h3>\n<div class=\"table-responsive mb-3\">\n<table class=\"table table-hover table-bordered\">\n<thead\/>\n<tbody>\n<tr>\n<td><b>CLIP<\/b><\/td>\n<td><b>SigLIP\u00a0<\/b><\/td>\n<td><b>Inference<\/b><\/td>\n<\/tr>\n<tr>\n<td>Softmax Based Loss<\/td>\n<td>Sigmoid Based Loss<\/td>\n<td>SigLIP is neither asymmetric nor dependent on a global normalization factor. As a result, the loss for each pair\u2014whether positive or negative\u2014is independent of other pairs in the mini-batch<\/td>\n<\/tr>\n<tr>\n<td>Each GPU stores an NxN matrix to compute all pairwise similarities<\/td>\n<td>No need to store NXN matrix\u00a0 as each positive\/negative pair operates independently.<\/td>\n<td>Reduces computational overhead due to memory-efficient loss calculation<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2 class=\"wp-block-heading\" id=\"h-siglip-2-over-siglip\">SigLIP 2 Over SigLIP<\/h2>\n<p>SigLIP 2 models outperform the previous SigLIP versions at all model scales in key areas such as zero-shot classification, image-text retrieval, and transfer performance when extracting visual representations for Vision-Language Models (VLMs). One standout feature is the dynamic resolution (naflex) version, which is especially useful for tasks sensitive to aspect ratio and resolution.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-features-of-siglip-2\">Key Features of SigLIP 2<\/h3>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"872\" height=\"367\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_BrltdKk.webp\" alt=\"key features of SigLIP 2\" class=\"wp-image-223432\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_BrltdKk.webp 872w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_BrltdKk-300x126.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_BrltdKk-768x323.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_BrltdKk-150x63.webp 150w\" sizes=\"auto, (max-width: 872px) 100vw, 872px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-training-with-sigmoid-amp-location-aware-captioners-locca-decoder\">Training with Sigmoid &amp; Location Aware Captioners (LocCa) Decoder<\/h4>\n<p>SigLIP 2 introduces a text decoder alongside the existing image and text vision encoders during training. For LocCa, a transformer decoder with cross-attention is added to the vision encoder to achieve two key goals:<\/p>\n<ol class=\"wp-block-list\">\n<li><b>Referring Expression (REF)<\/b>: Predicting bounding box coordinates for specific locations mentioned in textual descriptions.<\/li>\n<li><b>Grounded Captioning (GCAP):<\/b> Creating captions based on specific object locations within an image.<\/li>\n<\/ol>\n<h4 class=\"wp-block-heading\" id=\"h-improved-fine-grained-local-semantics\">Improved Fine-Grained Local Semantics<\/h4>\n<p>To improve fine-grained local semantics in image representation, SigLIP 2 adds two additional objectives: Global-Local Loss and Masked Prediction Loss.<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Self-Distillation: <\/b>Unlike traditional knowledge distillation, which uses a large \u201cteacher\u201d model to train a smaller \u201cstudent\u201d model, self-distillation uses the same model for both roles. It helps transfer knowledge from deeper network layers to shallower ones or from earlier training stages to later ones.<\/li>\n<li><b>Global-Local Loss: <\/b>This loss encourages local-to-global consistency. The vision encoder (acting as the student) processes small image patches and learns to match the full-image representation created by a teacher network.<\/li>\n<li><b>Masked Prediction Loss: <\/b>This loss works by replacing 50% of the embedded image patches with mask tokens, prompting the student model to match the teacher\u2019s features at the masked locations. This helps focus on individual per-patch features rather than the full image.<\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-better-adaptability-to-different-resolutions\">Better Adaptability to Different Resolutions<\/h4>\n<p>Since image models can be highly sensitive to changes in resolution and aspect ratio, SigLIP 2 introduces two approaches for handling this:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Fixed Resolution Variant:<\/b> In this version, training resumes from a checkpoint where the model has already learned most patterns (95% of training completed). The positional embeddings are resized to match the target sequence length, and training continues with the new resolution.<\/li>\n<li><b>Dynamic Resolution (NaFlex) Variant:<\/b> The NaFlex variant builds on concepts from FlexiViT and NaViT to enable a single model to handle multiple sequence lengths and maintain the native aspect ratio of images. This reduces aspect ratio distortion and is particularly useful for tasks like OCR and document image processing.<\/li>\n<\/ul>\n<p>Now that we have covered some of the key differentiating features of SigLIP 2, let us build an image retrieval system using it in Python.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-building-an-image-retrieval-system-using-siglip-2-and-comparison-with-siglip\">Building an Image Retrieval System Using SigLIP 2 and Comparison with SigLIP<\/h2>\n<p>In the following hands on tutorial, we will build an image retrieval system when user searches based on a image query. We will compare the responses from SigLIP 2 against SigLIP as well. We will be using the T4 GPU (free tier) on Google Colab for implementing this.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-step-1-installation-of-necessary-libraries\">Step 1. Installation of Necessary Libraries<\/h4>\n<pre class=\"wp-block-code\"><code>!pip install datasets sentencepiece\n!pip install faiss-cpu\n#update latest version of transformers\n!pip install git+https:\/\/github.com\/huggingface\/transformers<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-step-2-loading-siglip-models\">Step 2. Loading SigLIP Models<\/h4>\n<pre class=\"wp-block-code\"><code>import torch\nimport faiss\nfrom torchvision import transforms\n\nfrom PIL import Image\nfrom transformers import AutoProcessor, SiglipModel, AutoImageProcessor, AutoModel, AutoTokenizer\n\nimport numpy as np\nimport requests\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else \"cpu\")\n\nmodel = SiglipModel.from_pretrained(\"google\/siglip-base-patch16-384\").to(device)\nprocessor = AutoProcessor.from_pretrained(\"google\/siglip-base-patch16-384\")\ntokenizer = AutoTokenizer.from_pretrained(\"google\/siglip-base-patch16-384\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-step-3-functions-for-processing-input-images-generating-embeddings-amp-saving-it-in-faiss\">Step 3. Functions For Processing Input Images, Generating Embeddings &amp; Saving it in FAISS<\/h4>\n<pre class=\"wp-block-code\"><code>def add_vector(embedding, index):\n    vector = embedding.detach().cpu().numpy()\n    vector = np.float32(vector)\n    faiss.normalize_L2(vector)\n    index.add(vector)\n\ndef embed_siglip(image):\n    with torch.no_grad():\n        inputs = processor(images=image, return_tensors=\"pt\").to(device)\n        image_features = model.get_image_features(**inputs)\n        return image_features<\/code><\/pre>\n<p><b>add_vector: <\/b>This function takes a tensor embedding, normalizes it, and adds it to a FAISS index for efficient similarity searching.<\/p>\n<p><b>embed_siglip:<\/b> This function takes an image, processes it, passes it through a model to obtain its embedding (feature representation), and returns these features.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-step-4-loading-image-dataset\">Step 4. Loading Image Dataset<\/h4>\n<pre class=\"wp-block-code\"><code>API_TOKEN=\"\"\nheaders = {\"Authorization\": f\"Bearer {API_TOKEN}\"}\nAPI_URL = \"https:\/\/datasets-server.huggingface.co\/rows?dataset=ceyda\/fashion-products-small&amp;config=default&amp;split=train\"\n\ndef query():\n    response = requests.get(API_URL, headers=headers)\n    return response.json()\ndata = query()<\/code><\/pre>\n<p>We load an image <a href=\"https:\/\/huggingface.co\/datasets\/ceyda\/fashion-products-small\" target=\"_blank\" rel=\"nofollow noopener\">dataset <\/a>here and fetch it using the requests library for which we pre define the Hugging Face API token first. It is a dataset on Fashion products.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-step-5-storing-the-embeddings-in-faiss-vector-database\">Step 5. Storing the Embeddings in FAISS Vector Database<\/h4>\n<pre class=\"wp-block-code\"><code>index = faiss.IndexFlatL2(768)\n\n# read the image and add vector\nfor elem in data[\"rows\"]:\n  url = elem[\"row\"][\"image\"][\"src\"]\n  image = Image.open(requests.get(url, stream=True).raw)\n  #Generate Embedding of Image\n  clip_features = embed_siglip(image)\n  #Add vector to FAISS\n  add_vector(clip_features,index)\n\n#Save the index \nfaiss.write_index(index,\".\/siglip_70k.index\")<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-step-6-querying-the-model\">Step 6. Querying the Model<\/h4>\n<pre class=\"wp-block-code\"><code>url = \"https:\/\/encrypted-tbn0.gstatic.com\/images?q=tbn:ANd9GcRsZ4PhHTilpQ5zsG51SPZVrgEhdSfQ7_cg1g&amp;s\"\nimage = Image.open(requests.get(url, stream=True).raw)\n\nwith torch.no_grad():\n  inputs = processor(images=image, return_tensors=\"pt\").to(device)\n  input_features = model.get_image_features(**inputs)\n\ninput_features = input_features.detach().cpu().numpy()\ninput_features = np.float32(input_features)\nfaiss.normalize_L2(input_features)\ndistances, indices = index.search(input_features, 3)<\/code><\/pre>\n<p>Now that we\u2019ve built the model, let\u2019s test it out with a few prompts and see how it works.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-hands-on-retrieval-testing\">Hands-on Retrieval Testing<\/h2>\n<p>Since this is a fashion dataset, we want to query on some fashion products and check if the model is able to fetch similar looking products from the database.<\/p>\n<p>We will be first querying the model with this tan colored women\u2019s bag.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"304\" height=\"298\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_u8f1AWR.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223433\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_u8f1AWR.webp 304w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_u8f1AWR-300x294.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_u8f1AWR-150x147.webp 150w\" sizes=\"auto, (max-width: 304px) 100vw, 304px\"\/><\/figure>\n<p>Let us check the 3 most similar products fetched from the model based on this query now.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-testing-on-siglip-2-model\">Testing on SigLIP 2 Model<\/h3>\n<pre class=\"wp-block-code\"><code>#DISPLAYING SIMILAR IMAGE\nfor elem in indices[0]:\n  url = data[\"rows\"][elem][\"row\"][\"image\"][\"src\"]\n  image = Image.open(requests.get(url, stream=True).raw)\n  width = 300\n  ratio = (width \/ float(image.size[0]))\n  height = int((float(image.size[1]) * float(ratio)))\n  img = image.resize((width, height), Image.Resampling.LANCZOS)\n  display(img)<\/code><\/pre>\n<p><strong>Output from SigLIP 2 Model<\/strong><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"843\" height=\"318\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_NhLtrSB.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223434\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_NhLtrSB.webp 843w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_NhLtrSB-300x113.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_NhLtrSB-768x290.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_NhLtrSB-150x57.webp 150w\" sizes=\"auto, (max-width: 843px) 100vw, 843px\"\/><\/figure>\n<p>As seen from the output of the SigLIP 2 model, all the retrieved images of bags are close to our queried bag.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-testing-on-siglip-model\">Testing on SigLIP Model<\/h3>\n<p>Let us now check the same with SigLIP model. We can simply load this model in Step 2 using the following code<\/p>\n<pre class=\"wp-block-code\"><code>import torch\nimport faiss\nfrom torchvision import transforms\n\nfrom PIL import Image\nfrom transformers import AutoProcessor, SiglipModel, AutoImageProcessor, AutoModel, AutoTokenizer\n\nimport numpy as np\nimport requests\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else \"cpu\")\n\nmodel = SiglipModel.from_pretrained(\"google\/siglip-base-patch16-384\").to(device)\nprocessor = AutoProcessor.from_pretrained(\"google\/siglip-base-patch16-384\")\ntokenizer = AutoTokenizer.from_pretrained(\"google\/siglip-base-patch16-384\")<\/code><\/pre>\n<p>The other subsequent steps can be re-run as before.<\/p>\n<p><strong>Output from SigLIP Model<\/strong><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"844\" height=\"355\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_uqlW3Tt.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223435\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_uqlW3Tt.webp 844w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_uqlW3Tt-300x126.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_uqlW3Tt-768x323.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_uqlW3Tt-150x63.webp 150w\" sizes=\"auto, (max-width: 844px) 100vw, 844px\"\/><\/figure>\n<p>As seen from the output of the SigLIP model, two of the retrieved images of bags are similar to the retrieved images of bags from SigLIP 2 model. However, the third image retrieved from SigLIP model is not close to our query image as it is not close to the tan color.<\/p>\n<p>Let us check for another query with this input image.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"404\" height=\"300\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_WJOGSnW-thumbnail_webp-600x300-1.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223436\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_WJOGSnW-thumbnail_webp-600x300-1.webp 404w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_WJOGSnW-thumbnail_webp-600x300-1-300x223.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_WJOGSnW-thumbnail_webp-600x300-1-150x111.webp 150w\" sizes=\"auto, (max-width: 404px) 100vw, 404px\"\/><\/figure>\n<p><strong>Output from SigLIP 2 model<\/strong><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"819\" height=\"234\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aEkhhC4.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223437\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aEkhhC4.webp 819w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aEkhhC4-300x86.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aEkhhC4-768x219.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_aEkhhC4-150x43.webp 150w\" sizes=\"auto, (max-width: 819px) 100vw, 819px\"\/><\/figure>\n<p>As seen from the output of the SigLIP 2 model, all the retrieved images of the womens shoes are Canvas shoes and close to our queried shoe.<\/p>\n<p><strong>Output from SigLIP Model<\/strong><\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"844\" height=\"225\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_04eWWsw.webp\" alt=\"Image Search Using SigLIP 2\" class=\"wp-image-223438\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_04eWWsw.webp 844w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_04eWWsw-300x80.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_04eWWsw-768x205.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_04eWWsw-150x40.webp 150w\" sizes=\"auto, (max-width: 844px) 100vw, 844px\"\/><\/figure>\n<p>As seen from the output of the SigLIP model, two of the retrieved images of shoes are similar to the retrieved images of shoes from SigLIP 2 model. However, the third image retrieved from SigLIP model is not exactly like our query image as it is not a Canvas shoe.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>SigLIP 2 represents a significant step forward in the evolution of image-text retrieval and vision-language models. Its advanced features, such as dynamic resolution and improved fine-grained semantic understanding, make it a powerful tool for enhancing image search capabilities across various applications. By addressing key limitations of previous models, SigLIP 2 offers more accurate and efficient image retrieval, positioning it as a valuable asset in fields like e-commerce, digital asset management, and social media.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h4>\n<ul class=\"wp-block-list\">\n<li>SigLIP 2, developed by Google DeepMind, improves upon its predecessor by utilizing a unified training approach and sigmoid-based loss, offering more accurate and efficient image-text retrieval and zero-shot classification.<\/li>\n<li>Unlike CLIP, which uses a Softmax function that can struggle with nuanced image-text comparisons, SigLIP 2 employs a more effective sigmoid loss function that works independently on each image-text pair, enhancing performance.<\/li>\n<li>SigLIP 2 introduces the NaFlex variant, allowing the model to handle varying image resolutions and aspect ratios effectively, making it ideal for tasks such as OCR and document processing.<\/li>\n<li>Through the use of self-distillation and enhanced training techniques like Global-Local Loss and Masked Prediction Loss, SigLIP 2 offers better semantic understanding, making it more adept at capturing detailed visual features.<\/li>\n<li>SigLIP 2 features a Location Aware Captioners (LocCa) Decoder, enabling tasks like grounded captioning and predicting bounding box coordinates, further enhancing its capabilities for accurate image search and retrieval.<\/li>\n<\/ul>\n<p><strong>The media shown in this article is not owned by Analytics Vidhya and is used at the Author\u2019s discretion<\/strong>.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1740492334163\"><strong class=\"schema-faq-question\">Q1. What is SigLIP 2, and how does it improve image search capabilities?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. SigLIP 2 is a state-of-the-art multilingual vision-language encoder developed by Google DeepMind. It improves image search by enhancing semantic understanding, enabling better image-text retrieval and zero-shot classification. Its unified training approach and sigmoid-based loss function offer superior performance compared to previous models.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1740588609646\"><strong class=\"schema-faq-question\">Q2. What are the main features of SigLIP 2 that make it stand out?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. SigLIP 2 introduces features like Location Aware Captioners (LocCa) Decoder for predicting bounding box coordinates and grounded captioning. It also improves fine-grained local semantics through self-distillation, Global-Local Loss, and Masked Prediction Loss, which make it more adept at handling detailed visual information.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1740588679114\"><strong class=\"schema-faq-question\">Q3. What variants does SigLIP 2 come in?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. SigLIP 2 models come in two main variants:\u00a0FixRes\u00a0and\u00a0NaFlex. FixRes works with fixed resolution images, while NaFlex supports variable image aspect ratios and resolutions.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1740588721123\"><strong class=\"schema-faq-question\">Q4. What are the key improvements in SigLIP 2 over SigLIP?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">A. SigLIP 2 models outperform their predecessors in tasks like zero-shot classification, image-text retrieval, and localization tasks. They also offer better multilingual understanding and fairness due to a more diverse training dataset.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/mimi6\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_ZkJo4gb.webp\" width=\"48\" height=\"48\" alt=\"Nibedita Dutta\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Nibedita completed her master\u2019s in Chemical Engineering from IIT Kharagpur in 2014 and is currently working as a Senior Data Scientist. In her current capacity, she works on building intelligent ML-based solutions to improve business processes.               <\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Boosting image search capabilities has become a critical focus in the realm of digital asset management, e-commerce, and social media platforms. With the ever-increasing volume of visual content generated daily, the need for efficient and accurate image retrieval systems is more pressing than ever. Enter SigLIP 2 (Sigmoid Loss for Language-Image Pre-Training), a state-of-the-art multilingual [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":111733,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[14038,13739,15867,5476,50141],"dealstore":[],"offerexpiration":[],"class_list":["post-111732","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-boosting","tag-capabilities","tag-image","tag-search","tag-siglip"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Boosting Image Search Capabilities Using SigLIP 2 - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=111732\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Boosting Image Search Capabilities Using SigLIP 2 - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Boosting image search capabilities has become a critical focus in the realm of digital asset management, e-commerce, and social media platforms. With the ever-increasing volume of visual content generated daily, the need for efficient and accurate image retrieval systems is more pressing than ever. Enter SigLIP 2 (Sigmoid Loss for Language-Image Pre-Training), a state-of-the-art multilingual [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=111732\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-26T18:35:06+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"13 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=111732#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=111732\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Boosting Image Search Capabilities Using SigLIP 2\",\"datePublished\":\"2025-02-26T18:35:06+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=111732\"},\"wordCount\":2259,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=111732#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp\",\"keywords\":[\"boosting\",\"Capabilities\",\"Image\",\"search\",\"SigLIP\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=111732#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=111732\",\"url\":\"https:\/\/fivemor.com\/?p=111732\",\"name\":\"Boosting Image Search Capabilities Using SigLIP 2 - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=111732#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=111732#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp\",\"datePublished\":\"2025-02-26T18:35:06+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=111732#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=111732\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=111732#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=111732#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Boosting Image Search Capabilities Using SigLIP 2\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Boosting Image Search Capabilities Using SigLIP 2 - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=111732","og_locale":"en_US","og_type":"article","og_title":"Boosting Image Search Capabilities Using SigLIP 2 - Som2ny Network","og_description":"Boosting image search capabilities has become a critical focus in the realm of digital asset management, e-commerce, and social media platforms. With the ever-increasing volume of visual content generated daily, the need for efficient and accurate image retrieval systems is more pressing than ever. Enter SigLIP 2 (Sigmoid Loss for Language-Image Pre-Training), a state-of-the-art multilingual [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=111732","og_site_name":"Som2ny Network","article_published_time":"2025-02-26T18:35:06+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"13 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=111732#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=111732"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Boosting Image Search Capabilities Using SigLIP 2","datePublished":"2025-02-26T18:35:06+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=111732"},"wordCount":2259,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=111732#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp","keywords":["boosting","Capabilities","Image","search","SigLIP"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=111732#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=111732","url":"https:\/\/fivemor.com\/?p=111732","name":"Boosting Image Search Capabilities Using SigLIP 2 - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=111732#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=111732#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp","datePublished":"2025-02-26T18:35:06+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=111732#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=111732"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=111732#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/Boosting-Image-Search-Capabilities-using-SigLIP2.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=111732#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Boosting Image Search Capabilities Using SigLIP 2"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/111732","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=111732"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/111732\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/111733"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=111732"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=111732"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=111732"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=111732"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=111732"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}