{"id":202266,"date":"2025-04-24T07:42:40","date_gmt":"2025-04-24T07:42:40","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/how-to-use-google-gemini-models-for-computer-vision-tasks\/"},"modified":"2025-04-24T07:42:40","modified_gmt":"2025-04-24T07:42:40","slug":"how-to-use-google-gemini-models-for-computer-vision-tasks","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=202266","title":{"rendered":"How to Use Google Gemini Models for Computer Vision Tasks?"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Since the rise of AI chatbots, Google\u2019s Gemini has emerged as one of the most powerful players driving the evolution of intelligent systems. Beyond its conversational strength, Gemini also unlocks practical possibilities in computer vision, enabling machines to see, interpret, and describe the world around them.<\/p>\n<p>This guide walks you through the steps to leverage Google Gemini for computer vision, including how to set up your environment, send images with instructions, and interpret the model\u2019s outputs for <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2022\/03\/a-basic-introduction-to-object-detection\/\" target=\"_blank\" rel=\"noopener\">object detection<\/a>, caption generation, and <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2021\/06\/optical-character-recognitionocr-with-tesseract-opencv-and-python\/\" target=\"_blank\" rel=\"noopener\">OCR<\/a>. We\u2019ll also touch on data annotation tools (like those used with YOLO) to give context for custom training scenarios.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-google-gemini\">What is Google Gemini?<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/12\/what-is-google-gemini-features-usage-and-limitations\/\" target=\"_blank\" rel=\"noopener\">Google Gemini<\/a> is a family of AI models built to handle multiple data types, such as text, images, audio, and code together. This means they can process tasks that involve understanding both pictures and words.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-gemini-2-5-pro-features\">Gemini 2.5 Pro Features<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Multimodal Input:<\/strong> It accepts combinations of text and images in a single request.<\/li>\n<li><strong>Reasoning<\/strong>: The model can analyze information from the inputs to perform tasks like identifying objects or describing scenes.<\/li>\n<li><strong>Instruction Following<\/strong>: It responds to text instructions (prompts) that guide its analysis of the image.<\/li>\n<\/ul>\n<p>These features allow developers to use Gemini for vision-related tasks through an API without training a separate model for each job.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-the-role-of-data-annotation-the-yolo-annotator\">The Role of Data Annotation: The YOLO Annotator<\/h2>\n<p>While Gemini models provide powerful zero-shot or few-shot capabilities for these computer vision tasks, building highly specialized <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/computer-vision-models\/\" target=\"_blank\" rel=\"noreferrer noopener\">computer vision models<\/a> requires training on a dataset tailored to the specific problem. This is where data annotation becomes essential, particularly for supervised learning tasks like training a custom object detector.<\/p>\n<p>The YOLO Annotator (often referring to tools compatible with the YOLO format, like Labeling, CVAT, or Roboflow) is designed to create labeled datasets.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-what-is-data-annotation\">What is Data Annotation?<\/h3>\n<p>For object detection, annotation involves drawing bounding boxes around each object of interest in an image and assigning a class label (e.g., \u2018car\u2019, \u2018person\u2019, \u2018dog\u2019). This annotated data tells the model what to look for and where during training.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-features-of-annotation-tools-like-yolo-annotator\">Key Features of Annotation Tools (like YOLO Annotator)<\/h3>\n<ol class=\"wp-block-list\">\n<li><strong>User Interface:<\/strong> They provide graphical interfaces allowing users to load images, draw boxes (or polygons, keypoints, etc.), and assign labels efficiently.<\/li>\n<li><strong>Format Compatibility:<\/strong> Tools designed for YOLO models save annotations in a specific text file format that YOLO training scripts expect (typically one .txt file per image, containing class index and normalized bounding box coordinates).<\/li>\n<li><strong>Efficiency Features:<\/strong> Many tools include features like hotkeys, automatic saving, and sometimes model-assisted labeling to speed up the often time-consuming annotation process. Batch processing allows for more effective handling of large image sets.<\/li>\n<li><strong>Integration<\/strong>: Using standard formats like YOLO ensures that the annotated data can be easily used with popular training frameworks, including Ultralytics YOLO.<\/li>\n<\/ol>\n<p>While Google Gemini for Computer Vision, can detect general objects without prior annotation, if you needed a model to detect very specific, custom objects (e.g., unique types of industrial equipment, specific product defects), you would likely need to collect images and annotate them using a tool like a YOLO annotator to train a dedicated YOLO model.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-code-implementation-google-gemini-for-computer-vision\">Code Implementation \u2013 Google Gemini for Computer Vision<\/h2>\n<p>First, you need to install the necessary software libraries.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-1-install-the-prerequisites\">Step 1: Install the Prerequisites<\/h3>\n<h4 class=\"wp-block-heading\" id=\"h-1-install-libraries\">1. Install Libraries<\/h4>\n<p>Run this command in your terminal:<\/p>\n<pre class=\"wp-block-code\"><code>!uv pip install -U -q google-genai ultralytics<\/code><\/pre>\n<p>This command installs the <em>google-genai<\/em> library to communicate with the <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/03\/how-to-access-and-use-the-gemini-api-a-complete-guide\/\" target=\"_blank\" rel=\"noreferrer noopener\">Gemini API<\/a> and the <em>ultralytics<\/em> library, which contains helpful functions for handling images and drawing on them.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-2-import-modules\">2. Import Modules<\/h4>\n<p> Add these lines to your Python Notebook:<\/p>\n<pre class=\"wp-block-code\"><code>import json\n\nimport cv2\n\nimport ultralytics\n\nfrom google import genai\n\nfrom google.genai import types\n\nfrom PIL import Image\n\nfrom ultralytics.utils.downloads import safe_download\n\nfrom ultralytics.utils.plotting import Annotator, colors\n\nultralytics.checks()<\/code><\/pre>\n<p>This code imports libraries for tasks like reading images (<em>cv2<\/em>, <em>PIL<\/em>), handling JSON data (<em>json<\/em>), interacting with the API (<em>google.generativeai<\/em>), and utility functions (<em>ultralytics<\/em>).<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-3-configure-api-key\">3. Configure API Key<\/h4>\n<p>Initialize the client using your Google AI API key.<\/p>\n<pre class=\"wp-block-code\"><code># Replace \"your_api_key\" with your actual key\n\n# Use GenerativeModel for newer versions of the library\n\n# Initialize the Gemini client with your API key\n\nclient = genai.Client(api_key=\u201dyour_api_key\u201d)<\/code><\/pre>\n<p>This step prepares your script to send authenticated requests.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-2-function-to-interact-with-gemini\">Step 2: Function to Interact with Gemini<\/h3>\n<p>Create a function to send requests to the model. This function takes an image and a text prompt and returns the model\u2019s text output.<\/p>\n<pre class=\"wp-block-code\"><code>def inference(image, prompt, temp=0.5):\n\n\u00a0\u00a0\u00a0\"\"\"\n\n\u00a0\u00a0\u00a0Performs inference using Google Gemini 2.5 Pro Experimental model.\n\n\u00a0\u00a0\u00a0Args:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0image (str or genai.types.Blob): The image input, either as a base64-encoded string or Blob object.\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0prompt (str): A text prompt to guide the model's response.\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0temp (float, optional): Sampling temperature for response randomness. Default is 0.5.\n\n\u00a0\u00a0\u00a0Returns:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0str: The text response generated by the Gemini model based on the prompt and image.\n\n\u00a0\u00a0\u00a0\"\"\"\n\n\u00a0\u00a0\u00a0response = client.models.generate_content(\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0model=\"gemini-2.5-pro-exp-03-25\",\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0contents=[prompt, image],\u00a0 # Provide both the text prompt and image as input\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0config=types.GenerateContentConfig(\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0temperature=temp,\u00a0 # Controls creativity vs. determinism in output\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0),\n\n\u00a0\u00a0\u00a0)\n\n\u00a0\u00a0\u00a0return response.text\u00a0 # Return the generated textual response<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-explanation\">Explanation<\/h4>\n<ol class=\"wp-block-list\">\n<li>This function sends the image and your text instruction (prompt) to the Gemini model specified in the model_client.<\/li>\n<li>The temperature setting (temp) influences output randomness; lower values give more predictable results.<\/li>\n<\/ol>\n<h3 class=\"wp-block-heading\" id=\"h-step-3-preparing-image-data\">Step 3: Preparing Image Data<\/h3>\n<p>You need to load images correctly before sending them to the model. This function downloads an image if needed, reads it, converts the color format, and returns a <em>PIL Image<\/em> object and its dimensions.<\/p>\n<pre class=\"wp-block-code\"><code>def read_image(filename):\n\n\u00a0\u00a0\u00a0image_name = safe_download(filename)\n\n\u00a0\u00a0\u00a0# Read image with opencv\n\n\u00a0\u00a0\u00a0image = cv2.cvtColor(cv2.imread(f\"\/content\/{image_name}\"), cv2.COLOR_BGR2RGB)\n\n\u00a0\u00a0\u00a0# Extract width and height\n\n\u00a0\u00a0\u00a0h, w = image.shape[:2]\n\n\u00a0\u00a0\u00a0# # Read the image using OpenCV and convert it into the PIL format\n\n\u00a0\u00a0\u00a0return Image.fromarray(image), w, h<\/code><\/pre>\n<h4 class=\"wp-block-heading\" id=\"h-explanation-0\">Explanation<\/h4>\n<ol class=\"wp-block-list\">\n<li>This function uses OpenCV (cv2) to read the image file.<\/li>\n<li>It converts the image color order to RGB, which is standard.<\/li>\n<li>It returns the image as a PIL object, suitable for the inference function, and its width and height.<\/li>\n<\/ol>\n<h3 class=\"wp-block-heading\" id=\"h-step-4-result-formatting\">Step 4: Result formatting<\/h3>\n<pre class=\"wp-block-code\"><code>def clean_results(results):\n\n\u00a0\u00a0\u00a0\"\"\"Clean the results for visualization.\"\"\"\n\n\u00a0\u00a0\u00a0return results.strip().removeprefix(\"```json\").removesuffix(\"```\").strip()<\/code><\/pre>\n<p>This function formats the result into JSON format.\u00a0<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-task-1-object-detection\">Task 1: Object Detection<\/h2>\n<p>Gemini can find objects in an image and report their locations (bounding boxes) based on your text instructions.<\/p>\n<pre class=\"wp-block-code\"><code># Define the text prompt\n\nprompt = \"\"\"\n\nDetect the 2d bounding boxes of objects in image.\n\n\"\"\"\n\n# Fixed, plotting function depends on this.\n\noutput_prompt = \"Return just box_2d and labels, no additional text.\"\n\nimage, w, h = read_image(\"https:\/\/media-cldnry.s-nbcnews.com\/image\/upload\/t_fit-1000w,f_auto,q_auto:best\/newscms\/2019_02\/2706861\/190107-messy-desk-stock-cs-910a.jpg\")\u00a0 # Read img, extract width, height\n\nresults = inference(image, prompt + output_prompt)\u00a0 # Perform inference\n\ncln_results = json.loads(clean_results(results))\u00a0 # Clean results, list convert\n\nannotator = Annotator(image)\u00a0 # initialize Ultralytics annotator\n\nfor idx, item in enumerate(cln_results):\n\n\u00a0\u00a0\u00a0# By default, gemini model return output with y coordinates first.\n\n\u00a0\u00a0\u00a0# Scale normalized box coordinates (0\u20131000) to image dimensions\n\n\u00a0\u00a0\u00a0y1, x1, y2, x2 = item[\"box_2d\"]\u00a0 # bbox post processing,\n\n\u00a0\u00a0\u00a0y1 = y1 \/ 1000 * h\n\n\u00a0\u00a0\u00a0x1 = x1 \/ 1000 * w\n\n\u00a0\u00a0\u00a0y2 = y2 \/ 1000 * h\n\n\u00a0\u00a0\u00a0x2 = x2 \/ 1000 * w\n\n\u00a0\u00a0\u00a0if x1 &gt; x2:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0x1, x2 = x2, x1\u00a0 # Swap x-coordinates if needed\n\n\u00a0\u00a0\u00a0if y1 &gt; y2:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0y1, y2 = y2, y1\u00a0 # Swap y-coordinates if needed\n\n\u00a0\u00a0\u00a0annotator.box_label([x1, y1, x2, y2], label=item[\"label\"], color=colors(idx, True))\n\nImage.fromarray(annotator.result())\u00a0 # display the output<\/code><\/pre>\n<p><strong>Source Image:<\/strong> <a href=\"https:\/\/dynamichr.com\/why-a-cluttered-desk-kills-your-productivity-and-how-to-fix-it\/\" target=\"_blank\" rel=\"nofollow noopener\">Link<\/a><\/p>\n<p><strong>Output<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"626\" height=\"453\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Second.webp\" alt=\"Task 1 Output 1\" class=\"wp-image-230549\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Second.webp 626w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Second-300x217.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Second-150x109.webp 150w\" sizes=\"auto, (max-width: 626px) 100vw, 626px\"\/><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-explanation-1\">Explanation<\/h4>\n<ol class=\"wp-block-list\">\n<li>The prompt tells the model what to find and how to format the output (JSON)<\/li>\n<li>It converts the normalized box coordinates (0-1000) to pixel coordinates using the image width (w) and height (h).<\/li>\n<li>The Annotator tool draws the boxes and labels on a copy of the image<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-task-2-testing-reasoning-capabilities\">Task 2: Testing Reasoning Capabilities<\/h2>\n<p>With <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/05\/gemini-models-by-google\/\" target=\"_blank\" rel=\"noreferrer noopener\">Gemini models<\/a>, you can tackle complex tasks using advanced reasoning that understands context and delivers more precise results.<\/p>\n<pre class=\"wp-block-code\"><code># Define the text prompt\n\nprompt = \"\"\"\n\nDetect the 2d bounding box around:\n\nhighlight the area of morning light +\n\nPC on table\n\npotted plant\n\ncoffee cup on table\n\n\"\"\"\n\n# Fixed, plotting function depends on this.\n\noutput_prompt = \"Return just box_2d and labels, no additional text.\"\n\nimage, w, h = read_image(\"https:\/\/thumbs.dreamstime.com\/b\/modern-office-workspace-laptop-coffee-cup-cityscape-sunrise-sleek-desk-featuring-stationery-organized-neatly-city-345762953.jpg\")\u00a0 # Read image and extract width, height\n\nresults = inference(image, prompt + output_prompt)\n\n# Clean the results and load results in list format\n\ncln_results = json.loads(clean_results(results))\n\nannotator = Annotator(image)\u00a0 # initialize Ultralytics annotator\n\nfor idx, item in enumerate(cln_results):\n\n\u00a0\u00a0\u00a0# By default, gemini model return output with y coordinates first.\n\n\u00a0\u00a0\u00a0# Scale normalized box coordinates (0\u20131000) to image dimensions\n\n\u00a0\u00a0\u00a0y1, x1, y2, x2 = item[\"box_2d\"]\u00a0 # bbox post processing,\n\n\u00a0\u00a0\u00a0y1 = y1 \/ 1000 * h\n\n\u00a0\u00a0\u00a0x1 = x1 \/ 1000 * w\n\n\u00a0\u00a0\u00a0y2 = y2 \/ 1000 * h\n\n\u00a0\u00a0\u00a0x2 = x2 \/ 1000 * w\n\n\u00a0\u00a0\u00a0if x1 &gt; x2:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0x1, x2 = x2, x1\u00a0 # Swap x-coordinates if needed\n\n\u00a0\u00a0\u00a0if y1 &gt; y2:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0y1, y2 = y2, y1\u00a0 # Swap y-coordinates if needed\n\n\u00a0\u00a0\u00a0annotator.box_label([x1, y1, x2, y2], label=item[\"label\"], color=colors(idx, True))\n\nImage.fromarray(annotator.result())\u00a0 # display the output<\/code><\/pre>\n<p><strong>Source Image:<\/strong> <a href=\"https:\/\/www.dreamstime.com\/photos-images\/sunrise-cityscape-view-modern-office-desk.html\" target=\"_blank\" rel=\"nofollow noopener\">Link<\/a><\/p>\n<p><strong>Output<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"830\" height=\"460\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Third.webp\" alt=\"Task 1 Output 2\" class=\"wp-image-230551\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Third.webp 830w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Third-300x166.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Third-768x426.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Third-150x83.webp 150w\" sizes=\"auto, (max-width: 830px) 100vw, 830px\"\/><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-explanation-2\">Explanation<\/h4>\n<ol class=\"wp-block-list\">\n<li>This code block contains a complex prompt to test the model\u2019s reasoning capabilities.<\/li>\n<li>It converts the normalized box coordinates (0-1000) to pixel coordinates using the image width (w) and height (h).<\/li>\n<li>The Annotator tool draws the boxes and labels on a copy of the image.<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-task-3-image-captioning\">Task 3: Image Captioning<\/h2>\n<p>Gemini can create text descriptions for an image.<\/p>\n<pre class=\"wp-block-code\"><code># Define the text prompt\n\nprompt = \"\"\"\n\nWhat's inside the image, generate a detailed captioning in the form of short\n\nstory, Make 4-5 lines and start each sentence on a new line.\n\n\"\"\"\n\nimage, _, _ = read_image(\"https:\/\/cdn.britannica.com\/61\/93061-050-99147DCE\/Statue-of-Liberty-Island-New-York-Bay.jpg\")\u00a0 # Read image and extract width, height\n\nplt.imshow(image)\n\nplt.axis('off')\u00a0 # Hide axes\n\nplt.show()\n\nprint(inference(image, prompt))\u00a0 # Display the results<\/code><\/pre>\n<p><strong>Source Image:<\/strong> <a href=\"https:\/\/www.britannica.com\/story\/who-was-the-woman-behind-the-statue-of-liberty\" target=\"_blank\" rel=\"nofollow noopener\">Link<\/a><\/p>\n<p><strong>Output<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"829\" height=\"514\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fourth.webp\" alt=\"Task 2 Output\" class=\"wp-image-230552\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fourth.webp 829w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fourth-300x186.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fourth-768x476.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fourth-150x93.webp 150w\" sizes=\"auto, (max-width: 829px) 100vw, 829px\"\/><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-explanation-2\">Explanation<\/h4>\n<ol class=\"wp-block-list\">\n<li>This prompt asks for a specific style of description (narrative, 4 lines, new lines).<\/li>\n<li>The provided image is shown in the output.<\/li>\n<li>The function returns the generated text. This is useful for creating alt text or summaries.<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-task-4-optical-character-recognition-ocr\">Task 4: Optical Character Recognition (OCR)<\/h2>\n<p>Gemini can read text within an image and tell you where it found the text.<\/p>\n<pre class=\"wp-block-code\"><code># Define the text prompt\n\nprompt = \"\"\"\n\nExtract the text from the image\n\n\"\"\"\n\n# Fixed, plotting function depends on this.\n\noutput_prompt = \"\"\"\n\nReturn just box_2d which will be location of detected text areas + label\"\"\"\n\nimage, w, h = read_image(\"https:\/\/cdn.mos.cms.futurecdn.net\/4sUeciYBZHaLoMa5KiYw7h-1200-80.jpg\")\u00a0 # Read image and extract width, height\n\nresults = inference(image, prompt + output_prompt)\n\n# Clean the results and load results in list format\n\ncln_results = json.loads(clean_results(results))\n\nprint()\n\nannotator = Annotator(image)\u00a0 # initialize Ultralytics annotator\n\nfor idx, item in enumerate(cln_results):\n\n\u00a0\u00a0\u00a0# By default, gemini model return output with y coordinates first.\n\n\u00a0\u00a0\u00a0# Scale normalized box coordinates (0\u20131000) to image dimensions\n\n\u00a0\u00a0\u00a0y1, x1, y2, x2 = item[\"box_2d\"]\u00a0 # bbox post processing,\n\n\u00a0\u00a0\u00a0y1 = y1 \/ 1000 * h\n\n\u00a0\u00a0\u00a0x1 = x1 \/ 1000 * w\n\n\u00a0\u00a0\u00a0y2 = y2 \/ 1000 * h\n\n\u00a0\u00a0\u00a0x2 = x2 \/ 1000 * w\n\n\u00a0\u00a0\u00a0if x1 &gt; x2:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0x1, x2 = x2, x1\u00a0 # Swap x-coordinates if needed\n\n\u00a0\u00a0\u00a0if y1 &gt; y2:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0y1, y2 = y2, y1\u00a0 # Swap y-coordinates if needed\n\n\u00a0\u00a0\u00a0annotator.box_label([x1, y1, x2, y2], label=item[\"label\"], color=colors(idx, True))\n\nImage.fromarray(annotator.result())\u00a0 # display the output<\/code><\/pre>\n<p><strong>Source Image: <\/strong><a href=\"https:\/\/www.androidcentral.com\/apps-software\/ai\/googles-plan-to-steal-chatgpts-market-share-is-all-about-geminis-free-tier\" target=\"_blank\" rel=\"nofollow noopener\">Link<\/a><\/p>\n<p><strong>Output<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"831\" height=\"502\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fifth.webp\" alt=\"Task 3 Output\" class=\"wp-image-230553\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fifth.webp 831w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fifth-300x181.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fifth-768x464.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fifth-200x120.webp 200w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/Fifth-150x91.webp 150w\" sizes=\"auto, (max-width: 831px) 100vw, 831px\"\/><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-explanation-2\">Explanation<\/h4>\n<ol class=\"wp-block-list\">\n<li>This uses a prompt similar to object detection but asks for text (label) instead of object names.<\/li>\n<li>The code extracts the text and its location, printing the text and drawing boxes on the image.<\/li>\n<li>This is useful for digitizing documents or reading text from signs or labels in photos.<\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>Google Gemini for Computer Vision makes it easy to tackle tasks like object detection, image captioning, and OCR through simple API calls. By sending images along with clear text instructions, you can guide the model\u2019s understanding and get usable, real-time results.\u00a0<\/p>\n<p>That said, while Gemini is great for general-purpose tasks or quick experiments, it\u2019s not always the best fit for highly specialized use cases. Suppose you\u2019re working with niche objects or need tighter control over accuracy. In that case, the traditional route still holds strong: collect your dataset, annotate it with tools like YOLO labelers, and train a custom model tuned for your needs.<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/harsh9480979\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_0fBqNLi.webp\" width=\"48\" height=\"48\" alt=\"Harsh Mishra\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Harsh Mishra is an AI\/ML Engineer who spends more time talking to Large Language Models than actual humans. Passionate about GenAI, NLP, and making machines smarter (so they don\u2019t replace him just yet). When not optimizing models, he\u2019s probably optimizing his coffee intake. \ud83d\ude80\u2615<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Since the rise of AI chatbots, Google\u2019s Gemini has emerged as one of the most powerful players driving the evolution of intelligent systems. Beyond its conversational strength, Gemini also unlocks practical possibilities in computer vision, enabling machines to see, interpret, and describe the world around them. This guide walks you through the steps to leverage [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":202267,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[4507,20726,1531,8558,22004,8767],"dealstore":[],"offerexpiration":[],"class_list":["post-202266","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-computer","tag-gemini","tag-google","tag-models","tag-tasks","tag-vision"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>How to Use Google Gemini Models for Computer Vision Tasks? - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=202266\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"How to Use Google Gemini Models for Computer Vision Tasks? - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Since the rise of AI chatbots, Google\u2019s Gemini has emerged as one of the most powerful players driving the evolution of intelligent systems. Beyond its conversational strength, Gemini also unlocks practical possibilities in computer vision, enabling machines to see, interpret, and describe the world around them. This guide walks you through the steps to leverage [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=202266\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-04-24T07:42:40+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=202266#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=202266\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"How to Use Google Gemini Models for Computer Vision Tasks?\",\"datePublished\":\"2025-04-24T07:42:40+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=202266\"},\"wordCount\":1226,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=202266#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp\",\"keywords\":[\"Computer\",\"Gemini\",\"Google\",\"Models\",\"Tasks\",\"vision\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=202266#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=202266\",\"url\":\"https:\/\/fivemor.com\/?p=202266\",\"name\":\"How to Use Google Gemini Models for Computer Vision Tasks? - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=202266#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=202266#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp\",\"datePublished\":\"2025-04-24T07:42:40+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=202266#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=202266\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=202266#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=202266#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"How to Use Google Gemini Models for Computer Vision Tasks?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"How to Use Google Gemini Models for Computer Vision Tasks? - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=202266","og_locale":"en_US","og_type":"article","og_title":"How to Use Google Gemini Models for Computer Vision Tasks? - Som2ny Network","og_description":"Since the rise of AI chatbots, Google\u2019s Gemini has emerged as one of the most powerful players driving the evolution of intelligent systems. Beyond its conversational strength, Gemini also unlocks practical possibilities in computer vision, enabling machines to see, interpret, and describe the world around them. This guide walks you through the steps to leverage [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=202266","og_site_name":"Som2ny Network","article_published_time":"2025-04-24T07:42:40+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=202266#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=202266"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"How to Use Google Gemini Models for Computer Vision Tasks?","datePublished":"2025-04-24T07:42:40+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=202266"},"wordCount":1226,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=202266#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp","keywords":["Computer","Gemini","Google","Models","Tasks","vision"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=202266#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=202266","url":"https:\/\/fivemor.com\/?p=202266","name":"How to Use Google Gemini Models for Computer Vision Tasks? - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=202266#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=202266#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp","datePublished":"2025-04-24T07:42:40+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=202266#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=202266"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=202266#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/How-to-Use-Google-Gemini-Models-for-Computer-Vision-Tasks.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=202266#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"How to Use Google Gemini Models for Computer Vision Tasks?"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/202266","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=202266"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/202266\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/202267"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=202266"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=202266"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=202266"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=202266"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=202266"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}