{"id":89387,"date":"2025-02-15T09:05:18","date_gmt":"2025-02-15T09:05:18","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/enhancing-multimodal-rag-with-deepseek-janus-pro\/"},"modified":"2025-02-15T09:05:18","modified_gmt":"2025-02-15T09:05:18","slug":"enhancing-multimodal-rag-with-deepseek-janus-pro","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=89387","title":{"rendered":"Enhancing Multimodal RAG with Deepseek Janus Pro"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>DeepSeek Janus Pro 1B, launched on January 27, 2025, is an advanced multimodal AI model built to process and generate images from textual prompts. With its ability to comprehend and create images based on text, this 1 billion parameter version (1B) delivers efficient performance for a wide range of applications, including text-to-image generation and image understanding. Additionally, it excels at producing detailed captions from photos, making it a versatile tool for both creative and analytical tasks.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-learning-objectives\">Learning Objectives<\/h3>\n<ul class=\"wp-block-list\">\n<li>Analyzing its architecture and key features that enhance its capabilities.<\/li>\n<li>Exploring the underlying design and its impact on performance.<\/li>\n<li>A step-by-step guide to building a Retrieval-Augmented Generation (RAG) system.<\/li>\n<li>Utilizing the DeepSeek Janus Pro 1 billion model for real-world applications.<\/li>\n<li>Understanding how DeepSeek Janus Pro optimizes AI-driven solutions.<\/li>\n<\/ul>\n<p><em><strong>This article was published as a part of the\u00a0<\/strong><\/em><a href=\"https:\/\/www.analyticsvidhya.com\/datahack\/blogathon\" target=\"_blank\" rel=\"noreferrer noopener\"><em><strong>Data Science Blogathon.<\/strong><\/em><\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-deepseek-janus-pro\">What is DeepSeek Janus Pro?<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/01\/run-deepseek-models-locally\/\" target=\"_blank\" rel=\"noreferrer noopener\">DeepSeek<\/a> Janus Pro is a multimodal AI model that integrates text and image processing, capable of understanding and generating images from text prompts. The 1 billion parameter version (1B) is designed for efficient performance across applications like text-to-image generation and image understanding tasks.<\/p>\n<p>Under DeepSeek\u2019s Janus Pro series, the primary models available are<b>\u00a0\u201cJanus Pro 1B\u201d and \u201cJanus Pro 7B\u201d,<\/b> which differ mainly in their parameter size, with the 7B model being significantly larger and offering improved performance in text-to-image generation tasks;\u00a0both are considered multimodal models capable of handling both visual understanding and text generation based on visual context.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-features-and-design-aspects-of-janus-pro-1b\">Key Features and Design Aspects of Janus Pro 1B<\/h3>\n<ul class=\"wp-block-list\">\n<li><b>Architecture<\/b>: Janus Pro uses a unified transformer architecture but decouples visual encoding into separate pathways to improve performance in both image understanding and creation tasks.<\/li>\n<li><b>Capabilities<\/b>: It excels in tasks related to both understanding of images and the generation of new ones based on text prompts. It supports 384\u00d7384 image inputs.<\/li>\n<li><b>Image Encoders<\/b>: For image understanding tasks, Janus uses SigLIP to encode images. SigLIP is an image embedding model that uses CLIP\u2019s framework but replaces the loss function with a pairwise sigmoid loss. For image generation, Janus uses an existing encoder from LlamaGen, an autoregressive image generation mode. LlamaGen is a family of image-generation models that applies the next-token prediction paradigm of large language models to a visual generation<\/li>\n<li><b>Open Source:<\/b> It is available on GitHub under the MIT License, with model usage governed by the DeepSeek Model License.<\/li>\n<\/ul>\n<p>Also read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/01\/deepseek-janus-pro-7b\/\" target=\"_blank\" rel=\"noreferrer noopener\">How to Access DeepSeek Janus Pro 7B?<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-decoupled-architecture-for-image-understanding-amp-generation\">Decoupled Architecture For Image Understanding &amp; Generation<\/h2>\n<div class=\"wp-block-image figure mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"802\" height=\"208\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_O4tMZCG.webp\" alt=\"Architectural Features of Deepsee\" class=\"wp-image-221526\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_O4tMZCG.webp 802w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_O4tMZCG-300x78.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_O4tMZCG-768x199.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_O4tMZCG-150x39.webp 150w\" sizes=\"(max-width: 802px) 100vw, 802px\"\/><figcaption class=\"wp-element-caption\">Architectural Features of Deepsee<\/figcaption><\/figure>\n<\/div>\n<p>Janus-Pro diverges from previous multimodal models by employing separate, specialized pathways for visual encoding, rather than relying on a single visual encoder for both image understanding and generation.<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Image Understanding Encoder.<\/b> This pathway extracts semantic features from images.<\/li>\n<li><b>Image Generation Encoder.<\/b> This pathway synthesizes images based on text descriptions.<\/li>\n<\/ul>\n<p>This decoupled architecture facilitates task-specific optimizations, mitigating conflicts between interpretation and creative synthesis. The independent encoders interpret input features which are then processed by a unified autoregressive transformer. This allows both multimodal understanding and generation components to independently select their most suitable encoding methods.<\/p>\n<p>Also read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/01\/janus-pro-7b-vs-dall-e-3\/\" target=\"_blank\" rel=\"noreferrer noopener\">How DeepSeek\u2019s Janus Pro Stacks Up Against DALL-E 3?<\/a><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-features-of-model-architecture\">Key Features of Model Architecture<\/h3>\n<h3 class=\"wp-block-heading\" id=\"h-1-dual-pathway-architecture-for-visual-understanding-amp-generation\">1. Dual-pathway architecture for visual understanding &amp; generation<\/h3>\n<ul class=\"wp-block-list\">\n<li><b>Visual Understanding Pathway:<\/b>\u00a0For multimodal understanding tasks, Janus Pro uses SigLIP-L as the visual encoder, which supports image inputs of up to 384\u00d7384 resolution. This high-resolution support allows the model to capture more image details, thereby improving the accuracy of visual understanding.\u00a0\u00a0<\/li>\n<li><b>Visual Generation Pathway<\/b>:\u00a0For image generation tasks, Janus Pro uses LlamaGen Tokenizer with a downsampling rate of 16 to generate more detailed images.\u00a0\u00a0<\/li>\n<\/ul>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"790\" height=\"321\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_RdlmqXs.webp\" alt=\"DeepSeek Janus-Pro\" class=\"wp-image-221537\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_RdlmqXs.webp 790w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_RdlmqXs-300x122.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_RdlmqXs-768x312.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_RdlmqXs-150x61.webp 150w\" sizes=\"auto, (max-width: 790px) 100vw, 790px\"\/><figcaption class=\"wp-element-caption\">Fig 1. The architecture of our Janus-Pro. We decouple visual encoding for multimodal understanding and visual generation. \u201cUnd. Encoder\u201d and \u201cGen. Encoder\u201d are abbreviations for \u201cUnderstanding Encoder\u201d and \u201cGeneration Encoder\u201d, respectively. Source: <a href=\"https:\/\/github.com\/deepseek-ai\/Janus\/blob\/main\/janus_pro_tech_report.pdf\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">DeepSeek Janus-Pro<\/a><\/figcaption><\/figure>\n<h3 class=\"wp-block-heading\" id=\"h-2-unified-transformer-architecture\">2. Unified Transformer Architecture<\/h3>\n<p>A shared transformer backbone is used\u00a0for\u00a0text and image feature fusion. The independent encoding methods to convert the raw inputs into features are processed by a unified autoregressive transformer.\u00a0\u00a0<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-3-optimized-training-strategy\">3. Optimized Training Strategy<\/h3>\n<p>In Previous Janus training, there was a three-stage training process for the model. The first stage focused on training the adaptors and the image head. The second stage handled unified pretraining, during which all components except the understanding encoder and the generation encoder have their parameters updated. Stage III covered supervised fine-tuning, building upon Stage II by further unlocking the parameters of the understanding encoder during training.<\/p>\n<p>This was improved in Janus Pro:<\/p>\n<ul class=\"wp-block-list\">\n<li>By increasing the training steps in Stage I, allowing sufficient training on the ImageNet dataset.<\/li>\n<li>Additionally, in Stage II, for text-to-image generation training, the ImageNet data was dropped completely. Instead normal text-to-image data was utilized to train the model to generate images based on dense descriptions. This was found to improve the training efficiency and overall performance.<\/li>\n<\/ul>\n<p>Now, lets build Multimodal RAG with Deepseek Janus Pro:<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-multimodal-rag-with-deepseek-janus-pro-1b-model\">Multimodal RAG with Deepseek Janus Pro 1B model<\/h2>\n<p>In the following steps, we will build a multimodal RAG system to query on images based on the Deepseek Janus Pro 1B model.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-1-install-necessary-libraries\">Step 1. Install Necessary Libraries<\/h3>\n<pre class=\"wp-block-code\"><code>!pip install byaldi ollama pdf2image\n!sudo apt-get install -y poppler-utils\n!git clone https:\/\/github.com\/deepseek-ai\/Janus.git\n!pip install -e .\/Janus<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-step-2-model-for-saving-image-embeddings\">Step 2. Model For Saving Image Embeddings<\/h3>\n<pre class=\"wp-block-code\"><code>import os\nfrom pathlib import Path\nfrom byaldi import RAGMultiModalModel\nimport ollama\n# Initialize RAGMultiModalModel\nmodel1 = RAGMultiModalModel.from_pretrained(\"vidore\/colqwen2-v0.1\")<\/code><\/pre>\n<p>Byaldi gives an easy-to-use framework for setting up multimodal RAG systems. As seen from the above code, we load Colqwen2, which is a model designed for efficient document indexing using visual features.\u00a0<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-3-loading-the-image-pdf\">Step 3. Loading the Image PDF<\/h3>\n<pre class=\"wp-block-code\"><code># Use ColQwen2 to index and store the presentation\nindex_name = \"image_index\"\nmodel1.index(input_path=Path(\"\/content\/PublicWaterMassMailing.pdf\"),\n    index_name=index_name,\n    store_collection_with_index=True, # Stores base64 images along with the vectors\n    overwrite=True\n)<\/code><\/pre>\n<p>We use this <a href=\"https:\/\/investor.accenture.com\/~\/media\/Files\/A\/Accenture-IR-V3\/events-and-presentations\/accenture-investor-and-analyst-conference-cfo-slides.pdf\" target=\"_blank\" rel=\"nofollow noopener\">PDF<\/a> to query and build an RAG system on in the next steps. In the above code, we store the image PDF along with the vectors.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-4-querying-amp-retrieval-from-saved-images\">Step 4. Querying &amp; Retrieval From Saved Images<\/h3>\n<pre class=\"wp-block-code\"><code>query = \"How many clients drive more than 50% revenue?\"\nreturned_page = model1.search(query, k=1)[0]\nimport base64\n# Example Base64 string (truncated for brevity)\nbase64_string = returned_page['base64']\n\n# Decode the Base64 string\nimage_data = base64.b64decode(base64_string)\nwith open('output_image.png', 'wb') as image_file:\n    image_file.write(image_data)<\/code><\/pre>\n<p>The relevant page from the pages of the PDF is retrieved and saved as output_image.png based on the query.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-5-load-janus-pro-model\">Step 5. Load Janus Pro Model<\/h3>\n<pre class=\"wp-block-code\"><code>import os\nos.chdir(r\"\/content\/Janus\")\n\nfrom janus.models import VLChatProcessor\nfrom transformers import AutoConfig, AutoModelForCausalLM\nimport torch\nfrom janus.utils.io import load_pil_images\nfrom PIL import Image\n\nprocessor= VLChatProcessor.from_pretrained(\"deepseek-ai\/Janus-Pro-1B\")\ntokenizer = processor.tokenizer\nvl_gpt = AutoModelForCausalLM.from_pretrained(\n    \"deepseek-ai\/Janus-Pro-1B\", trust_remote_code=True\n)\n\nconversation = [\n    {\n        \"role\": \"\",\n        \"content\": f\"<image_placeholder>\\n{query}\",\n        \"images\": ['\/content\/output_image.png'],\n    },\n    {\"role\": \"\", \"content\": \"\"},\n]\n\n# load images and prepare for inputs\npil_images = load_pil_images(conversation)\ninputs = processor(conversations=conversation, images=pil_images)\n\n# # run image encoder to get the image embeddings\ninputs_embeds = vl_gpt.prepare_inputs_embeds(**inputs)<\/image_placeholder><\/code><\/pre>\n<ul class=\"wp-block-list\">\n<li><b>VLChatProcessor.from_pretrained(\u201cdeepseek-ai\/Janus-Pro-1B\u201d)<\/b> loads a pretrained processor for handling multimodal inputs (images and text). This processor will process and prepare input data (like text and images) for the model.<\/li>\n<li>The tokenizer is extracted from the VLChatProcessor. It will tokenize the text input, converting text into a format suitable for the model.<\/li>\n<li><b>AutoModelForCausalLM.from_pretrained(\u201cdeepseek-ai\/Janus-Pro-1B\u201d) <\/b>loads the pre-trained Janus Pro model, specifically for causal language modelling.<\/li>\n<li>Also, <b>a multimodal conversation format <\/b>is set up where the user inputs both text and an image.<\/li>\n<li>The <b>load_pil_images(conversation) <\/b>is a function that likely loads the images listed in the conversation object and converts them into PIL Image format, which is commonly used for image processing in Python.<\/li>\n<li>The <b>processor<\/b> here is an instance of a multimodal processor (the <b>VLChatProcessor <\/b>from the DeepSeek Janus Pro model), which takes both text and image data as input.<\/li>\n<li><b>prepare_inputs_embeds(inputs)<\/b> is a method that takes the processed inputs (inputs contain both the text and image)\u00a0, and prepares the embeddings required for the model to generate a response.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-step-6-output-generation\">Step 6. Output Generation<\/h3>\n<pre class=\"wp-block-code\"><code>outputs =  vl_gpt.language_model.generate(\n    inputs_embeds=inputs_embeds,\n    attention_mask=inputs.attention_mask,\n    pad_token_id=tokenizer.eos_token_id,\n    bos_token_id=tokenizer.bos_token_id,\n    eos_token_id=tokenizer.eos_token_id,\n    max_new_tokens=512,\n    do_sample=False,\n    use_cache=True,\n)\n\nanswer = tokenizer.decode(outputs[0].cpu().tolist(), skip_special_tokens=True)\nprint(answer)<\/code><\/pre>\n<p>The code generates a response from the DeepSeek Janus Pro 1B model using the prepared input embeddings (text and image). It uses several configuration settings like padding, start\/end tokens, max token length, and whether to use caching and sampling. After the response is generated, it decodes the token IDs back into human-readable text using the tokenizer. The decoded output is stored in the answer variable.\u00a0<\/p>\n<p>The whole code is present in this <a href=\"https:\/\/colab.research.google.com\/drive\/1vwYH8gl_KL1rnf994vX6KojD3a9wM6VM?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">colab notebook<\/a>.\u00a0\u00a0<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-output-for-the-query\">Output For the Query<\/h4>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"190\" height=\"28\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_76_YG7oUtq.webp\" alt=\"output\" class=\"wp-image-221539\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_76_YG7oUtq.webp 190w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_76_YG7oUtq-150x22.webp 150w\" sizes=\"auto, (max-width: 190px) 100vw, 190px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-output-for-another-query\">Output For Another Query<\/h4>\n<p><em>\u201cWhat has been the revenue in France?\u201d<\/em><\/p>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"73\" height=\"39\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_mJHQSUu.webp\" alt=\"output\" class=\"wp-image-221522\"\/><\/figure>\n<p>The above response is not accurate even though the relevant page was retrieved by the\u00a0colqwen2 retriever, the DeepSeek Janus Pro 1B model could not generate the accurate answer from the page. The exact answer should be $2B.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-output-for-another-query-0\">Output For Another Query<\/h4>\n<p><em>\u201c\u201dWhat has been the number of promotions since beginning of FY20?\u201d<\/em><\/p>\n<figure class=\"wp-block-image size-full figure mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"84\" height=\"33\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/image_Hsk2vXk.webp\" alt=\"output\" class=\"wp-image-221523\"\/><\/figure>\n<p>The above response is correct as it matches with the text mentioned in the PDF.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusions\">Conclusions<\/h2>\n<p>In conclusion, the DeepSeek Janus Pro 1B model represents a significant advancement in multimodal AI, with its decoupled architecture that optimizes both image understanding and generation tasks. By utilizing separate visual encoders for these tasks and refining its training strategy, Janus Pro offers enhanced performance in text-to-image generation and image analysis. This innovative approach (Multimodal RAG with Deepseek Janus Pro), combined with its open-source accessibility, makes it a powerful tool for various applications in AI-driven visual comprehension and creation.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h2>\n<ol class=\"wp-block-list\">\n<li><b>Multimodal AI with Dual Pathways<\/b>: Janus Pro 1B integrates both text and image processing, using separate encoders for image understanding (SigLIP) and image generation (LlamaGen), enhancing task-specific performance.<\/li>\n<li><b>Decoupled Architecture:<\/b> The model separates visual encoding into distinct pathways, enabling independent optimization for image understanding and generation, thus minimizing conflicts in processing tasks.<\/li>\n<li><b>Unified Transformer Backbone<\/b>: A shared transformer architecture merges the features of text and images, streamlining multimodal data fusion for more effective AI performance.<\/li>\n<li><b>Improved Training Strategy:<\/b> Janus Pro\u2019s optimized training approach includes increased steps in Stage I and the use of specialized text-to-image data in Stage II, significantly boosting training efficiency and output quality.<\/li>\n<li><b>Open-Source Accessibility: <\/b>Janus Pro 1B is available on GitHub under the MIT License, encouraging widespread use and adaptation in various AI-driven applications.<\/li>\n<\/ol>\n<p><strong>The media shown in this article is not owned by Analytics Vidhya and is used at the Author\u2019s discretion.<\/strong><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/karthik3852845\/\"\/><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1739519714218\"><strong class=\"schema-faq-question\">Q1. What is DeepSeek Janus Pro 1B?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. DeepSeek Janus Pro 1B is a multimodal AI model designed to integrate both text and image processing, capable of understanding and generating images from text descriptions. It features 1 billion parameters for efficient performance in tasks like text-to-image generation and image understanding.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739519727698\"><strong class=\"schema-faq-question\">Q2. How does the architecture of Janus Pro 1B work?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Janus Pro uses a unified transformer architecture with decoupled visual encoding. This means it employs separate pathways for image understanding and generation, allowing task-specific optimization for each task.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739519743524\"><strong class=\"schema-faq-question\">Q3. How does the training process of Janus Pro differ from previous versions?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Janus Pro improves on previous training strategies by increasing training steps, dropping the ImageNet dataset in favor of specialized text-to-image data, and focusing on better fine-tuning for enhanced efficiency and performance.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739519762277\"><strong class=\"schema-faq-question\">Q4. What kind of applications can benefit from using Janus Pro 1B?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Janus Pro 1B is particularly useful for tasks involving text-to-image generation, image understanding, and multimodal AI applications that require both image and text processing capabilities<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739519787216\"><strong class=\"schema-faq-question\">Q5. How does Janus-Pro compare to other models like DALL-E 3?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Janus-Pro-7B outperforms DALL-E 3 in benchmarks such as GenEval and DPG-Bench, according to DeepSeek. Janus-Pro separates understanding\/generation, scales data\/models for stable image generation, and maintains a unified, flexible, and cost-efficient structure. While both models perform text-to-image generation, Janus-Pro also offers image captioning, which DALL-E 3 does not.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/mimi6\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_ZkJo4gb.webp\" width=\"48\" height=\"48\" alt=\"Nibedita Dutta\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Nibedita completed her master\u2019s in Chemical Engineering from IIT Kharagpur in 2014 and is currently working as a Senior Data Scientist. In her current capacity, she works on building intelligent ML-based solutions to improve business processes.               <\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>DeepSeek Janus Pro 1B, launched on January 27, 2025, is an advanced multimodal AI model built to process and generate images from textual prompts. With its ability to comprehend and create images based on text, this 1 billion parameter version (1B) delivers efficient performance for a wide range of applications, including text-to-image generation and image [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":51392,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5815,29466,16626,30678,20383,1190,32726],"dealstore":[],"offerexpiration":[],"class_list":["post-89387","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-blogathon","tag-deepseek","tag-enhancing","tag-janus","tag-multimodal","tag-pro","tag-rag"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Enhancing Multimodal RAG with Deepseek Janus Pro - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=89387\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Enhancing Multimodal RAG with Deepseek Janus Pro - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"DeepSeek Janus Pro 1B, launched on January 27, 2025, is an advanced multimodal AI model built to process and generate images from textual prompts. With its ability to comprehend and create images based on text, this 1 billion parameter version (1B) delivers efficient performance for a wide range of applications, including text-to-image generation and image [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=89387\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-15T09:05:18+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"493\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=89387#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=89387\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Enhancing Multimodal RAG with Deepseek Janus Pro\",\"datePublished\":\"2025-02-15T09:05:18+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=89387\"},\"wordCount\":1784,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=89387#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp\",\"keywords\":[\"Blogathon\",\"DeepSeek\",\"Enhancing\",\"Janus\",\"Multimodal\",\"Pro\",\"RAG\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=89387#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=89387\",\"url\":\"https:\/\/fivemor.com\/?p=89387\",\"name\":\"Enhancing Multimodal RAG with Deepseek Janus Pro - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=89387#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=89387#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp\",\"datePublished\":\"2025-02-15T09:05:18+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=89387#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=89387\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=89387#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp\",\"width\":872,\"height\":493},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=89387#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Enhancing Multimodal RAG with Deepseek Janus Pro\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Enhancing Multimodal RAG with Deepseek Janus Pro - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=89387","og_locale":"en_US","og_type":"article","og_title":"Enhancing Multimodal RAG with Deepseek Janus Pro - Som2ny Network","og_description":"DeepSeek Janus Pro 1B, launched on January 27, 2025, is an advanced multimodal AI model built to process and generate images from textual prompts. With its ability to comprehend and create images based on text, this 1 billion parameter version (1B) delivers efficient performance for a wide range of applications, including text-to-image generation and image [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=89387","og_site_name":"Som2ny Network","article_published_time":"2025-02-15T09:05:18+00:00","og_image":[{"width":872,"height":493,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=89387#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=89387"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Enhancing Multimodal RAG with Deepseek Janus Pro","datePublished":"2025-02-15T09:05:18+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=89387"},"wordCount":1784,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=89387#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp","keywords":["Blogathon","DeepSeek","Enhancing","Janus","Multimodal","Pro","RAG"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=89387#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=89387","url":"https:\/\/fivemor.com\/?p=89387","name":"Enhancing Multimodal RAG with Deepseek Janus Pro - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=89387#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=89387#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp","datePublished":"2025-02-15T09:05:18+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=89387#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=89387"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=89387#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/DeepSeek-R1-OpenAIs-o1-Biggest-Competitor-is-HERE.webp.webp","width":872,"height":493},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=89387#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Enhancing Multimodal RAG with Deepseek Janus Pro"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/89387","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=89387"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/89387\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/51392"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=89387"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=89387"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=89387"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=89387"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=89387"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}