{"id":147786,"date":"2025-03-21T13:50:03","date_gmt":"2025-03-21T13:50:03","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/can-smoldocling-make-document-parsing-more-efficient\/"},"modified":"2025-03-21T13:50:03","modified_gmt":"2025-03-21T13:50:03","slug":"can-smoldocling-make-document-parsing-more-efficient","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=147786","title":{"rendered":"Can SmolDocling Make Document Parsing More Efficient?"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Digital documents have long presented a dual challenge for both human readers and automated systems: preserving rich structural nuances while converting content into machine-processable formats. Traditional methods, whether relying on complex ensemble pipelines or massive foundational models, often struggle to balance accuracy with computational efficiency. SmolDocling emerges as a game-changing solution, offering an ultra-compact 256M-parameter vision-language model that performs end-to-end document conversion with remarkable precision and speed.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-the-challenge-of-document-conversion\">The Challenge of Document Conversion<\/h2>\n<p>For decades, converting complex layouts ranging from business documents to academic papers into structured representations has been a difficult task. Common issues include:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Layout Variability:<\/strong> Documents present a wide array of layouts and styles.<\/li>\n<li><strong>Opaque Formats:<\/strong> Formats like PDF are optimized for printing rather than semantic parsing, obscuring the underlying structure.<\/li>\n<li><strong>Resource Demands:<\/strong> Traditional large-scale models or ensemble solutions require extensive computational resources and intricate tuning.<\/li>\n<\/ul>\n<p>These challenges have led to a lot of research, but finding a solution that is both efficient and accurate is still difficult.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-introducing-smoldocling\">Introducing SmolDocling<\/h2>\n<p>SmolDocling addresses these hurdles head-on by leveraging a unified approach:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>End-to-End Conversion:<\/strong> Instead of piecing together multiple specialized models, SmolDocling processes entire document pages in one go.<\/li>\n<li><strong>Compact yet Powerful:<\/strong> With just 256M parameters, it delivers performance comparable to models up to 27 times larger.<\/li>\n<li><strong>Robust Multi-Modal Capabilities:<\/strong> Whether dealing with code listings, tables, equations, or complex charts, SmolDocling adapts seamlessly across diverse document types.<\/li>\n<\/ul>\n<p>At its core, the model introduces a novel markup format known as DocTags\u2014a universal standard that meticulously captures every element\u2019s content, structure, and spatial context.<\/p>\n<p>DocTags revolutionize the way document elements are represented:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Structured Vocabulary:<\/strong> Inspired by earlier work like OTSL, DocTags use XML-style tags to explicitly differentiate between text, images, tables, code, and more.<\/li>\n<li><strong>Spatial Awareness:<\/strong> Each element is annotated with precise bounding box coordinates, ensuring that layout context is preserved.<\/li>\n<li><strong>Unified Representation:<\/strong> Whether processing a full-page document or an isolated element (like a cropped table), the format remains consistent, boosting the model\u2019s ability to learn and generalize.<\/li>\n<\/ul>\n<ul class=\"wp-block-list\">\n<li><strong><picture\/><\/strong> \u2013 Represents an image or visual content in the document.<\/li>\n<li><strong><flow_chart\/><\/strong> \u2013 Likely represents a diagram or structured graphical representation.<\/li>\n<li><strong><br \/>\n<caption\/><\/strong> \u2013 Provides a description or annotation for an image or diagram.<\/li>\n<li><strong><otsl\/><\/strong> \u2013 Possibly represents a structured document format for tables or layouts.<\/li>\n<li><strong><loc_xx\/><\/strong> \u2013 Indicates the position of an element within the document.<\/li>\n<li><strong><ched\/><\/strong> \u2013 Likely a shorthand for \u201cheader\u201d or \u201ccategorical header\u201d within a table.<\/li>\n<li><strong><fcel\/><\/strong> \u2013 Probably refers to \u201cformatted cell,\u201d indicating specific cell content in tables.<\/li>\n<li><strong><nl\/><\/strong> \u2013 Represents a new line or a break in text.<\/li>\n<li><strong><section_header_level_1\/><\/strong> \u2013 Marks a major section heading in the document.<\/li>\n<li><strong><text\/><\/strong> \u2013 Defines general text content within the document.<\/li>\n<li><strong><unordered_list\/><\/strong> \u2013 Represents a bulleted or unordered list.<\/li>\n<li><strong><list_item\/><\/strong> \u2013 Specifies an individual item within a list.<\/li>\n<li><strong><code\/><\/strong> \u2013 Contains programming or script-related content, formatted for readability.<\/li>\n<\/ul>\n<p>This clear, structured format minimizes ambiguity, a common issue with direct conversion methods to formats like HTML or Markdown.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-deep-dive-dataset-training-and-model-architecture\">Deep Dive: Dataset Training and Model Architecture<\/h2>\n<h3 class=\"wp-block-heading\" id=\"h-dataset-training\"><strong>Dataset Training<\/strong><\/h3>\n<p>A key pillar of SmolDocling\u2019s success is its rich, diverse training data:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Pre-training Data:<\/strong><strong><br \/><\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>DocLayNet-PT:<\/strong> A 1.4M page dataset extracted from unique PDF documents sourced from CommonCrawl, Wikipedia, and business documents. This dataset is enriched with weak annotations covering layout elements, table structures, language, topics, and figure classifications.<\/li>\n<li><strong>DocMatix:<\/strong> Adapted using a similar weak annotation strategy as DocLayNet-PT, this dataset includes multi-task document conversion tasks.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Task-Specific Data:<\/strong><strong><br \/><\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Layout &amp; Structure:<\/strong> High-quality annotated pages from DocLayNet v2, WordScape, and synthetically generated pages from SynthDocNet ensure robust layout and table structure learning.<\/li>\n<li><strong>Charts, Code, and Equations:<\/strong> Custom-generated datasets provide extensive visual diversity. For instance, over 2.5 million charts are generated using three different visualization libraries, while 9.3M rendered code snippets and 5.5M formulas provide detailed coverage of technical document elements.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Instruction Tuning:<\/strong> To reinforce the recognition of different page elements and introduce document-related features and no-code pipelines, rule-based techniques and the <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/05\/ibm-granite-code-models\/\" target=\"_blank\" rel=\"noreferrer noopener\">Granite-3.1-2b-instruct<\/a> LLM were leveraged. Using samples from DocLayNet-PT pages, one instruction was generated by randomly sampling layout elements from a page. These instructions included tasks such as:\n<ul class=\"wp-block-list\">\n<li>\u201cPerform OCR at bbox\u201d<\/li>\n<li>\u201cIdentify page element type at bbox\u201d<\/li>\n<li>\u201cExtract all section headers from the page\u201d<\/li>\n<\/ul>\n<\/li>\n<li>Additionally, training with the Cauldron dataset helps avoid catastrophic forgetting due to the introduction of numerous conversation datasets.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-model-architecture-of-smoldocling\">Model Architecture of SmolDocling<\/h3>\n<p>SmolDocling builds upon the SmolVLM framework and incorporates several innovative techniques to ensure efficiency and effectiveness:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Vision Encoder with SigLIP Backbone:<\/strong> The model uses a <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/10\/googles-siglip\/\" target=\"_blank\" rel=\"noreferrer noopener\">SigLIP<\/a> base 16\/512 encoder (93M parameters) which applies an aggressive pixel shuffle strategy. This compresses each 512\u00d7512 image patch into 64 visual tokens, significantly reducing the number of image hidden states.<\/li>\n<li><strong>Enhanced Tokenization:<\/strong> By increasing the pixel-to-token ratio (up to 4096 pixels per token) and introducing special tokens for sub-image separation, tokenization efficiency is markedly improved. This design ensures that both full-page documents and cropped elements are processed uniformly.<\/li>\n<li><strong>Curriculum Learning Approach:<\/strong> Training begins with freezing the vision encoder, focusing on aligning the language model with the new DocTags format. Once the model is familiar with the output structure, the vision encoder is unfrozen and fine-tuned along with task-specific datasets, ensuring comprehensive learning.<\/li>\n<li><strong>Efficient Inference:<\/strong> With a maximum sequence length of 8,192 tokens and the ability to process up to three pages at a time, SmolDocling achieves page conversion times of just 0.35 seconds using VLLM on an A100 GPU, while occupying only 0.489 GB of VRAM.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-comparative-analysis-smoldocling-versus-other-models\">Comparative Analysis: SmolDocling Versus Other Models<\/h2>\n<p>A thorough evaluation of SmolDocling against leading vision-language models highlights its competitive edge:<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-text-recognition-ocr-and-document-formatting\">Text Recognition (OCR) and Document Formatting<\/h3>\n<figure class=\"wp-block-table\">\n<table class=\"table table-bordered border-black table-striped\">\n<tbody>\n<tr>\n<td><strong>Method<\/strong><\/td>\n<td><strong>Model Size<\/strong><\/td>\n<td><strong>Edit Distance \u2193<\/strong><\/td>\n<td><strong>F1-score \u2191<\/strong><\/td>\n<td><strong>Precision \u2191<\/strong><\/td>\n<td><strong>Recall \u2191<\/strong><\/td>\n<td><strong>BLEU \u2191<\/strong><\/td>\n<td><strong>METEOR \u2191<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen2.5 VL [9]<\/td>\n<td>7B<\/td>\n<td>0.56<\/td>\n<td>0.72<\/td>\n<td>0.80<\/td>\n<td>0.70<\/td>\n<td>0.46<\/td>\n<td>0.57<\/td>\n<\/tr>\n<tr>\n<td>GOT [89]<\/td>\n<td>580M<\/td>\n<td>0.61<\/td>\n<td>0.69<\/td>\n<td>0.71<\/td>\n<td>0.73<\/td>\n<td>0.48<\/td>\n<td>0.59<\/td>\n<\/tr>\n<tr>\n<td>Nougat (base) [12]<\/td>\n<td>350M<\/td>\n<td>0.62<\/td>\n<td>0.66<\/td>\n<td>0.72<\/td>\n<td>0.67<\/td>\n<td>0.44<\/td>\n<td>0.54<\/td>\n<\/tr>\n<tr>\n<td><strong>SmolDocling (Ours)<\/strong><\/td>\n<td><strong>256M<\/strong><\/td>\n<td><strong>0.48<\/strong><\/td>\n<td><strong>0.80<\/strong><\/td>\n<td><strong>0.89<\/strong><\/td>\n<td><strong>0.79<\/strong><\/td>\n<td><strong>0.58<\/strong><\/td>\n<td><strong>0.67<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<p><strong>Insights:<\/strong> SmolDocling outperforms larger models across all key metrics in full-page transcription. The significant improvements in F1-score, precision, and recall reflect its superior capability in accurately reproducing textual elements and preserving reading order.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-specialized-tasks-code-listings-and-equations\">Specialized Tasks: Code Listings and Equations<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Code Listings:<\/strong> For tasks like code listing transcription, SmolDocling exhibits an impressive F1-score of 0.92 and precision of 0.94, highlighting its expertise at handling indentation and syntax that carry semantic significance.<\/li>\n<li><strong>Equations:<\/strong> In the domain of equation recognition, SmolDocling closely matches or exceeds the performance of models like Qwen2.5 VL and GOT, achieving an F1-score of 0.95 and precision of 0.96.<\/li>\n<\/ul>\n<p>These results underscore SmolDocling\u2019s ability to not only match but often surpass the performance of models that are significantly larger in size, affirming that a compact model can be both efficient and effective when built with a focused architecture and optimized training strategies.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-code-demonstration-and-output-visualization\">Code Demonstration and Output Visualization<\/h2>\n<p>To provide a practical glimpse into how SmolDocling operates, the following section includes a sample code snippet along with an illustration of the expected output. This example demonstrates how to convert a document image into the DocTags markup format.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-example-1-sample-code-snippet\">Example 1: Sample Code Snippet<\/h3>\n<pre class=\"wp-block-code\"><code>!pip install docling_core\n\n!pip install flash-attn\n\nimport torch\n\nfrom docling_core.types.doc import DoclingDocument\n\nfrom docling_core.types.doc.document import DocTagsDocument\n\nfrom transformers import AutoProcessor, AutoModelForVision2Seq\n\nfrom transformers.image_utils import load_image\n\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n\n# Load images\n\n# Initialize processor and model\n\nprocessor = AutoProcessor.from_pretrained(\"ds4sd\/SmolDocling-256M-preview\")\n\nmodel = AutoModelForVision2Seq.from_pretrained(\n\n\u00a0\u00a0\u00a0\u00a0\"ds4sd\/SmolDocling-256M-preview\",\n\n\u00a0\u00a0\u00a0\u00a0torch_dtype=torch.bfloat16,\n\n\u00a0\u00a0\u00a0\u00a0_attn_implementation=\"flash_attention_2\"# if DEVICE == \"cuda\" else \"eager\",\n\n).to(DEVICE)\n\nmodel.device\n\n# Load images\n\nimage = load_image(\"https:\/\/user-images.githubusercontent.com\/12294956\/47312583-697cfe00-d65a-11e8-930a-e15fd67a5bb1.png\")\n\n# Create input messages\n\nmessages = [\n\n\u00a0\u00a0\u00a0\u00a0{\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"role\": \"user\",\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\"content\": [\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{\"type\": \"image\"},\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0{\"type\": \"text\", \"text\": \"Convert this page to docling.\"}\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0]\n\n\u00a0\u00a0\u00a0\u00a0},\n\n]\n\n# Prepare inputs\n\nprompt = processor.apply_chat_template(messages, add_generation_prompt=True)\n\ninputs = processor(text=prompt, images=[image], return_tensors=\"pt\")\n\ninputs = inputs.to(DEVICE)\n\n# Generate outputs\n\ngenerated_ids = model.generate(**inputs, max_new_tokens=8192)\n\nprompt_length = inputs.input_ids.shape[1]\n\ntrimmed_generated_ids = generated_ids[:, prompt_length:]\n\ndoctags = processor.batch_decode(\n\n\u00a0\u00a0\u00a0\u00a0trimmed_generated_ids,\n\n\u00a0\u00a0\u00a0\u00a0skip_special_tokens=False,\n\n)[0].lstrip()\n\n# Populate document\n\ndoctags_doc = DocTagsDocument.from_doctags_and_image_pairs([doctags], [image])\n\nprint(doctags)\n\n# create a docling document\n\ndoc = DoclingDocument(name=\"Document\")\n\ndoc.load_from_doctags(doctags_doc)\n\nfrom IPython.display import display, Markdown\n\ndisplay(Markdown(doc.export_to_markdown()))<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-input-image\">Input Image<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1847\" height=\"980\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-99.webp\" alt=\"Input\" class=\"wp-image-227313\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-99.webp 1847w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-99-300x159.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-99-768x407.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-99-1536x815.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-99-150x80.webp 150w\" sizes=\"auto, (max-width: 1847px) 100vw, 1847px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-output\">Output<\/h3>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"534\" height=\"264\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-100.webp\" alt=\"Output\" class=\"wp-image-227314\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-100.webp 534w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-100-300x148.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-100-150x74.webp 150w\" sizes=\"auto, (max-width: 534px) 100vw, 534px\"\/><\/figure>\n<p>This output illustrates how various document elements\u2014text blocks, tables, and code listings are precisely marked with their content and spatial information, making them ready for further processing or analysis. But the model is unable to convert all the text DocTags markup format. As you can see, model didn\u2019t read the human written text.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-example-2-sample-code-snippet\">Example 2: Sample Code Snippet<\/h3>\n<pre class=\"wp-block-code\"><code>!curl -L -o image2.png https:\/\/i.imgur.com\/BFN038S.png<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"551\" height=\"1271\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-101-1.webp\" alt=\"receipt\" class=\"wp-image-227312\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-101-1.webp 551w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-101-1-130x300.webp 130w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-101-1-150x346.webp 150w\" sizes=\"auto, (max-width: 551px) 100vw, 551px\"\/><\/figure>\n<\/div>\n<p>The input image is receipt and now we are extracting the text from it. <\/p>\n<pre class=\"wp-block-code\"><code>image = load_image(\".\/image2.png\")<\/code><\/pre>\n<pre class=\"wp-block-code\"><code># Create input messages\nmessages = [\n    {\n        \"role\": \"user\",\n        \"content\": [\n            {\"type\": \"image\"},\n            {\"type\": \"text\", \"text\": \"Convert this page to docling.\"}\n        ]\n    },\n]\n\n# Prepare inputs\nprompt1 = processor.apply_chat_template(messages, add_generation_prompt=True)\ninputs1 = processor(text=prompt1, images=[image], return_tensors=\"pt\")\ninputs1 = inputs1.to(DEVICE)\n\n# Generate outputs\ngenerated_ids = model.generate(**inputs1, max_new_tokens=8192)\nprompt_length = inputs1.input_ids.shape[1]\ntrimmed_generated_ids = generated_ids[:, prompt_length:]\ndoctags = processor.batch_decode(\n    trimmed_generated_ids,\n    skip_special_tokens=False,\n)[0].lstrip()\n\n# Populate document\ndoctags_doc = DocTagsDocument.from_doctags_and_image_pairs([doctags], [image])\nprint(doctags)\n# create a docling document\ndoc = DoclingDocument(name=\"Document\")\ndoc.load_from_doctags(doctags_doc)\n\n# export as any format\n# HTML\n# doc.save_as_html(output_file)\n# MD\nprint(doc.export_to_markdown())<\/code><\/pre>\n<pre class=\"wp-block-code\"><code>from IPython.display import display, Markdown\n\ndisplay(Markdown(doc.export_to_markdown()))<\/code><\/pre>\n<h3 class=\"wp-block-heading\" id=\"h-output-0\">Output<\/h3>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"357\" height=\"965\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-102.webp\" alt=\"Output\" class=\"wp-image-227311\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-102.webp 357w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-102-111x300.webp 111w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/03\/image-102-150x405.webp 150w\" sizes=\"auto, (max-width: 357px) 100vw, 357px\"\/><\/figure>\n<\/div>\n<p>It is quite impressive as the model extracted all the content from the receipt and it is better than the obove given example.<\/p>\n<p>Notebook with full code: <a href=\"https:\/\/colab.research.google.com\/drive\/1Y1Kt713ErVnjX4gpScZNafU1j_HxJBi8?usp=sharing\">Click Here<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion-and-future-directions\">Conclusion and Future Directions<\/h2>\n<p>SmolDocling sets a new benchmark in document conversion by proving that smaller, more efficient models can rival and even surpass the capabilities of their larger counterparts. Its innovative use of DocTags and an end-to-end conversion strategy provide a compelling blueprint for the next generation of vision-language models. It works well with receipts overall and performs acceptably with other documents, though not always perfectly this serves as a consequence of its memory-saving model design.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Efficiency:<\/strong> With a compact 256M parameter architecture, SmolDocling achieves rapid page conversion with minimal computational overhead.<\/li>\n<li><strong>Robustness:<\/strong> Extensive pre-training and task-specific datasets, along with a curriculum learning approach, ensure that the model generalizes well across diverse document types.<\/li>\n<li><strong>Comparative Superiority:<\/strong> Through rigorous evaluations, SmolDocling has demonstrated superior performance in OCR, code listing transcription, and equation recognition compared to larger models.<\/li>\n<\/ul>\n<p>As the research community continues to refine techniques for element localization and multimodal understanding, SmolDocling provides a clear pathway toward more resource-efficient and versatile document processing solutions. With plans to release the accompanying datasets publicly, this work paves the way for further advancements and collaborations in the field.<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/shaik8558834\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_An81zCg.webp\" width=\"48\" height=\"48\" alt=\"Shaik Hamzah\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>GenAI Intern @ Analytics Vidhya | Final Year @ VIT Chennai<br \/>Passionate about AI and machine learning, I&#8217;m eager to dive into roles as an AI\/ML Engineer or Data Scientist where I can make a real impact. With a knack for quick learning and a love for teamwork, I&#8217;m excited to bring innovative solutions and cutting-edge advancements to the table. My curiosity drives me to explore AI across various fields and take the initiative to delve into data engineering, ensuring I stay ahead and deliver impactful projects.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Digital documents have long presented a dual challenge for both human readers and automated systems: preserving rich structural nuances while converting content into machine-processable formats. Traditional methods, whether relying on complex ensemble pipelines or massive foundational models, often struggle to balance accuracy with computational efficiency. SmolDocling emerges as a game-changing solution, offering an ultra-compact 256M-parameter [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":147787,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[10635,20668,62083,62082],"dealstore":[],"offerexpiration":[],"class_list":["post-147786","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-document","tag-efficient","tag-parsing","tag-smoldocling"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Can SmolDocling Make Document Parsing More Efficient? - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=147786\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Can SmolDocling Make Document Parsing More Efficient? - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Digital documents have long presented a dual challenge for both human readers and automated systems: preserving rich structural nuances while converting content into machine-processable formats. Traditional methods, whether relying on complex ensemble pipelines or massive foundational models, often struggle to balance accuracy with computational efficiency. SmolDocling emerges as a game-changing solution, offering an ultra-compact 256M-parameter [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=147786\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-03-21T13:50:03+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"9 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=147786#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=147786\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Can SmolDocling Make Document Parsing More Efficient?\",\"datePublished\":\"2025-03-21T13:50:03+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=147786\"},\"wordCount\":833,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=147786#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp\",\"keywords\":[\"Document\",\"Efficient\",\"Parsing\",\"SmolDocling\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=147786#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=147786\",\"url\":\"https:\/\/fivemor.com\/?p=147786\",\"name\":\"Can SmolDocling Make Document Parsing More Efficient? - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=147786#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=147786#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp\",\"datePublished\":\"2025-03-21T13:50:03+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=147786#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=147786\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=147786#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=147786#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Can SmolDocling Make Document Parsing More Efficient?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Can SmolDocling Make Document Parsing More Efficient? - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=147786","og_locale":"en_US","og_type":"article","og_title":"Can SmolDocling Make Document Parsing More Efficient? - Som2ny Network","og_description":"Digital documents have long presented a dual challenge for both human readers and automated systems: preserving rich structural nuances while converting content into machine-processable formats. Traditional methods, whether relying on complex ensemble pipelines or massive foundational models, often struggle to balance accuracy with computational efficiency. SmolDocling emerges as a game-changing solution, offering an ultra-compact 256M-parameter [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=147786","og_site_name":"Som2ny Network","article_published_time":"2025-03-21T13:50:03+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"9 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=147786#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=147786"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Can SmolDocling Make Document Parsing More Efficient?","datePublished":"2025-03-21T13:50:03+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=147786"},"wordCount":833,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=147786#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp","keywords":["Document","Efficient","Parsing","SmolDocling"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=147786#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=147786","url":"https:\/\/fivemor.com\/?p=147786","name":"Can SmolDocling Make Document Parsing More Efficient? - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=147786#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=147786#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp","datePublished":"2025-03-21T13:50:03+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=147786#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=147786"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=147786#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/SmolDocling-High-Accuracy-Document-Parsing-with-a-Small-Footprint-.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=147786#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Can SmolDocling Make Document Parsing More Efficient?"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/147786","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=147786"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/147786\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/147787"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=147786"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=147786"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=147786"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=147786"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=147786"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}