{"id":196577,"date":"2025-04-21T12:14:50","date_gmt":"2025-04-21T12:14:50","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/14-powerful-techniques-defining-the-evolution-of-embedding\/"},"modified":"2025-04-21T12:14:50","modified_gmt":"2025-04-21T12:14:50","slug":"14-powerful-techniques-defining-the-evolution-of-embedding","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=196577","title":{"rendered":"14 Powerful Techniques Defining the Evolution of Embedding"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p><strong>Summary:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><em>Evolution of Embeddings from basic count-based methods (TF-IDF, Word2Vec) to context-aware models like BERT and ELMo, which capture nuanced semantics by analyzing entire sentences bidirectionally.<\/em><\/li>\n<li><em>Leaderboards such as MTEB benchmark embeddings for tasks like retrieval and classification.<\/em><\/li>\n<li><em>Open-source platforms (Hugging Face) allow developers to access cutting-edge embeddings and deploy models tailored to different use cases.<\/em><\/li>\n<\/ul>\n<p>You know how, back in the day, we used simple word\u2010count tricks to represent text? Well, things have come a long way since then. Now, when we talk about the evolution of embeddings, we mean numerical snapshots that capture not just which words appear but what they really mean, how they relate to each other in context, and even how they tie into images and other media. Embeddings power everything from search engines that understand your intent to recommendation systems that seem to read your mind. They\u2019re at the heart of cutting\u2010edge AI and machine\u2010learning applications, too. So, let\u2019s take a stroll through this evolution from raw counts to semantic vectors, exploring how each approach works, what it brings to the table, and where it falls short.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-ranking-of-embeddings-in-mteb-leaderboards\">Ranking of Embeddings in MTEB Leaderboards <\/h2>\n<p>Most modern LLMs generate embeddings as intermediate outputs of their architectures. These can be extracted and fine-tuned for various downstream tasks, making LLM-based embeddings one of the most versatile tools available today.<\/p>\n<p>To keep up with the fast-moving landscape, platforms like <a href=\"https:\/\/huggingface.co\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Hugging Face<\/a> have introduced resources like the Massive Text Embedding Benchmark (MTEB) Leaderboard. This leaderboard ranks embedding models based on their performance across a wide range of tasks, including classification, clustering, retrieval, and more. This is substantially helping practitioners identify the best models for their use cases.<\/p>\n<div class=\"wp-block-image figure  mt-2 mb-2 d-table mx-auto\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1562\" height=\"712\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/02-image9.webp\" alt=\"Ranking of Embeddings in MTEB Leaderboards\" class=\"wp-image-231252\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/02-image9.webp 1562w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/02-image9-300x137.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/02-image9-768x350.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/02-image9-1536x700.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/02-image9-150x68.webp 150w\" sizes=\"auto, (max-width: 1562px) 100vw, 1562px\"\/><\/figure>\n<\/div>\n<p>Armed with these leaderboard insights, let\u2019s roll up our sleeves and dive into the vectorization toolbox \u2013 count vectors, TF\u2013IDF, and other classic methods, which still serve as the essential building blocks for today\u2019s sophisticated embeddings.<\/p>\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1386\" height=\"568\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/03-image39.webp\" alt=\"Ranking of Embeddings in MTEB Leaderboards\" class=\"wp-image-231253\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/03-image39.webp 1386w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/03-image39-300x123.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/03-image39-768x315.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/03-image39-150x61.webp 150w\" sizes=\"auto, (max-width: 1386px) 100vw, 1386px\"\/><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-1-count-vectorization\">1. Count Vectorization<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2022\/09\/implementing-count-vectorizer-and-tf-idf-in-nlp-using-pyspark\/\" target=\"_blank\" rel=\"noreferrer noopener\">Count Vectorization<\/a> is one of the simplest techniques for representing text. It emerged from the need to convert raw text into numerical form so that machine learning models could process it. In this method, each document is transformed into a vector that reflects the count of each word appearing in it. This straightforward approach laid the groundwork for more complex representations and is still useful in scenarios where interpretability is key.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li>The text corpus is first tokenized into words. A vocabulary is built from all unique tokens.<\/li>\n<li>Each document is represented as a vector where each dimension corresponds to the word\u2019s respective vector in the vocabulary.<\/li>\n<li>The value in each dimension is simply the frequency or count of a certain word in the document.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Example:<\/strong> For a vocabulary [\u201c<em>apple<\/em>\u201c, \u201c<em>banana<\/em>\u201c, \u201c<em>cherry<\/em>\u201c], the document \u201c<em>apple apple cherry<\/em>\u201d becomes [<em>2, 0, 1<\/em>].<\/li>\n<li><strong>Additional Detail:<\/strong> Count Vectorization serves as the foundation for many other approaches. Its simplicity does not capture any contextual or semantic information, but it remains an essential preprocessing step in many NLP pipelines.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>from sklearn.feature_extraction.text import CountVectorizer\n\nimport pandas as pd\n\n# Sample text documents with repeated words\n\ndocuments = [\n\n\t\"Natural Language Processing is fun and natural natural natural\",\n\n\t\"I really love love love Natural Language Processing Processing Processing\",\n\n\t\"Machine Learning is a part of AI AI AI AI\",\n\n\t\"AI and NLP NLP NLP are closely related related\"\n\n]\n\n# Initialize CountVectorizer\n\nvectorizer = CountVectorizer()\n\n# Fit and transform the text data\n\nX = vectorizer.fit_transform(documents)\n\n# Get feature names (unique words)\n\nfeature_names = vectorizer.get_feature_names_out()\n\n# Convert to DataFrame for better visualization\n\ndf = pd.DataFrame(X.toarray(), columns=feature_names)\n\n# Print the matrix\n\nprint(df)<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1078\" height=\"165\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/04-image24.webp\" alt=\"Count Vectorization Output\" class=\"wp-image-231254\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/04-image24.webp 1078w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/04-image24-300x46.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/04-image24-768x118.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/04-image24-150x23.webp 150w\" sizes=\"auto, (max-width: 1078px) 100vw, 1078px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Simplicity and Interpretability:<\/strong> Easy to implement and understand.<\/li>\n<li><strong>Deterministic:<\/strong> Produces a fixed representation that is easy to analyze.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>High Dimensionality and Sparsity:<\/strong> Vectors are often large and mostly zero, leading to inefficiencies.<\/li>\n<li><strong>Lack of Semantic Context:<\/strong> Does not capture meaning or relationships between words.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-2-one-hot-encoding\">2. One-Hot Encoding<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/12\/how-to-do-one-hot-encoding\/\" target=\"_blank\" rel=\"noreferrer noopener\">One-hot encoding<\/a> is one of the earliest approaches to representing words as vectors. Developed alongside early digital computing techniques in the 1950s and 1960s, it transforms categorical data, such as words, into binary vectors. Each word is represented uniquely, ensuring that no two words share similar representations, though this comes at the expense of capturing semantic similarity.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-0\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li>Every word in the vocabulary is assigned a vector whose length equals the size of the vocabulary.<\/li>\n<li>In each vector, all elements are 0 except for a single 1 in the position corresponding to that word.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Example: <\/strong>With a vocabulary [\u201c<em>apple<\/em>\u201c, \u201c<em>banana<\/em>\u201c, \u201c<em>cherry<\/em>\u201c], the word \u201c<em>banana<\/em>\u201d is represented as [<em>0, 1, 0<\/em>].<\/li>\n<li><strong>Additional Detail: <\/strong>One-hot vectors are completely orthogonal, which means that the cosine similarity between two different words is zero. This approach is simple and unambiguous but fails to capture any similarity (e.g., \u201capple\u201d and \u201corange\u201d appear equally dissimilar to \u201capple\u201d and \u201ccar\u201d).<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-0\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>from sklearn.feature_extraction.text import CountVectorizer\n\nimport pandas as pd\n\n# Sample text documents\n\ndocuments = [\n\n\u00a0\u00a0\u00a0\"Natural Language Processing is fun and natural natural natural\",\n\n\u00a0\u00a0\u00a0\"I really love love love Natural Language Processing Processing Processing\",\n\n\u00a0\u00a0\u00a0\"Machine Learning is a part of AI AI AI AI\",\n\n\u00a0\u00a0\u00a0\"AI and NLP NLP NLP are closely related related\"\n\n]\n\n# Initialize CountVectorizer with binary=True for One-Hot Encoding\n\nvectorizer = CountVectorizer(binary=True)\n\n# Fit and transform the text data\n\nX = vectorizer.fit_transform(documents)\n\n# Get feature names (unique words)\n\nfeature_names = vectorizer.get_feature_names_out()\n\n# Convert to DataFrame for better visualization\n\ndf = pd.DataFrame(X.toarray(), columns=feature_names)\n\n# Print the one-hot encoded matrix\n\nprint(df)<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1081\" height=\"172\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/05-image19.webp\" alt=\"One-Hot Encoding Output\" class=\"wp-image-231255\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/05-image19.webp 1081w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/05-image19-300x48.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/05-image19-768x122.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/05-image19-150x24.webp 150w\" sizes=\"auto, (max-width: 1081px) 100vw, 1081px\"\/><\/figure>\n<\/div>\n<p>So, basically, you can view the difference between Count Vectorizer and One Hot Encoding. Count Vectorizer counts how many times a certain word exists in a sentence, whereas One Hot Encoding labels the word as 1 if it exists in a certain sentence\/document.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"742\" height=\"285\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/06-image22.webp\" alt=\"One-Hot Encoding \" class=\"wp-image-231256\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/06-image22.webp 742w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/06-image22-300x115.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/06-image22-150x58.webp 150w\" sizes=\"auto, (max-width: 742px) 100vw, 742px\"\/><\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\" id=\"h-when-to-use-what\">When to Use What?<\/h4>\n<ul class=\"wp-block-list\">\n<li>Use <strong>CountVectorizer<\/strong> when the number of times a word appears is important (e.g., spam detection, document similarity).<\/li>\n<li>Use <strong>One-Hot Encoding<\/strong> when you only care about whether a word appears at least once (e.g., categorical feature encoding for ML models).<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-0\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Clarity and Uniqueness:<\/strong> Each word has a distinct and non-overlapping representation<\/li>\n<li><strong>Simplicity:<\/strong> Easy to implement with minimal computational overhead for small vocabularies.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-0\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Inefficiency with Large Vocabularies:<\/strong> Vectors become extremely high-dimensional and sparse.<\/li>\n<li><strong>No Semantic Similarity:<\/strong> Does not allow for any relationships between words; all non-identical words are equally distant.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-3-tf-idf-term-frequency-inverse-document-frequency\">3. TF-IDF (Term Frequency-Inverse Document Frequency)<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2020\/02\/quick-introduction-bag-of-words-bow-tf-idf\/\" target=\"_blank\" rel=\"noreferrer noopener\">TF-IDF<\/a> was developed to improve upon raw count methods by counting word occurrences and weighing words based on their overall importance in a corpus. Introduced in the early 1970s, TF-IDF is a cornerstone in information retrieval systems and text mining applications. It helps highlight terms that are significant in individual documents while downplaying words that are common across all documents.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-1\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Term Frequency (TF):<\/strong> Measures how often a word appears in a document.<\/li>\n<li><strong>Inverse Document Frequency (IDF):<\/strong> Scales the importance of a word by considering how common or rare it is across all documents.<\/li>\n<li>The final TF-IDF score is the product of TF and IDF.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<ul class=\"wp-block-list\">\n<li><strong>Example:<\/strong> Common words like \u201cthe\u201d receive low scores, whereas more unique words receive higher scores, making them stand out in document analysis. Hence, we normally omit the frequent terms, which are also called Stopwords, in NLP tasks.<\/li>\n<li><strong>Additional Detail:<\/strong> TF-IDF transforms raw frequency counts into a measure that can effectively differentiate between important keywords and commonly used words. It has become a standard method in search engines and document clustering.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-1\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>from sklearn.feature_extraction.text import TfidfVectorizer\n\nimport pandas as pd\n\nimport numpy as np\n\n# Sample short sentences\n\ndocuments = [\n\n\u00a0\u00a0\u00a0\"cat sits here\",\n\n\u00a0\u00a0\u00a0\"dog barks loud\",\n\n\u00a0\u00a0\u00a0\"cat barks loud\"\n\n]\n\n# Initialize TfidfVectorizer to get both TF and IDF\n\nvectorizer = TfidfVectorizer()\n\n# Fit and transform the text data\n\nX = vectorizer.fit_transform(documents)\n\n# Extract feature names (unique words)\n\nfeature_names = vectorizer.get_feature_names_out()\n\n# Get TF matrix (raw term frequencies)\n\ntf_matrix = X.toarray()\n\n# Compute IDF values manually\n\nidf_values = vectorizer.idf_\n\n# Compute TF-IDF manually (TF * IDF)\n\ntfidf_matrix = tf_matrix * idf_values\n\n# Convert to DataFrames for better visualization\n\ndf_tf = pd.DataFrame(tf_matrix, columns=feature_names)\n\ndf_idf = pd.DataFrame([idf_values], columns=feature_names)\n\ndf_tfidf = pd.DataFrame(tfidf_matrix, columns=feature_names)\n\n# Print tables\n\nprint(\"\\n\ud83d\udd39 Term Frequency (TF) Matrix:\\n\", df_tf)\n\nprint(\"\\n\ud83d\udd39 Inverse Document Frequency (IDF) Values:\\n\", df_idf)\n\nprint(\"\\n\ud83d\udd39 TF-IDF Matrix (TF * IDF):\\n\", df_tfidf)<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"414\" height=\"295\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-image40.webp\" alt=\"Evolution of Embeddings\" class=\"wp-image-231259\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-image40.webp 414w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-image40-300x214.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/10-image40-150x107.webp 150w\" sizes=\"auto, (max-width: 414px) 100vw, 414px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-1\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Enhanced Word Importance:<\/strong> Emphasizes content-specific words.<\/li>\n<li><strong>Reduces Dimensionality:<\/strong> Filters out common words that add little value.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-1\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Sparse Representation:<\/strong> Despite weighting, the resulting vectors are still sparse.<\/li>\n<li><strong>Lack of Context:<\/strong> Does not capture word order or deeper semantic relationships.<\/li>\n<\/ul>\n<p>Also Read:<a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2022\/09\/implementing-count-vectorizer-and-tf-idf-in-nlp-using-pyspark\/\" target=\"_blank\" rel=\"noreferrer noopener\"> Implementing Count Vectorizer and TF-IDF in NLP using PySpark<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-4-okapi-bm25\">4. Okapi BM25<\/h2>\n<p>Okapi BM25, developed in the 1990s, is a probabilistic model designed primarily for ranking documents in information retrieval systems rather than as an embedding method per se. BM25 is an enhanced version of TF-IDF, commonly used in search engines and information retrieval. It improves upon TF-IDF by considering document length normalization and saturation of term frequency (i.e., diminishing returns for repeated words).<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-2\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Probabilistic Framework:<\/strong> This framework estimates the relevance of a document based on the frequency of query terms, adjusted by document length.<\/li>\n<li>Uses parameters to control the influence of term frequency and to dampen the effect of very high counts.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p>Here we will be looking into the BM25 scoring mechanism:<\/p>\n<p>BM25 introduces two parameters, k1 and b, which allow fine-tuning of the term frequency saturation and the length normalization, respectively. These parameters are crucial for optimizing the BM25 algorithm\u2019s performance in various search contexts.<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Example:<\/strong> BM25 assigns higher relevance scores to documents that contain rare query terms with moderate frequency while adjusting for document length and vice versa.<\/li>\n<li><strong>Additional Detail:<\/strong> Although BM25 does not produce vector embeddings, it has deeply influenced text retrieval systems by improving upon the shortcomings of TF-IDF in ranking documents.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-2\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>import numpy as np\n\nimport pandas as pd\n\nfrom sklearn.feature_extraction.text import CountVectorizer\n\n# Sample documents\n\ndocuments = [\n\n\u00a0\u00a0\u00a0\"cat sits here\",\n\n\u00a0\u00a0\u00a0\"dog barks loud\",\n\n\u00a0\u00a0\u00a0\"cat barks loud\"\n\n]\n\n# Compute Term Frequency (TF) using CountVectorizer\n\nvectorizer = CountVectorizer()\n\nX = vectorizer.fit_transform(documents)\n\ntf_matrix = X.toarray()\n\nfeature_names = vectorizer.get_feature_names_out()\n\n# Compute Inverse Document Frequency (IDF) for BM25\n\nN = len(documents)\u00a0 # Total number of documents\n\ndf = np.sum(tf_matrix &gt; 0, axis=0)\u00a0 # Document Frequency (DF) for each term\n\nidf = np.log((N - df + 0.5) \/ (df + 0.5) + 1)\u00a0 # BM25 IDF formula\n\n# Compute BM25 scores\n\nk1 = 1.5\u00a0 # Smoothing parameter\n\nb = 0.75\u00a0 # Length normalization parameter\n\navgdl = np.mean([len(doc.split()) for doc in documents])\u00a0 # Average document length\n\ndoc_lengths = np.array([len(doc.split()) for doc in documents])\n\nbm25_matrix = np.zeros_like(tf_matrix, dtype=np.float64)\n\nfor i in range(N):\u00a0 # For each document\n\n\u00a0\u00a0\u00a0for j in range(len(feature_names)):\u00a0 # For each term\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0term_freq = tf_matrix[i, j]\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0num = term_freq * (k1 + 1)\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0denom = term_freq + k1 * (1 - b + b * (doc_lengths[i] \/ avgdl))\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0bm25_matrix[i, j] = idf[j] * (num \/ denom)\n\n# Convert to DataFrame for better visualization\n\ndf_tf = pd.DataFrame(tf_matrix, columns=feature_names)\n\ndf_idf = pd.DataFrame([idf], columns=feature_names)\n\ndf_bm25 = pd.DataFrame(bm25_matrix, columns=feature_names)\n\n# Display the results\n\nprint(\"\\n\ud83d\udd39 Term Frequency (TF) Matrix:\\n\", df_tf)\n\nprint(\"\\n\ud83d\udd39 BM25 Inverse Document Frequency (IDF):\\n\", df_idf)\n\nprint(\"\\n\ud83d\udd39 BM25 Scores:\\n\", df_bm25)<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"536\" height=\"267\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-image13.webp\" alt=\"BN 25 Output\" class=\"wp-image-231262\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-image13.webp 536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-image13-300x149.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/13-image13-150x75.webp 150w\" sizes=\"auto, (max-width: 536px) 100vw, 536px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-info-retrieval\">Code Implementation (Info Retrieval)<\/h3>\n<pre class=\"wp-block-code\"><code>!pip install bm25s\n\nimport bm25s\n\n# Create your corpus here\n\ncorpus = [\n\n\u00a0\u00a0\u00a0\"a cat is a feline and likes to purr\",\n\n\u00a0\u00a0\u00a0\"a dog is the human's best friend and loves to play\",\n\n\u00a0\u00a0\u00a0\"a bird is a beautiful animal that can fly\",\n\n\u00a0\u00a0\u00a0\"a fish is a creature that lives in water and swims\",\n\n]\n\n# Create the BM25 model and index the corpus\n\nretriever = bm25s.BM25(corpus=corpus)\n\nretriever.index(bm25s.tokenize(corpus))\n\n# Query the corpus and get top-k results\n\nquery = \"does the fish purr like a cat?\"\n\nresults, scores = retriever.retrieve(bm25s.tokenize(query), k=2)\n\n# Let's see what we got!\n\ndoc, score = results[0, 0], scores[0, 0]\n\nprint(f\"Rank {i+1} (score: {score:.2f}): {doc}\")<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"494\" height=\"45\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-image11.webp\" alt=\"BN25 Output\" class=\"wp-image-231263\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-image11.webp 494w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-image11-300x27.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/14-image11-150x14.webp 150w\" sizes=\"auto, (max-width: 494px) 100vw, 494px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-2\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Improved Relevance Ranking:<\/strong> Better handles document length and term saturation.<\/li>\n<li><strong>Widely Adopted:<\/strong> Standard in many modern search engines and IR systems.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-2\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Not a True Embedding:<\/strong> It scores documents rather than producing a continuous vector space representation.<\/li>\n<li><strong>Parameter Sensitivity:<\/strong> Requires careful tuning for optimal performance.<\/li>\n<\/ul>\n<p>Also Read: <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2021\/05\/build-your-own-nlp-based-search-engine-using-bm25\/\" target=\"_blank\" rel=\"noreferrer noopener\">How to Create NLP Search Engine With BM25?<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-5-word2vec-cbow-and-skip-gram\">5. Word2Vec (CBOW and Skip-gram)<\/h2>\n<p>Introduced by Google in 2013, <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2021\/07\/word2vec-for-word-embeddings-a-beginners-guide\/\" target=\"_blank\" rel=\"noreferrer noopener\">Word2Vec<\/a> revolutionized NLP by learning dense, low-dimensional vector representations of words. It moved beyond counting and weighting by training shallow neural networks that capture semantic and syntactic relationships based on word context. Word2Vec comes in two flavors: Continuous Bag-of-Words (CBOW) and Skip-gram.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-3\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>CBOW (Continuous Bag-of-Words):<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong> Predicts a target word based on the surrounding context words.<\/li>\n<li><strong>Process:<\/strong> Takes multiple context words (ignoring the order) and learns to predict the central word.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Skip-gram:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong> Uses the target word to predict its surrounding context words.<\/li>\n<li><strong>Process:<\/strong> Particularly effective for learning representations of rare words by focusing on their contexts.<br \/><img decoding=\"async\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/15-image27.webp\" alt=\"Evolution of Embeddings\"\/><\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<\/strong> Both architectures use a neural network with one hidden layer and employ optimization tricks such as negative sampling or hierarchical softmax to manage computational complexity. The resulting embeddings capture nuanced semantic relationships for instance, \u201cking\u201d minus \u201cman\u201d plus \u201cwoman\u201d approximates \u201cqueen.\u201d<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-3\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>!pip install numpy==1.24.3\n\nfrom gensim.models import Word2Vec\n\nimport networkx as nx\n\nimport matplotlib.pyplot as plt\n\n# Sample corpus\n\nsentences = [\n\n\t[\"I\", \"love\", \"deep\", \"learning\"],\n\n\t[\"Natural\", \"language\", \"processing\", \"is\", \"fun\"],\n\n\t[\"Word2Vec\", \"is\", \"a\", \"great\", \"tool\"],\n\n\t[\"AI\", \"is\", \"the\", \"future\"],\n\n]\n\n# Train Word2Vec models\n\ncbow_model = Word2Vec(sentences, vector_size=10, window=2, min_count=1, sg=0)\u00a0 # CBOW\n\nskipgram_model = Word2Vec(sentences, vector_size=10, window=2, min_count=1, sg=1)\u00a0 # Skip-gram\n\n# Get word vectors\n\nword = \"is\"\n\nprint(f\"CBOW Vector for '{word}':\\n\", cbow_model.wv[word])\n\nprint(f\"\\nSkip-gram Vector for '{word}':\\n\", skipgram_model.wv[word])\n\n# Get most similar words\n\nprint(\"\\n\ud83d\udd39 CBOW Most Similar Words:\", cbow_model.wv.most_similar(word))\n\nprint(\"\\n\ud83d\udd39 Skip-gram Most Similar Words:\", skipgram_model.wv.most_similar(word))\n<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1093\" height=\"204\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/17-image21.webp\" alt=\"Word2vec Output\" class=\"wp-image-231266\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/17-image21.webp 1093w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/17-image21-300x56.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/17-image21-768x143.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/17-image21-150x28.webp 150w\" sizes=\"auto, (max-width: 1093px) 100vw, 1093px\"\/><\/figure>\n<\/div>\n<p>Visualizing the CBOW and Skip-gram:<\/p>\n<pre class=\"wp-block-code\"><code>def visualize_cbow():\n\n\u00a0\u00a0\u00a0G = nx.DiGraph()\n\n\u00a0\u00a0\u00a0# Nodes\n\n\u00a0\u00a0\u00a0context_words = [\"Natural\", \"is\", \"fun\"]\n\n\u00a0\u00a0\u00a0target_word = \"learning\"\n\n\u00a0\u00a0\u00a0for word in context_words:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0G.add_edge(word, \"Hidden Layer\")\n\n\u00a0\u00a0\u00a0G.add_edge(\"Hidden Layer\", target_word)\n\n\u00a0\u00a0\u00a0# Draw the network\n\n\u00a0\u00a0\u00a0pos = nx.spring_layout(G)\n\n\u00a0\u00a0\u00a0plt.figure(figsize=(6, 4))\n\n\u00a0\u00a0\u00a0nx.draw(G, pos, with_labels=True, node_size=3000, node_color=\"lightblue\", edge_color=\"gray\")\n\n\u00a0\u00a0\u00a0plt.title(\"CBOW Model Visualization\")\n\n\u00a0\u00a0\u00a0plt.show()\n\nvisualize_cbow()<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"619\" height=\"442\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/18-image18.webp\" alt=\"CBOW Model Visualization\" class=\"wp-image-231267\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/18-image18.webp 619w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/18-image18-300x214.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/18-image18-150x107.webp 150w\" sizes=\"auto, (max-width: 619px) 100vw, 619px\"\/><\/figure>\n<\/div>\n<pre class=\"wp-block-code\"><code>def visualize_skipgram():\n\n\u00a0\u00a0\u00a0G = nx.DiGraph()\n\n\u00a0\u00a0\u00a0# Nodes\n\n\u00a0\u00a0\u00a0target_word = \"learning\"\n\n\u00a0\u00a0\u00a0context_words = [\"Natural\", \"is\", \"fun\"]\n\n\u00a0\u00a0\u00a0G.add_edge(target_word, \"Hidden Layer\")\n\n\u00a0\u00a0\u00a0for word in context_words:\n\n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0G.add_edge(\"Hidden Layer\", word)\n\n\u00a0\u00a0\u00a0# Draw the network\n\n\u00a0\u00a0\u00a0pos = nx.spring_layout(G)\n\n\u00a0\u00a0\u00a0plt.figure(figsize=(6, 4))\n\n\u00a0\u00a0\u00a0nx.draw(G, pos, with_labels=True, node_size=3000, node_color=\"lightgreen\", edge_color=\"gray\")\n\n\u00a0\u00a0\u00a0plt.title(\"Skip-gram Model Visualization\")\n\n\u00a0\u00a0\u00a0plt.show()\n\nvisualize_skipgram()<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"619\" height=\"442\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/19-image5.webp\" alt=\"Skip-gram Model Visualization\" class=\"wp-image-231268\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/19-image5.webp 619w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/19-image5-300x214.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/19-image5-150x107.webp 150w\" sizes=\"auto, (max-width: 619px) 100vw, 619px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-3\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Semantic Richness:<\/strong> Learns meaningful relationships between words.<\/li>\n<li><strong>Efficient Training:<\/strong> Can be trained on large corpora relatively quickly.<\/li>\n<li><strong>Dense Representations:<\/strong> Uses low-dimensional, continuous vectors that facilitate downstream processing.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-3\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Static Representations:<\/strong> Provides one embedding per word regardless of context.<\/li>\n<li><strong>Context Limitations:<\/strong> Cannot disambiguate polysemous words that have different meanings in different contexts.<\/li>\n<\/ul>\n<p>To read more about Word2Vec read <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2021\/07\/word2vec-for-word-embeddings-a-beginners-guide\/\">this<\/a> blog.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-6-glove-global-vectors-for-word-representation\">6. GloVe (Global Vectors for Word Representation)<\/h2>\n<p>GloVe, developed at Stanford in 2014, builds on the ideas of Word2Vec by combining global co-occurrence statistics with local context information. It was designed to produce word embeddings that capture overall corpus-level statistics, offering improved consistency across different contexts.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-4\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Co-occurrence Matrix:<\/strong> Constructs a matrix capturing how frequently pairs of words appear together across the entire corpus.<br \/><img decoding=\"async\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/20-image28.webp\" alt=\"\"\/>\n<p>This logic of Co-occurence matrices are also widely used in Computer Vision too, especially under the topic of GLCM(Gray-Level Co-occurrence Matrix). It is a statistical method used in image processing and computer vision for texture analysis that considers the spatial relationship between pixels.<\/p>\n<\/li>\n<li><strong>Matrix Factorization:<\/strong> Factorizes this matrix to derive word vectors that capture global statistical information.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<ul class=\"wp-block-list\">\n<li><strong>Additional Detail:<br \/><\/strong>Unlike Word2Vec\u2019s purely predictive model, GloVe\u2019s approach allows the model to learn the ratios of word co-occurrences, which some studies have found to be more robust in capturing semantic similarities and analogies.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-4\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>import numpy as np\n\n# Load pre-trained GloVe embeddings\n\nglove_model = api.load(\"glove-wiki-gigaword-50\")\u00a0 # You can use \"glove-twitter-25\", \"glove-wiki-gigaword-100\", etc.\n\n# Example words\n\nword = \"king\"\n\nprint(f\"\ud83d\udd39 Vector representation for '{word}':\\n\", glove_model[word])\n\n# Find similar words\n\nsimilar_words = glove_model.most_similar(word, topn=5)\n\nprint(\"\\n\ud83d\udd39 Words similar to 'king':\", similar_words)\n\nword1 = \"king\"\n\nword2 = \"queen\"\n\nsimilarity = glove_model.similarity(word1, word2)\n\nprint(f\"\ud83d\udd39 Similarity between '{word1}' and '{word2}': {similarity:.4f}\")<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1527\" height=\"227\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/21-image36.webp\" alt=\"GloVe (Global Vectors for Word Representation)\" class=\"wp-image-231269\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/21-image36.webp 1527w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/21-image36-300x45.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/21-image36-768x114.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/21-image36-150x22.webp 150w\" sizes=\"auto, (max-width: 1527px) 100vw, 1527px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"416\" height=\"26\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/22-image23.webp\" alt=\"GloVe (Global Vectors for Word Representation) | Evolution of Embeddings\" class=\"wp-image-231270\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/22-image23.webp 416w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/22-image23-300x19.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/22-image23-150x9.webp 150w\" sizes=\"auto, (max-width: 416px) 100vw, 416px\"\/><\/figure>\n<\/div>\n<p>This image will help you understand how this similarity looks like when plotted:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"314\" height=\"268\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/23-image8.webp\" alt=\"GloVe (Global Vectors for Word Representation)\" class=\"wp-image-231271\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/23-image8.webp 314w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/23-image8-300x256.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/23-image8-150x128.webp 150w\" sizes=\"auto, (max-width: 314px) 100vw, 314px\"\/><\/figure>\n<\/div>\n<p>Do refer to <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2017\/06\/word-embeddings-count-word2veec\/\">this<\/a> for more in-depth information.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-4\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Global Context Integration:<\/strong> Uses entire corpus statistics to improve representation.<\/li>\n<li><strong>Stability:<\/strong> Often yields more consistent embeddings across different contexts.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-4\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Resource Demanding:<\/strong> Building and factorizing large matrices can be computationally expensive.<\/li>\n<li><strong>Static Nature:<\/strong> Similar to Word2Vec, it does not generate context-dependent embeddings.<\/li>\n<\/ul>\n<p>GloVe learns embeddings from word co-occurrence matrices.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-7-fasttext\">7. FastText<\/h2>\n<p>FastText, released by Facebook in 2016, extends Word2Vec by incorporating subword (character n-gram) information. This innovation helps the model handle rare words and morphologically rich languages by breaking words down into smaller units, thereby capturing internal structure.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-5\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Subword Modeling:<\/strong> Represents each word as a sum of its character n-gram vectors.<\/li>\n<li><strong>Embedding Learning:<\/strong> Trains a model that uses these subword vectors to produce a final word embedding.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<br \/><\/strong>This method is particularly useful for languages with rich morphology and for dealing with out-of-vocabulary words. By decomposing words, FastText can generalize better across similar word forms and misspellings.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-5\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>import gensim.downloader as api\n\nfasttext_model = api.load(\"fasttext-wiki-news-subwords-300\")\n\n# Example word\n\nword = \"king\"\n\nprint(f\"\ud83d\udd39 Vector representation for '{word}':\\n\", fasttext_model[word])\n\n# Find similar words\n\nsimilar_words = fasttext_model.most_similar(word, topn=5)\n\nprint(\"\\n\ud83d\udd39 Words similar to 'king':\", similar_words)\n\nword1 = \"king\"\n\nword2 = \"queen\"\n\nsimilarity = fasttext_model.similarity(word1, word2)\n\nprint(f\"\ud83d\udd39 Similarity between '{word1}' and '{word2}': {similarity:.4f}\")<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"725\" height=\"301\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/24-image20.webp\" alt=\"FastText | Evolution of Embeddings\" class=\"wp-image-231272\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/24-image20.webp 725w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/24-image20-300x125.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/24-image20-150x62.webp 150w\" sizes=\"auto, (max-width: 725px) 100vw, 725px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1580\" height=\"151\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/25-image34.webp\" alt=\"FastText | Evolution of Embeddings\" class=\"wp-image-231273\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/25-image34.webp 1580w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/25-image34-300x29.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/25-image34-768x73.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/25-image34-1536x147.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/25-image34-150x14.webp 150w\" sizes=\"auto, (max-width: 1580px) 100vw, 1580px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"413\" height=\"29\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/26-image2.webp\" alt=\"FastText | Evolution of Embeddings\" class=\"wp-image-231274\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/26-image2.webp 413w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/26-image2-300x21.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/26-image2-150x11.webp 150w\" sizes=\"auto, (max-width: 413px) 100vw, 413px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-5\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Handling OOV(Out of Vocabulary) Words:<\/strong> Improves performance when words are infrequent or unseen. Can say that the test dataset has some labels which do not exist in our train dataset.<\/li>\n<li><strong>Morphological Awareness:<\/strong> Captures the internal structure of words.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-5\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Increased Complexity:<\/strong> The inclusion of subword information adds to computational overhead.<\/li>\n<li><strong>Still Static or Fixed:<\/strong> Despite the improvements, FastText does not adjust embeddings based on a sentence\u2019s surrounding context.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-8-doc2vec\">8. Doc2Vec<\/h2>\n<p>Doc2Vec extends Word2Vec\u2019s ideas to larger bodies of text, such as sentences, paragraphs, or entire documents. Introduced in 2014, it provides a means to obtain fixed-length vector representations for variable-length texts, enabling more effective document classification, clustering, and retrieval.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-6\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism<\/strong>:\n<ul class=\"wp-block-list\">\n<li><strong>Distributed Memory (DM) Model:<\/strong> Augments the Word2Vec architecture by adding a unique document vector that, along with context words, predicts a target word.<\/li>\n<li><strong>Distributed Bag-of-Words (DBOW) Model:<\/strong> Learns document vectors by predicting words randomly sampled from the document.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<br \/><\/strong>These models learn document-level embeddings that capture the overall semantic content of the text. They are especially useful for tasks where the structure and theme of the entire document are important.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-6\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>import gensim\n\nfrom gensim.models.doc2vec import Doc2Vec, TaggedDocument\n\nimport nltk\n\nnltk.download('punkt_tab')\n\n# Sample documents\n\ndocuments = [\n\n\t\"Machine learning is amazing\",\n\n\t\"Natural language processing enables AI to understand text\",\n\n\t\"Deep learning advances artificial intelligence\",\n\n\t\"Word embeddings improve NLP tasks\",\n\n\t\"Doc2Vec is an extension of Word2Vec\"\n\n]\n\n# Tokenize and tag documents\n\ntagged_data = [TaggedDocument(words=nltk.word_tokenize(doc.lower()), tags=[str(i)]) for i, doc in enumerate(documents)]\n\n# Print tagged data\n\nprint(tagged_data)\n\n# Define model parameters\n\nmodel = Doc2Vec(vector_size=50, window=2, min_count=1, workers=4, epochs=100)\n\n# Build vocabulary\n\nmodel.build_vocab(tagged_data)\n\n# Train the model\n\nmodel.train(tagged_data, total_examples=model.corpus_count, epochs=model.epochs)\n\n# Test a document by generating its vector\n\ntest_doc = \"Artificial intelligence uses machine learning\"\n\ntest_vector = model.infer_vector(nltk.word_tokenize(test_doc.lower()))\n\nprint(f\"\ud83d\udd39 Vector representation of test document:\\n{test_vector}\")\n\n# Find most similar documents to the test document\n\nsimilar_docs = model.dv.most_similar([test_vector], topn=3)\n\nprint(\"\ud83d\udd39 Most similar documents:\")\n\nfor tag, score in similar_docs:\n\n\tprint(f\"Document {tag} - Similarity Score: {score:.4f}\")<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"630\" height=\"178\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/27-image15.webp\" alt=\"Doc2Vec\" class=\"wp-image-231276\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/27-image15.webp 630w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/27-image15-300x85.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/27-image15-150x42.webp 150w\" sizes=\"auto, (max-width: 630px) 100vw, 630px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"328\" height=\"80\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/28-image35.webp\" alt=\"Doc2Vec\" class=\"wp-image-231278\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/28-image35.webp 328w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/28-image35-300x73.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/28-image35-150x37.webp 150w\" sizes=\"auto, (max-width: 328px) 100vw, 328px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-6\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Document-Level Representation:<\/strong> Effectively captures thematic and contextual information of larger texts.<\/li>\n<li><strong>Versatility:<\/strong> Useful in a variety of tasks, from recommendation systems to clustering and summarization.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-6\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Training Sensitivity:<\/strong> Requires significant data and careful tuning to produce high-quality docent vectors.<\/li>\n<li><strong>Static Embeddings:<\/strong> Each document is represented by one vector regardless of the internal variability of content.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-9-infersent\">9. InferSent<\/h2>\n<p>InferSent, developed by Facebook in 2017, was designed to generate high-quality sentence embeddings through supervised learning on natural language inference (NLI) datasets. It aims to capture semantic nuances at the sentence level, making it highly effective for tasks like semantic similarity and textual entailment.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-7\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Supervised Training:<\/strong> Uses labeled NLI data to learn sentence representations that reflect the logical relationships between sentences.<\/li>\n<li><strong>Bidirectional LSTMs:<\/strong> Employs recurrent neural networks that process sentences from both directions to capture context.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<br \/><\/strong>The model leverages supervised understanding to refine embeddings so that semantically similar sentences are closer together in the vector space, greatly enhancing performance on tasks like sentiment analysis and paraphrase detection.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-7\">Code Implementation<\/h3>\n<p>You can follow <a href=\"https:\/\/www.kaggle.com\/code\/jacksoncrow\/infersent-demo\">this<\/a> Kaggle Notebook to implement this.<\/p>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"421\" height=\"379\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/29-image25.webp\" alt=\"InferSent\" class=\"wp-image-231279\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/29-image25.webp 421w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/29-image25-300x270.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/29-image25-150x135.webp 150w\" sizes=\"auto, (max-width: 421px) 100vw, 421px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-7\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Rich Semantic Capturing:<\/strong> Provides deep, contextually nuanced sentence representations.<\/li>\n<li><strong>Task-Optimized:<\/strong> Excels at capturing relationships required for semantic inference tasks.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-7\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Dependence on Labeled Data:<\/strong> Requires extensively annotated datasets for training.<\/li>\n<li><strong>Computationally Intensive:<\/strong> More resource-demanding than unsupervised methods.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-10-universal-sentence-encoder-use\">10. Universal Sentence Encoder (USE)<\/h2>\n<p>The Universal Sentence Encoder (USE) is a model developed by Google to create high-quality, general-purpose sentence embeddings. Released in 2018, USE has been designed to work well across a variety of NLP tasks with minimal fine-tuning, making it a versatile tool for applications ranging from semantic search to text classification.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-8\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Architecture Options:<\/strong> USE can be implemented using Transformer architectures or Deep Averaging Networks (DANs) to encode sentences.<\/li>\n<li><strong>Pretraining:<\/strong> Trained on large, diverse datasets to capture broad language patterns, it maps sentences into a fixed-dimensional space.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<br \/><\/strong>USE provides robust embeddings across domains and tasks, making it an excellent \u201cout-of-the-box\u201d solution. Its design balances performance and efficiency, offering high-level embeddings without the need for extensive task-specific tuning.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-8\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>import tensorflow_hub as hub\n\nimport tensorflow as tf\n\nimport numpy as np\n\n# Load the model (this may take a few seconds on first run)\n\nembed = hub.load(\"https:\/\/tfhub.dev\/google\/universal-sentence-encoder\/4\")\n\nprint(\"\u2705 USE model loaded successfully!\")\n\n# Sample sentences\n\nsentences = [\n\n\t\"Machine learning is fun.\",\n\n\t\"Artificial intelligence and machine learning are related.\",\n\n\t\"I love playing football.\",\n\n\t\"Deep learning is a subset of machine learning.\"\n\n]\n\n# Get sentence embeddings\n\nembeddings = embed(sentences)\n\n# Convert to NumPy for easier manipulation\n\nembeddings_np = embeddings.numpy()\n\n# Display shape and first vector\n\nprint(f\"\ud83d\udd39 Embedding shape: {embeddings_np.shape}\")\n\nprint(f\"\ud83d\udd39 First sentence embedding (truncated):\\n{embeddings_np[0][:10]} ...\")\n\nfrom sklearn.metrics.pairwise import cosine_similarity\n\n# Compute pairwise cosine similarities\n\nsimilarity_matrix = cosine_similarity(embeddings_np)\n\n# Display similarity matrix\n\nimport pandas as pd\n\nsimilarity_df = pd.DataFrame(similarity_matrix, index=sentences, columns=sentences)\n\nprint(\"\ud83d\udd39 Sentence Similarity Matrix:\\n\")\n\nprint(similarity_df.round(2))\n\nimport matplotlib.pyplot as plt\n\nfrom sklearn.decomposition import PCA\n\n# Reduce to 2D\n\npca = PCA(n_components=2)\n\nreduced = pca.fit_transform(embeddings_np)\n\n# Plot\n\nplt.figure(figsize=(8, 6))\n\nplt.scatter(reduced[:, 0], reduced[:, 1], color=\"blue\")\n\nfor i, sentence in enumerate(sentences):\n\n\tplt.annotate(f\"Sentence {i+1}\", (reduced[i, 0]+0.01, reduced[i, 1]+0.01))\n\nplt.title(\"\ud83d\udcca Sentence Embeddings (PCA projection)\")\n\nplt.xlabel(\"PCA 1\")\n\nplt.ylabel(\"PCA 2\")\n\nplt.grid(True)\n\nplt.show()<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"663\" height=\"87\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/30-image26.webp\" alt=\"Universal Sentence Encoder (USE)\" class=\"wp-image-231282\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/30-image26.webp 663w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/30-image26-300x39.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/30-image26-150x20.webp 150w\" sizes=\"auto, (max-width: 663px) 100vw, 663px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1043\" height=\"484\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/31-image37.webp\" alt=\"Universal Sentence Encoder (USE)\" class=\"wp-image-231283\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/31-image37.webp 1043w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/31-image37-300x139.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/31-image37-768x356.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/31-image37-150x70.webp 150w\" sizes=\"auto, (max-width: 1043px) 100vw, 1043px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"756\" height=\"547\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/32-image31.webp\" alt=\"Universal Sentence Encoder (USE)\" class=\"wp-image-231284\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/32-image31.webp 756w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/32-image31-300x217.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/32-image31-150x109.webp 150w\" sizes=\"auto, (max-width: 756px) 100vw, 756px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-8\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Versatility:<\/strong> Well-suited for a broad range of applications without additional training.<\/li>\n<li><strong>Pretrained Convenience:<\/strong> Ready for immediate use, saving time and computational resources.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-8\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Fixed Representations:<\/strong> Produces a single embedding per sentence without dynamically adjusting to different contexts.<\/li>\n<li><strong>Model Size:<\/strong> Some variants are quite large, which can affect deployment in resource-limited environments.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-11-node2vec\">11. Node2Vec<\/h2>\n<p>Node2Vec is a method originally designed for learning node embeddings in graph structures. While not a text representation method per se, it is increasingly applied in NLP tasks that involve network or graph data, such as social networks or knowledge graphs. Introduced around 2016, it helps capture structural relationships in graph data.<\/p>\n<p><strong>Use Cases: <\/strong>Node classification, link prediction, graph clustering, recommendation systems.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-9\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Random Walks:<\/strong> Performs biased random walks on a graph to generate sequences of nodes.<\/li>\n<li><strong>Skip-gram Model:<\/strong> Applies a strategy similar to Word2Vec on these sequences to learn low-dimensional embeddings for nodes.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<br \/><\/strong>By simulating the sentences within the nodes, Node2Vec effectively captures the local and global structure of the graphs. It is highly adaptive and can be used for various downstream tasks, such as clustering, classification or recommendation systems in networked data.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-9\">Code Implementation<\/h3>\n<p>We will use this ready-made graph from NetworkX to view our Node2Vec implementation.To learn more about the Karate Club Graph, click <a href=\"https:\/\/networkx.org\/documentation\/stable\/reference\/generated\/networkx.generators.social.karate_club_graph.html\" target=\"_blank\" rel=\"nofollow noopener\">here<\/a>.<\/p>\n<pre class=\"wp-block-code\"><code>!pip install numpy==1.24.3 # Adjust version if needed\n\nimport networkx as nx\n\nimport numpy as np\n\nfrom node2vec import Node2Vec\n\nimport matplotlib.pyplot as plt\n\nfrom sklearn.decomposition import PCA\n\n# Create a simple graph\n\nG = nx.karate_club_graph()\u00a0 # A famous test graph with 34 nodes\n\n# Visualize original graph\n\nplt.figure(figsize=(6, 6))\n\nnx.draw(G, with_labels=True, node_color=\"skyblue\", edge_color=\"gray\", node_size=500)\n\nplt.title(\"Original Karate Club Graph\")\n\nplt.show()\n\n# Initialize Node2Vec model\n\nnode2vec = Node2Vec(G, dimensions=64, walk_length=30, num_walks=200, workers=2)\n\n# Train the model (Word2Vec under the hood)\n\nmodel = node2vec.fit(window=10, min_count=1, batch_words=4)\n\n# Get the vector for a specific node\n\nnode_id = 0\n\nvector = model.wv[str(node_id)]\u00a0 # Note: Node IDs are stored as strings\n\nprint(f\"\ud83d\udd39 Embedding for node {node_id}:\\n{vector[:10]}...\")\u00a0 # Truncated\n\n# Get all embeddings\n\nnode_ids = model.wv.index_to_key\n\nembeddings = np.array([model.wv[node] for node in node_ids])\n\n# Reduce dimensions to 2D\n\npca = PCA(n_components=2)\n\nreduced = pca.fit_transform(embeddings)\n\n# Plot embeddings\n\nplt.figure(figsize=(8, 6))\n\nplt.scatter(reduced[:, 0], reduced[:, 1], color=\"orange\")\n\nfor i, node in enumerate(node_ids):\n\n\tplt.annotate(node, (reduced[i, 0] + 0.05, reduced[i, 1] + 0.05))\n\nplt.title(\"\ud83d\udcca Node2Vec Embeddings (PCA Projection)\")\n\nplt.xlabel(\"PCA 1\")\n\nplt.ylabel(\"PCA 2\")\n\nplt.grid(True)\n\nplt.show()\n\n# Find most similar nodes to node 0\n\nsimilar_nodes = model.wv.most_similar(str(0), topn=5)\n\nprint(\"\ud83d\udd39 Nodes most similar to node 0:\")\n\nfor node, score in similar_nodes:\n\n\tprint(f\"Node {node} \u2192 Similarity Score: {score:.4f}\")<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"619\" height=\"642\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/33-image29.webp\" alt=\"Original Karate Club Graph\" class=\"wp-image-231287\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/33-image29.webp 619w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/33-image29-289x300.webp 289w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/33-image29-150x156.webp 150w\" sizes=\"auto, (max-width: 619px) 100vw, 619px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"678\" height=\"69\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/34-image4-1.webp\" alt=\"Ouput\" class=\"wp-image-231288\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/34-image4-1.webp 678w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/34-image4-1-300x31.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/34-image4-1-150x15.webp 150w\" sizes=\"auto, (max-width: 678px) 100vw, 678px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"706\" height=\"547\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/35-image32.webp\" alt=\"Node2Vec Embeddings\" class=\"wp-image-231289\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/35-image32.webp 706w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/35-image32-300x232.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/35-image32-150x116.webp 150w\" sizes=\"auto, (max-width: 706px) 100vw, 706px\"\/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"330\" height=\"119\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/36-image33.webp\" alt=\"Output\" class=\"wp-image-231290\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/36-image33.webp 330w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/36-image33-300x108.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/36-image33-150x54.webp 150w\" sizes=\"auto, (max-width: 330px) 100vw, 330px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-9\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Graph Structure Capture:<\/strong> Excels at embedding nodes with rich relational information.<\/li>\n<li><strong>Flexibility:<\/strong> Can be applied to any graph-structured data, not just language.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-9\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Domain Specificity:<\/strong> Less applicable to plain text unless represented as a graph.<\/li>\n<li><strong>Parameter Sensitivity:<\/strong> The quality of embeddings is sensitive to the parameters used in random walks.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-12-elmo-embeddings-from-language-models\">12. ELMo (Embeddings from Language Models)<\/h2>\n<p>ELMo, introduced by the Allen Institute for AI in 2018, marked a breakthrough by providing deep contextualized word representations. Unlike earlier models that generate a single vector per word, ELMo produces dynamic embeddings that change based on a sentence\u2019s context, capturing both syntactic and semantic nuances.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-10\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Bidirectional LSTMs:<\/strong> Processes text in both forward and backward directions to capture full contextual information.<\/li>\n<li><strong>Layered Representations:<\/strong> Combines representations from multiple layers of the neural network, each capturing different aspects of language.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<br \/><\/strong>The key innovation is that the same word can have different embeddings depending on its usage, allowing ELMo to handle ambiguity and polysemy more effectively. This context sensitivity leads to improvements in many downstream NLP tasks. It operates through customizable parameters, including <strong>dimensions<\/strong> (embedding vector size), <strong>walk_length<\/strong> (nodes per random walk), <strong>num_walks<\/strong> (walks per node), and bias parameters <strong>p<\/strong> (return factor) and <strong>q<\/strong> (in-out factor) that control walk behavior by balancing breadth-first (BFS) and depth-first (DFS) search tendencies. The methodology combines <strong>biased random walks<\/strong>, which explore node neighborhoods with tunable search strategies, with <strong>Word2Vec\u2019s Skip-gram architecture<\/strong> to learn embeddings preserving network structure and node relationships. Node2Vec enables effective node classification, link prediction, and graph clustering by capturing both local network patterns and broader structures in the embedding space.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-10\">Code Implementation<\/h3>\n<p>To implement and understand more about ELMo, you can refer to this article <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2019\/03\/learn-to-use-elmo-to-extract-features-from-text\/\" target=\"_blank\" rel=\"noopener\">here<\/a>.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-10\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Context-Awareness:<\/strong> Provides word embeddings that vary in accordance with the context.<\/li>\n<li><strong>Enhanced Performance:<\/strong> Improves results based on a variety of tasks, including sentiment analysis, question answering, and machine translation.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-10\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Computationally Demanding:<\/strong> Requires more resources for training and inference.<\/li>\n<li><strong>Complex Architecture:<\/strong> Challenging to implement and fine-tune compared to other simpler models.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-13-bert-and-its-variants\">13. BERT and Its Variants<\/h2>\n<h3 class=\"wp-block-heading\" id=\"h-what-is-bert\">What is BERT?<\/h3>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2019\/09\/demystifying-bert-groundbreaking-nlp-framework\/\" target=\"_blank\" rel=\"noreferrer noopener\">BERT<\/a> or Bidirectional Encoder Representations from Transformers, released by Google in 2018, revolutionized NLP by introducing a transformer-based architecture that captures bidirectional context. Unlike previous models that processed text in a unidirectional manner, BERT considers both the left and right context of each word. This deep, contextual understanding enables BERT to excel at tasks ranging from question answering and sentiment analysis to named entity recognition.<\/p>\n<p><strong>How It Works:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Transformer Architecture: <\/strong>BERT is built on a multi-layer transformer network that uses a self-attention mechanism to capture dependencies between all words in a sentence simultaneously. This allows the model to weigh the dependency of each word on every other word.<\/li>\n<li><strong>Masked Language Modeling: <\/strong>During pre-training, BERT randomly masks certain words in the input and then predicts them based on their context. This forces the model to learn bidirectional context and develop a robust understanding of language patterns.<\/li>\n<li><strong>Next Sentence Prediction: <\/strong>BERT is also trained on pairs of sentences, learning to predict whether one sentence logically follows another. This helps it capture relationships between sentences, an essential feature for tasks like document classification and natural language inference.<\/li>\n<\/ul>\n<p><strong>Additional Detail:<\/strong> BERT\u2019s architecture allows it to learn intricate patterns of language, including syntax and semantics. Fine-tuning on downstream tasks is straightforward, leading to state-of-the-art performance across many benchmarks.<\/p>\n<p><strong>Benefits:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Deep Contextual Understanding:<\/strong> By considering both past and future context, BERT generates richer, more nuanced word representations.<\/li>\n<li><strong>Versatility:<\/strong> BERT can be fine-tuned with relatively little additional training for a wide range of downstream tasks.<\/li>\n<\/ul>\n<p><strong>Shortcomings:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Heavy Computational Load:<\/strong> The model requires significant computational resources during both training and inference.<\/li>\n<li><strong>Large Model Size:<\/strong> BERT\u2019s large number of parameters can make it challenging to deploy in resource-constrained environments.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-sbert-sentence-bert\">SBERT (Sentence-BERT)<\/h3>\n<p>Sentence-BERT (SBERT) was introduced in 2019 to address a key limitation of BERT\u2014its inefficiency in generating semantically meaningful sentence embeddings for tasks like semantic similarity, clustering, and information retrieval. SBERT adapts BERT\u2019s architecture to produce fixed-size sentence embeddings that are optimized for comparing the meaning of sentences directly.<\/p>\n<p><strong>How It Works<\/strong>:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Siamese Network Architecture: <\/strong>SBERT modifies the original BERT structure by employing a siamese (or triplet) network architecture. This means it processes two (or more) sentences in parallel through identical BERT-based encoders, allowing the model to learn embeddings such that semantically similar sentences are close together in vector space.<\/li>\n<li><strong>Pooling Operation: <\/strong>After processing sentences through BERT, SBERT applies a pooling strategy (commonly meaning pooling) on the token embeddings to produce a fixed-size vector for each sentence.<\/li>\n<li><strong>Fine-Tuning with Sentence Pairs: <\/strong>SBERT is fine-tuned on tasks involving sentence pairs using contrastive or triplet loss. This training objective encourages the model to place similar sentences closer together and dissimilar ones further apart in the embedding space.<\/li>\n<\/ul>\n<p><strong>Benefits<\/strong>:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Efficient Sentence Comparisons:<\/strong> SBERT is optimized for tasks like semantic search and clustering. Due to its fixed size and semantically rich sentence embeddings, comparing tens of thousands of sentences becomes computationally feasible.<\/li>\n<li><strong>Versatility in Downstream Tasks:<\/strong> SBERT embeddings are effective for a variety of applications, such as paraphrase detection, semantic textual similarity, and information retrieval.<\/li>\n<\/ul>\n<p><strong>Shortcomings:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Dependence on Fine-Tuning Data:<\/strong> The quality of SBERT embeddings can be heavily influenced by the domain and quality of the training data used during fine-tuning.<\/li>\n<li><strong>Resource Intensive Training:<\/strong> Although inference is efficient, the initial fine-tuning process requires considerable computational resources.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-distilbert\">DistilBERT<\/h3>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2022\/11\/introduction-to-distilbert-in-student-model\/\" target=\"_blank\" rel=\"noreferrer noopener\">DistilBERT<\/a>, introduced by Hugging Face in 2019, is a lighter and faster variant of BERT that retains much of its performance. It was created using a technique called knowledge distillation, where a smaller model (student) is trained to mimic the behavior of a larger, pre-trained model (teacher), in this case, BERT.<\/p>\n<p><strong>How It Works:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Knowledge Distillation:<\/strong> DistilBERT is trained to match the output distributions of the original BERT model while using fewer parameters. It removes some layers (e.g., 6 instead of 12 in the BERT-base) but maintains crucial learning behavior.<\/li>\n<li><strong>Loss Function:<\/strong> The training uses a combination of language modeling loss and distillation loss (KL divergence between teacher and student logits).<\/li>\n<li><strong>Speed Optimization:<\/strong> DistilBERT is optimized to be 60% faster during inference while retaining ~97% of BERT\u2019s performance on downstream tasks.<\/li>\n<\/ul>\n<p><strong>Benefits<\/strong>:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Lightweight and Fast:<\/strong> Ideal for real-time or mobile applications due to reduced computational demands.<\/li>\n<li><strong>Competitive Performance:<\/strong> Achieves near-BERT accuracy with significantly lower resource usage.<\/li>\n<\/ul>\n<p><strong>Shortcomings<\/strong>:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Slight Drop in Accuracy:<\/strong> While very close, it might slightly underperform compared to the full BERT model in complex tasks.<\/li>\n<li><strong>Limited Fine-Tuning Flexibility:<\/strong> It may not generalize as well in niche domains as full-sized models.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-roberta\">RoBERTa<\/h3>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/04\/training-an-adapter-for-roberta-model-for-sequence-classification-task\/\" target=\"_blank\" rel=\"noreferrer noopener\">RoBERTa<\/a> or Robustly Optimized BERT Pretraining Approach was introduced by Facebook AI in 2019 as a robust enhancement over BERT. It tweaks the pretraining methodology to improve performance significantly across a wide range of tasks.<\/p>\n<p><strong>How It Works:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Training<\/strong> <strong>Enhancements<\/strong>:\n<ul class=\"wp-block-list\">\n<li>Removes the Next Sentence Prediction (NSP) objective, which was found to hurt performance in some settings.<\/li>\n<li>Trains on much <strong>larger datasets<\/strong> (e.g., Common Crawl) and for <strong>longer durations<\/strong>.<\/li>\n<li>Uses <strong>larger mini-batches<\/strong> and <strong>more training steps<\/strong> to stabilize and optimize learning.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Dynamic Masking:<\/strong> This method applies masking on the fly during each training epoch, exposing the model to more diverse masking patterns than BERT\u2019s static masking.<\/li>\n<\/ul>\n<p><strong>Benefits:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Superior Performance:<\/strong> Outperforms BERT on several benchmarks, including GLUE and SQuAD.<\/li>\n<li><strong>Robust Learning:<\/strong> Better generalization across domains due to improved training data and strategies.<\/li>\n<\/ul>\n<p><strong>Shortcomings<\/strong>:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Resource Intensive:<\/strong> Even more computationally demanding than BERT.<\/li>\n<li><strong>Overfitting Risk:<\/strong> With extensive training and large datasets, there\u2019s a risk of overfitting if not handled carefully.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-11\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>from transformers import AutoTokenizer, AutoModel\n\nimport torch\n\n# Input sentence for embedding\n\nsentence = \"Natural Language Processing is transforming how machines understand humans.\"\n\n# Choose device (GPU if available)\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# =============================\n\n# 1. BERT Base Uncased\n\n# =============================\n\n# model_name = \"bert-base-uncased\"\n\n# =============================\n\n# 2. SBERT - Sentence-BERT\n\n# =============================\n\n# model_name = \"sentence-transformers\/all-MiniLM-L6-v2\"\n\n# =============================\n\n# 3. DistilBERT\n\n# =============================\n\n# model_name = \"distilbert-base-uncased\"\n\n# =============================\n\n# 4. RoBERTa\n\n# =============================\n\nmodel_name = \"roberta-base\"\u00a0 # Only RoBERTa is active now uncomment other to test other models\n\n# Load tokenizer and model\n\ntokenizer = AutoTokenizer.from_pretrained(model_name)\n\nmodel = AutoModel.from_pretrained(model_name).to(device)\n\nmodel.eval()\n\n# Tokenize input\n\ninputs = tokenizer(sentence, return_tensors=\"pt\", truncation=True, padding=True).to(device)\n\n# Forward pass to get embeddings\n\nwith torch.no_grad():\n\n\u00a0\u00a0\u00a0\u00a0outputs = model(**inputs)\n\n# Get token embeddings\n\ntoken_embeddings = outputs.last_hidden_state\u00a0 # (batch_size, seq_len, hidden_size)\n\n# Mean Pooling for sentence embedding\n\nsentence_embedding = torch.mean(token_embeddings, dim=1)\n\nprint(f\"Sentence embedding from {model_name}:\")\n\nprint(sentence_embedding)<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1657\" height=\"611\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/37-image12.webp\" alt=\"Output\" class=\"wp-image-231291\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/37-image12.webp 1657w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/37-image12-300x111.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/37-image12-768x283.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/37-image12-1536x566.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/37-image12-150x55.webp 150w\" sizes=\"auto, (max-width: 1657px) 100vw, 1657px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-summary\">Summary<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>BERT<\/strong> provides deep, bidirectional contextualized embeddings ideal for a wide range of NLP tasks. It captures intricate language patterns through transformer-based self-attention but produces token-level embeddings that need to be aggregated for sentence-level tasks.<\/li>\n<li><strong>SBERT<\/strong> extends BERT by transforming it into a model that directly produces meaningful sentence embeddings. With its siamese network architecture and contrastive learning objectives, SBERT excels at tasks requiring fast and accurate semantic comparisons between sentences, such as semantic search, paraphrase detection, and sentence clustering.<\/li>\n<li><strong>DistilBERT<\/strong> offers a lighter, faster alternative to BERT by using knowledge distillation. It retains most of BERT\u2019s performance while being more suitable for real-time or resource-constrained applications. It is ideal when inference speed and efficiency are key concerns, though it may slightly underperform in complex scenarios.<\/li>\n<li><strong>RoBERTa<\/strong> improves upon BERT by modifying its pre-training regime, removing the next sentence prediction task by using larger datasets, and applying dynamic masking. These changes result in better generalization and performance across benchmarks, though at the cost of increased computational resources.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-other-notable-bert-variants\">Other Notable BERT Variants<\/h3>\n<p>While BERT and its direct descendants like SBERT, DistilBERT, and RoBERTa have made a significant impact in NLP, several other powerful variants have emerged to address different limitations and enhance specific capabilities:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>ALBERT (A Lite BERT)<\/strong><strong><br \/><\/strong>ALBERT is a more efficient version of BERT that reduces the number of parameters through two key innovations: <em>factorized embedding parameterization<\/em> (which separates the size of the vocabulary embedding from the hidden layers) and <em>cross-layer parameter sharing<\/em> (which reuses weights across transformer layers). These changes make ALBERT faster and more memory-efficient while preserving performance on many NLP benchmarks.<\/li>\n<li><strong>XLNet<br \/><\/strong>Unlike BERT, which relies on masked language modeling, <strong>XLNet<\/strong> adopts a <em>permutation-based autoregressive<\/em> training strategy. This allows it to capture bidirectional context without relying on data corruption like masking. XLNet also integrates ideas from Transformer-XL, which enables it to model longer-term dependencies and outperform BERT on several NLP tasks.<\/li>\n<li><strong>T5 (Text-to-Text Transfer Transformer)<\/strong><br \/>Developed by Google Research, <strong>T5<\/strong> frames every NLP task, from translation to classification, as a text-to-text problem. For example, instead of producing a classification label directly, T5 learns to <em>generate<\/em> the label as a word or phrase. This unified approach makes it highly flexible and powerful, capable of tackling a broad spectrum of NLP challenges.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-14-clip-and-blip\">14. CLIP and BLIP<\/h2>\n<p>Modern multimodal models like <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2021\/01\/openais-future-of-vision-contrastive-language-image-pre-trainingclip\/\" target=\"_blank\" rel=\"noreferrer noopener\">CLIP<\/a> (Contrastive Language-Image Pretraining) and <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/03\/salesforce-blip-revolutionizing-image-captioning\/\" target=\"_blank\" rel=\"noreferrer noopener\">BLIP<\/a> (Bootstrapping Language-Image Pre-training) represent the latest frontier in embedding techniques. They bridge the gap between textual and visual data, enabling tasks that involve both language and images. These models have become essential for applications such as image search, captioning, and visual question answering.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-how-it-works-11\">How It Works<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>CLIP:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong> Trains on large datasets of image-text pairs, using contrastive learning to align image embeddings with corresponding text embeddings.<\/li>\n<li><strong>Process:<\/strong> The model learns to map images and text into a shared vector space where related pairs are closer together.<\/li>\n<\/ul>\n<\/li>\n<li><strong>BLIP:<\/strong>\n<ul class=\"wp-block-list\">\n<li><strong>Mechanism:<\/strong> Uses a bootstrapping approach to refine the alignment between language and vision through iterative training.<\/li>\n<li><strong>Process:<\/strong> Improves upon initial alignments to achieve more accurate multimodal representations.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Additional Detail:<\/strong><strong><br \/><\/strong>These models harness the power of transformers for text and convolutional or transformer-based networks for images. Their ability to jointly reason about text and visual content has opened up new possibilities in multimodal AI research.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-code-implementation-12\">Code Implementation<\/h3>\n<pre class=\"wp-block-code\"><code>from transformers import CLIPProcessor, CLIPModel\n\n# from transformers import BlipProcessor, BlipModel\u00a0 # Uncomment to use BLIP\n\nfrom PIL import Image\n\nimport torch\n\nimport requests\n\n# Choose device\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# Load a sample image and text\n\nimage_url = \"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/datasets\/cat_style_layout.png\"\n\nimage = Image.open(requests.get(image_url, stream=True).raw).convert(\"RGB\")\n\ntext = \"a cute puppy\"\n\n# ===========================\n\n# 1. CLIP (for Embeddings)\n\n# ===========================\n\nclip_model_name = \"openai\/clip-vit-base-patch32\"\n\nclip_model = CLIPModel.from_pretrained(clip_model_name).to(device)\n\nclip_processor = CLIPProcessor.from_pretrained(clip_model_name)\n\n# Preprocess input\n\ninputs = clip_processor(text=[text], images=image, return_tensors=\"pt\", padding=True).to(device)\n\n# Get text and image embeddings\n\nwith torch.no_grad():\n\n\u00a0\u00a0\u00a0\u00a0text_embeddings = clip_model.get_text_features(input_ids=inputs[\"input_ids\"])\n\n\u00a0\u00a0\u00a0\u00a0image_embeddings = clip_model.get_image_features(pixel_values=inputs[\"pixel_values\"])\n\n# Normalize embeddings (optional)\n\ntext_embeddings = text_embeddings \/ text_embeddings.norm(dim=-1, keepdim=True)\n\nimage_embeddings = image_embeddings \/ image_embeddings.norm(dim=-1, keepdim=True)\n\nprint(\"Text Embedding Shape (CLIP):\", text_embeddings.shape)\n\nprint(\"Image Embedding Shape (CLIP):\", image_embeddings)\n\n# ===========================\n\n# 2. BLIP (commented)\n\n# ===========================\n\n# blip_model_name = \"Salesforce\/blip-image-text-matching-base\"\n\n# blip_processor = BlipProcessor.from_pretrained(blip_model_name)\n\n# blip_model = BlipModel.from_pretrained(blip_model_name).to(device)\n\n# inputs = blip_processor(images=image, text=text, return_tensors=\"pt\").to(device)\n\n# with torch.no_grad():\n\n# \u00a0 \u00a0 text_embeddings = blip_model.text_encoder(input_ids=inputs[\"input_ids\"]).last_hidden_state[:, 0, :]\n\n# \u00a0 \u00a0 image_embeddings = blip_model.vision_model(pixel_values=inputs[\"pixel_values\"]).last_hidden_state[:, 0, :]\n\n# print(\"Text Embedding Shape (BLIP):\", text_embeddings.shape)\n\n# print(\"Image Embedding Shape (BLIP):\", image_embeddings)<\/code><\/pre>\n<p><strong>Output:<\/strong><\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"923\" height=\"417\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/38-image7.webp\" alt=\"Output\" class=\"wp-image-231300\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/38-image7.webp 923w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/38-image7-300x136.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/38-image7-768x347.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/38-image7-150x68.webp 150w\" sizes=\"auto, (max-width: 923px) 100vw, 923px\"\/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\" id=\"h-benefits-11\">Benefits<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>Cross-Modal Understanding:<\/strong> Provides powerful representations that work across text and images.<\/li>\n<li><strong>Wide Applicability:<\/strong> Useful in image retrieval, captioning, and other multimodal tasks.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-shortcomings-11\">Shortcomings<\/h3>\n<ul class=\"wp-block-list\">\n<li><strong>High Complexity:<\/strong> Training requires large, well-curated datasets of paired data.<\/li>\n<li><strong>Heavy Resource Requirements:<\/strong> Multimodal models are among the most computationally demanding.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-comparison-of-embeddings\">Comparison of Embeddings<\/h2>\n<figure class=\"wp-block-table\">\n<table class=\"table table-bordered border-black table-striped\">\n<thead>\n<tr>\n<th><strong>Embedding<\/strong><\/th>\n<th><strong>Type<\/strong><\/th>\n<th><strong>Model Architecture \/ Approach<\/strong><\/th>\n<th><strong>Common Use Cases<\/strong><\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Count Vectorizer<\/strong><\/td>\n<td>Context-independent, No ML<\/td>\n<td>Count-based (Bag of Words)<\/td>\n<td>Sentence embeddings for search, chatbots, and semantic similarity<\/td>\n<\/tr>\n<tr>\n<td><strong>One-Hot Encoding<\/strong><\/td>\n<td>Context-independent, No ML<\/td>\n<td>Manual encoding<\/td>\n<td>Baseline models, rule-based systems<\/td>\n<\/tr>\n<tr>\n<td><strong>TF-IDF<\/strong><\/td>\n<td>Context-independent, No ML<\/td>\n<td>Count + Inverse Document Frequency<\/td>\n<td>Document ranking, text similarity, keyword extraction<\/td>\n<\/tr>\n<tr>\n<td><strong>Okapi BM25<\/strong><\/td>\n<td>Context-independent, Statistical Ranking<\/td>\n<td>Probabilistic IR model<\/td>\n<td>Search engines, information retrieval<\/td>\n<\/tr>\n<tr>\n<td><strong>Word2Vec (CBOW, SG)<\/strong><\/td>\n<td>Context-independent, ML-based<\/td>\n<td>Neural network (shallow)<\/td>\n<td>Sentiment analysis, word similarity, NLP pipelines<\/td>\n<\/tr>\n<tr>\n<td><strong>GloVe<\/strong><\/td>\n<td>Context-independent, ML-based<\/td>\n<td>Global co-occurrence matrix + ML<\/td>\n<td>Word similarity, embedding initialization<\/td>\n<\/tr>\n<tr>\n<td><strong>FastText<\/strong><\/td>\n<td>Context-independent, ML-based<\/td>\n<td>Word2Vec + Subword embeddings<\/td>\n<td>Morphologically rich languages, OOV word handling<\/td>\n<\/tr>\n<tr>\n<td><strong>Doc2Vec<\/strong><\/td>\n<td>Context-independent, ML-based<\/td>\n<td>Extension of Word2Vec for documents<\/td>\n<td>Document classification, clustering<\/td>\n<\/tr>\n<tr>\n<td><strong>InferSent<\/strong><\/td>\n<td>Context-dependent, RNN-based<\/td>\n<td>BiLSTM with supervised learning<\/td>\n<td>Semantic similarity, NLI tasks<\/td>\n<\/tr>\n<tr>\n<td><strong>Universal Sentence Encoder<\/strong><\/td>\n<td>Context-dependent, Transformer-based<\/td>\n<td>Transformer \/ DAN (Deep Averaging Net)<\/td>\n<td>Sentence embeddings for search, chatbots, semantic similarity<\/td>\n<\/tr>\n<tr>\n<td><strong>Node2Vec<\/strong><\/td>\n<td>Graph-based embedding<\/td>\n<td>Random walk + Skipgram<\/td>\n<td>Graph representation, recommendation systems, link prediction<\/td>\n<\/tr>\n<tr>\n<td><strong>ELMo<\/strong><\/td>\n<td>Context-dependent, RNN-based<\/td>\n<td>Bi-directional LSTM<\/td>\n<td>Named Entity Recognition, Question Answering, Coreference Resolution<\/td>\n<\/tr>\n<tr>\n<td><strong>BERT &amp; Variants<\/strong><\/td>\n<td>Context-dependent, Transformer-based<\/td>\n<td>Q&amp;A, sentiment analysis, summarization, and semantic search<\/td>\n<td>Q&amp;A, sentiment analysis, summarization, semantic search<\/td>\n<\/tr>\n<tr>\n<td><strong>CLIP<\/strong><\/td>\n<td>Multimodal, Transformer-based<\/td>\n<td>Vision + Text encoders (Contrastive)<\/td>\n<td>Image captioning, cross-modal search, text-to-image retrieval<\/td>\n<\/tr>\n<tr>\n<td><strong>BLIP<\/strong><\/td>\n<td>Multimodal, Transformer-based<\/td>\n<td>Vision-Language Pretraining (VLP)<\/td>\n<td>Image captioning, VQA (Visual Question Answering)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>The journey of embeddings has come a long way from basic count-based methods like one-hot encoding to today\u2019s powerful, context-aware, and even multimodal models like BERT and CLIP. Each step has been about pushing past the limitations of the last, helping us better understand and represent human language. Nowadays, thanks to platforms like Hugging Face and Ollama, we have access to a growing library of cutting-edge embedding models making it easier than ever to tap into this new era of language intelligence.<\/p>\n<p>But beyond knowing how these techniques work, it\u2019s worth considering how they fit our real-world goals. Whether you\u2019re building a chatbot, a semantic search engine, a recommender system, or a document summarization system, there\u2019s an embedding out there that brings our ideas to life. After all, in today\u2019s world of language tech, there\u2019s truly a vector for every vision.<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/shaik8558834\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_An81zCg.webp\" width=\"48\" height=\"48\" alt=\"Shaik Hamzah\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>GenAI Intern @ Analytics Vidhya | Final Year @ VIT Chennai<br \/>Passionate about AI and machine learning, I&#8217;m eager to dive into roles as an AI\/ML Engineer or Data Scientist where I can make a real impact. With a knack for quick learning and a love for teamwork, I&#8217;m excited to bring innovative solutions and cutting-edge advancements to the table. My curiosity drives me to explore AI across various fields and take the initiative to delve into data engineering, ensuring I stay ahead and deliver impactful projects.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Summary: Evolution of Embeddings from basic count-based methods (TF-IDF, Word2Vec) to context-aware models like BERT and ELMo, which capture nuanced semantics by analyzing entire sentences bidirectionally. Leaderboards such as MTEB benchmark embeddings for tasks like retrieval and classification. Open-source platforms (Hugging Face) allow developers to access cutting-edge embeddings and deploy models tailored to different use [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":196578,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[20338,59412,7471,7289,7755],"dealstore":[],"offerexpiration":[],"class_list":["post-196577","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-defining","tag-embedding","tag-evolution","tag-powerful","tag-techniques"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>14 Powerful Techniques Defining the Evolution of Embedding - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=196577\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"14 Powerful Techniques Defining the Evolution of Embedding - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Summary: Evolution of Embeddings from basic count-based methods (TF-IDF, Word2Vec) to context-aware models like BERT and ELMo, which capture nuanced semantics by analyzing entire sentences bidirectionally. Leaderboards such as MTEB benchmark embeddings for tasks like retrieval and classification. Open-source platforms (Hugging Face) allow developers to access cutting-edge embeddings and deploy models tailored to different use [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=196577\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-04-21T12:14:50+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"36 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=196577#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=196577\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"14 Powerful Techniques Defining the Evolution of Embedding\",\"datePublished\":\"2025-04-21T12:14:50+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=196577\"},\"wordCount\":4977,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=196577#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp\",\"keywords\":[\"Defining\",\"Embedding\",\"evolution\",\"Powerful\",\"Techniques\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=196577#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=196577\",\"url\":\"https:\/\/fivemor.com\/?p=196577\",\"name\":\"14 Powerful Techniques Defining the Evolution of Embedding - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=196577#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=196577#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp\",\"datePublished\":\"2025-04-21T12:14:50+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=196577#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=196577\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=196577#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=196577#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"14 Powerful Techniques Defining the Evolution of Embedding\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"14 Powerful Techniques Defining the Evolution of Embedding - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=196577","og_locale":"en_US","og_type":"article","og_title":"14 Powerful Techniques Defining the Evolution of Embedding - Som2ny Network","og_description":"Summary: Evolution of Embeddings from basic count-based methods (TF-IDF, Word2Vec) to context-aware models like BERT and ELMo, which capture nuanced semantics by analyzing entire sentences bidirectionally. Leaderboards such as MTEB benchmark embeddings for tasks like retrieval and classification. Open-source platforms (Hugging Face) allow developers to access cutting-edge embeddings and deploy models tailored to different use [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=196577","og_site_name":"Som2ny Network","article_published_time":"2025-04-21T12:14:50+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"36 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=196577#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=196577"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"14 Powerful Techniques Defining the Evolution of Embedding","datePublished":"2025-04-21T12:14:50+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=196577"},"wordCount":4977,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=196577#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp","keywords":["Defining","Embedding","evolution","Powerful","Techniques"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=196577#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=196577","url":"https:\/\/fivemor.com\/?p=196577","name":"14 Powerful Techniques Defining the Evolution of Embedding - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=196577#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=196577#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp","datePublished":"2025-04-21T12:14:50+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=196577#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=196577"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=196577#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/The-Evolution-of-Embeddings-A-Comprehensive-Journey.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=196577#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"14 Powerful Techniques Defining the Evolution of Embedding"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/196577","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=196577"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/196577\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/196578"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=196577"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=196577"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=196577"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=196577"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=196577"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}