{"id":94664,"date":"2025-02-18T00:54:17","date_gmt":"2025-02-18T00:54:17","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/from-policy-gradient-to-grpo\/"},"modified":"2025-02-18T00:54:17","modified_gmt":"2025-02-18T00:54:17","slug":"from-policy-gradient-to-grpo","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=94664","title":{"rendered":"From Policy Gradient to GRPO"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>For decades,\u00a0Reinforcement Learning (RL)\u00a0has been the driving force behind breakthroughs in\u00a0robotics, game-playing AI (AlphaGo, OpenAI Five), and control systems. RL\u2019s strength lies in its ability to optimize decision-making by\u00a0maximizing long-term rewards, making it ideal for problems requiring sequential reasoning. However,\u00a0large language models (LLMs)\u00a0initially relied on\u00a0supervised learning, where models were fine-tuned on static datasets. This approach lacked adaptability\u2014while LLMs could mimic human text, they struggled with nuanced human preference alignment, leading to inconsistencies in conversational AI. The introduction of\u00a0RLHF (Reinforcement Learning with Human Feedback)\u00a0changed everything. By integrating RL into LLM fine-tuning, models like\u00a0ChatGPT, DeepSeek, Gemini, and Claude\u00a0could optimize (LLM Optimization) their responses based on user feedback. <\/p>\n<p>However, standard\u00a0PPO-based RLHF\u00a0had inefficiencies, requiring expensive reward modeling and iterative training. Enter\u00a0DeepSeek\u2019s Group Relative Policy Optimization (GRPO)\u2014a breakthrough that\u00a0eliminated the need for explicit reward modeling\u00a0by directly optimizing\u00a0preference rankings. To fully grasp the significance of GRPO, we must first explore the\u00a0fundamental policy optimization techniques (LLM optimization)\u00a0that power modern reinforcement learning.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-learning-objectives-nbsp\">Learning Objectives\u00a0<\/h3>\n<ul class=\"wp-block-list\">\n<li>Understand why RL-based techniques are crucial for optimizing LLMs like ChatGPT, DeepSeek, Claude, and Gemini.<\/li>\n<li>Learn the fundamentals of policy optimization, including PG, TRPO, and PPO.Explore DPO and GRPO for preference-based LLM training without explicit reward models.<\/li>\n<li>Compare PG, TRPO, PPO, DPO, and GRPO to determine the best approach for RL and LLM fine-tuning.<\/li>\n<li>Gain hands-on experience with Python implementations of policy optimization algorithms.<\/li>\n<li>Evaluate fine-tuning impact using training loss curves and probability distributions.<\/li>\n<li>Apply DPO and GRPO to enhance LLM safety, alignment, and reliability.<\/li>\n<\/ul>\n<p><em><strong>This article was published as a part of the\u00a0<\/strong><\/em><a href=\"https:\/\/www.analyticsvidhya.com\/datahack\/blogathon\" target=\"_blank\" rel=\"noreferrer noopener\"><em><strong>Data Science Blogathon.<\/strong><\/em><\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-primer-on-policy-optimization-techniques\">Primer on Policy Optimization Techniques<\/h2>\n<p>Before diving into DeepSeek\u2019s GRPO, it\u2019s crucial to understand the policy optimization techniques that form the foundation of reinforcement learning (RL) in both traditional control tasks and <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/12\/fine-tuning-llama-3-2-3b-for-rag\/\" target=\"_blank\" rel=\"noreferrer noopener\">LLM fine-tuning<\/a>. Policy optimization refers to the process of improving an AI agent\u2019s decision-making strategy (policy) to maximize expected rewards. While early methods like vanilla <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2020\/11\/baseline-for-policy-gradients\/\" target=\"_blank\" rel=\"noreferrer noopener\">policy gradient (PG)<\/a> laid the groundwork, more sophisticated techniques such as TRPO, PPO, DPO, and GRPO evolved to address issues like stability, efficiency, and preference alignment.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-what-is-policy-optimization\">What is Policy Optimization?<\/h3>\n<p>\u00a0 At its core, policy optimization is about learning the optimal policy <b>\u03c0_\u03b8(a\u2223s)<\/b>, which maps a state <i>s<\/i>\u00a0to an action <i>a<\/i>\u00a0while maximizing long-term rewards. The objective function in RL is typically formulated as:\u00a0\u00a0<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"227\" height=\"60\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/Eq1.webp\" alt=\"Formula\" class=\"wp-image-221863\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/Eq1.webp 227w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/Eq1-150x40.webp 150w\" sizes=\"auto, (max-width: 227px) 100vw, 227px\"\/><\/figure>\n<p>Where <b>R(\u03c4)<\/b>\u00a0is the total reward collected in a trajectory \u03c4, and the expectation is taken over all possible trajectories following policy <b>\u03c0_\u03b8<\/b>.<\/p>\n<p>There are three major approaches to policy optimization:<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-1-nbsp-gradient-based-optimization-policy-gradient-methods\">1.\u00a0Gradient-Based Optimization (Policy Gradient Methods)<\/h4>\n<ul class=\"wp-block-list\">\n<li>These methods directly compute gradients of expected reward and update policy parameters using gradient ascent.<\/li>\n<li>Example: REINFORCE algorithm (Vanilla Policy Gradient).<\/li>\n<li>Pros: Simple, works with continuous and discrete actions.<\/li>\n<li>Cons: High variance, requires tricks like baseline subtraction.<\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-2-nbsp-trust-region-optimization-trpo-ppo\">2.\u00a0Trust-Region Optimization (TRPO, PPO)<\/h4>\n<ul class=\"wp-block-list\">\n<li>Introduces constraints (KL divergence) to ensure policy updates are stable and not too drastic.<\/li>\n<li>Example: TRPO ensures updates stay within a \u201ctrust region\u201d; PPO simplifies this with clipping.<\/li>\n<li>Pros: More stable than raw policy gradients.<\/li>\n<li>Cons: Computationally expensive (TRPO), hyperparameter-sensitive (PPO).<\/li>\n<\/ul>\n<h4 class=\"wp-block-heading\" id=\"h-3-nbsp-preference-based-optimization-dpo-grpo\">3.\u00a0Preference-Based Optimization (DPO, GRPO)<\/h4>\n<ul class=\"wp-block-list\">\n<li>Optimizes directly from ranked human preferences instead of rewards.<\/li>\n<li>Example: DPO learns from preferred vs. rejected responses; GRPO generalizes to groups.<\/li>\n<li>Pros: Eliminates the need for reward models and better aligns LLMs with human intent.<\/li>\n<li>Cons: Requires high-quality preference data.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-mathematical-foundations-required-for-all-methods\">Mathematical Foundations (Required for All Methods)<\/h2>\n<h3 class=\"wp-block-heading\" id=\"h-a-markov-decision-process-mdp\">A. Markov Decision Process (MDP)<\/h3>\n<p>RL is typically formulated as a <b>Markov Decision Process<\/b> (MDP), represented as:<\/p>\n<figure class=\"wp-block-image size-full is-resized figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"167\" height=\"48\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/MDP.webp\" alt=\"Formula\" class=\"wp-image-221864\" style=\"width:167px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/MDP.webp 167w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/MDP-150x43.webp 150w\" sizes=\"auto, (max-width: 167px) 100vw, 167px\"\/><\/figure>\n<p>where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>S<\/b> is the state space,<\/li>\n<li><b>A<\/b> is the action space,<\/li>\n<li><b>P(s\u2032\u2223s,a)<\/b> is the transition probability to state s\u2032,<\/li>\n<li><b>R(s,a)<\/b> is the reward function,<\/li>\n<li><b>\u03b3<\/b> is the discount factor (how much future rewards are valued).<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-nbsp-b-expected-return-j-\u03b8\">\u00a0 B. Expected Return J(\u03b8)<\/h3>\n<p>\u00a0 The <b>Expected Return<\/b> (ER) measures how much cumulative reward we expect from following policy <b>\u03c0_\u03b8<\/b>:\u00a0\u00a0<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"234\" height=\"99\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/ER.webp\" alt=\"Formula\" class=\"wp-image-221865\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/ER.webp 234w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/ER-150x63.webp 150w\" sizes=\"auto, (max-width: 234px) 100vw, 234px\"\/><\/figure>\n<p>\u00a0 where <b>\u03b3<\/b> (0 \u2264<b> \u03b3<\/b> \u2264 1) determines how much future rewards contribute.\u00a0\u00a0<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-c-policy-gradient-theorem\">C. Policy Gradient Theorem<\/h3>\n<p>Policy gradient (PG) methods update the policy using gradients of expected rewards. The key equation:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"383\" height=\"47\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG.webp\" alt=\"Formula\" class=\"wp-image-221859\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG.webp 383w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG-300x37.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG-150x18.webp 150w\" sizes=\"auto, (max-width: 383px) 100vw, 383px\"\/><\/figure>\n<p>where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>A(s,a)<\/b> is the advantage function (how good action <b>a<\/b> is compared to average actions in state <b>s<\/b>).<\/li>\n<li><b> log\u03c0_\u03b8\u200b<\/b> ensures we increase the probabilities of better actions.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-d-advantage-function-a-s-a\">D. Advantage Function A(s,a)<\/h3>\n<p>To reduce variance in gradient estimates, we use the advantage function:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"275\" height=\"65\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/AdvFunc.webp\" alt=\"Formula\" class=\"wp-image-221858\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/AdvFunc.webp 275w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/AdvFunc-150x35.webp 150w\" sizes=\"auto, (max-width: 275px) 100vw, 275px\"\/><\/figure>\n<p>where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>Q(s,a)<\/b> is the expected return for taking action <b>a<\/b> at state <b>s<\/b>.<\/li>\n<li><b>V(s) <\/b>is the expected return following policy <b>\u03c0<\/b> from <b>s<\/b>.<\/li>\n<\/ul>\n<p>Using <b>A(s,a)<\/b>\u00a0helps make updates more stable and efficient.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-policy-gradient-pg-the-foundation\">Policy Gradient (PG) \u2013 The Foundation<\/h2>\n<p>The Policy Gradient (PG) method is the most fundamental approach to reinforcement learning. Instead of learning a value function, PG directly parameterizes the policy <b>\u03c0_\u03b8(a\u2223s)\u00a0<\/b>and updates it using gradient ascent. This allows learning in continuous action spaces, making it effective for tasks like robotics, game AI, and LLM fine-tuning.<\/p>\n<p>However, PG methods suffer from high variance due to their reliance on sampling full trajectories. More advanced methods like TRPO, PPO, and GRPO build upon PG to improve stability.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-the-policy-gradient-theorem\">The Policy Gradient Theorem<\/h2>\n<p>\u00a0 The goal of policy optimization is to find policy parameters <b>\u03b8<\/b>\u00a0that maximize expected return:\u00a0\u00a0<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"222\" height=\"53\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_1.webp\" alt=\"The Policy Gradient Theorem\" class=\"wp-image-221857\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_1.webp 222w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_1-150x36.webp 150w\" sizes=\"auto, (max-width: 222px) 100vw, 222px\"\/><\/figure>\n<p>Using the log-derivative trick, we obtain the Policy Gradient Theorem:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"377\" height=\"59\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_2.webp\" alt=\"The Policy Gradient Theorem\" class=\"wp-image-221856\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_2.webp 377w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_2-300x47.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_2-150x23.webp 150w\" sizes=\"auto, (max-width: 377px) 100vw, 377px\"\/><\/figure>\n<p>where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>\u2207\u03b8\u200blog\u03c0\u03b8\u200b(a\u2223s) <\/b>is the gradient of the log-probability of taking action aaa.<\/li>\n<li><b>A(s,a)<\/b> (Advantage function) determines how much better action aaa is compared to others.<\/li>\n<li>We perform gradient ascent to increase the probability of good actions.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-code-example-reinforce-algorithm\">Code Example: REINFORCE Algorithm<\/h2>\n<p>The REINFORCE algorithm is the simplest form of PG. It samples trajectories, computes rewards, and updates the policy parameters. Below is the main training loop (only the key function is shown to limit the scope; the full notebook is <a href=\"https:\/\/colab.research.google.com\/drive\/1oqRPR_mg8tpw6JxJZlFBdHfyJ6Sp7_t7?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">linked<\/a>).<\/p>\n<pre class=\"wp-block-code\"><code>def train_policy_gradient(env, policy, optimizer, num_episodes=500, gamma=0.99):\n    \"\"\"Train a policy using the REINFORCE algorithm\"\"\"\n    reward_history = []\n\n    for episode in range(num_episodes):\n        state, _ = env.reset()\n        log_probs = []\n        rewards = []\n        done = False\n\n        while not done:\n            state = torch.FloatTensor(state).unsqueeze(0)\n            action_probs = policy(state)\n            action_dist = torch.distributions.Categorical(action_probs)\n            action = action_dist.sample()\n\n            log_probs.append(action_dist.log_prob(action))\n            next_state, reward, done, _, _ = env.step(action.item())\n            rewards.append(reward)\n            state = next_state\n\n        # Compute discounted rewards\n        returns = []\n        G = 0\n        for r in reversed(rewards):\n            G = r + gamma * G\n            returns.insert(0, G)\n\n        returns = torch.tensor(returns)\n        returns = (returns - returns.mean()) \/ (returns.std() + 1e-9)  # Normalize for stability\n\n        # Compute policy gradient loss\n        loss = []\n        for log_prob, G in zip(log_probs, returns):\n            loss.append(-log_prob * G)  # Gradient ascent on expected return\n        loss = torch.stack(loss).sum()\n\n        # Optimize policy\n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n\n        reward_history.append(sum(rewards))\n\n    return reward_history<\/code><\/pre>\n<p>\u00a0 \ud83d\udd17 <a href=\"https:\/\/colab.research.google.com\/drive\/1oqRPR_mg8tpw6JxJZlFBdHfyJ6Sp7_t7?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">Full implementation available here\u00a0<\/a><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-code-explanation\">Code Explanation<\/h3>\n<p>The <b><i>train_policy_gradient <\/i><\/b>function implements the <b><i>REINFORCE <\/i><\/b>algorithm, which optimizes policy parameters using Monte Carlo updates. The training begins by initializing the environment and iterating over multiple episodes, collecting state-action-reward trajectories. For each step in an episode, an action is sampled from the policy, executed in the environment, and its corresponding reward is stored. After completing an episode, the discounted rewards are computed using the <b><i>compute_discounted_rewards <\/i><\/b>function, ensuring that future rewards contribute appropriately to policy updates. These rewards are then normalized to reduce variance, making training more stable. The policy loss is calculated by multiplying the log probabilities of actions by their respective discounted rewards. Finally, the policy is updated using gradient descent, which maximizes the expected return by reinforcing actions that led to higher rewards.<\/p>\n<h4 class=\"wp-block-heading\" id=\"h-expected-outcomes-amp-justification\">Expected Outcomes &amp; Justification<\/h4>\n<p>The training plot demonstrates how the total episode rewards evolve over 500 episodes. Initially, the agent performs poorly, as seen in the low reward values in early episodes (e.g., Episode 50: 20.0). However, as training progresses, the agent learns more effective strategies, leading to higher rewards (Episode 100: 134.0, Episode 150: 229.0). The performance peaks when the agent successfully balances the pole for the maximum time, reaching 500 rewards per episode (Episode 200, 350, and 450). However, instability is evident, as seen in the sharp reward drop in Episode 250 (26.0) and Episode 500 (9.0). This behaviour arises due to the high variance of PG methods, where updates can occasionally lead to suboptimal policies before stabilizing.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"571\" height=\"455\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OP1.webp\" alt=\"Policy Gradient (REINFORCE)\" class=\"wp-image-221855\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OP1.webp 571w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OP1-300x239.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OP1-150x120.webp 150w\" sizes=\"auto, (max-width: 571px) 100vw, 571px\"\/><\/figure>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"323\" height=\"220\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OpP2.webp\" alt=\"Policy Gradient (REINFORCE)\" class=\"wp-image-221853\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OpP2.webp 323w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OpP2-300x204.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PG_OpP2-150x102.webp 150w\" sizes=\"auto, (max-width: 323px) 100vw, 323px\"\/><\/figure>\n<p>The overall trend shows increasing average rewards, indicating that the policy is improving. However, fluctuations in rewards highlight the limitation of vanilla PG methods, which motivates the need for more stable techniques like TRPO and PPO.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-trust-region-policy-optimization-trpo-nbsp\">Trust Region Policy Optimization (TRPO)\u00a0<\/h2>\n<p>While Policy Gradient (PG) methods like REINFORCE are effective, they suffer from high variance and instability in updates. One bad update can drastically collapse the learned policy. TRPO (Trust Region Policy Optimization) improves upon PG by ensuring updates are constrained within a trust region, preventing abrupt changes that could harm performance.<\/p>\n<p>Instead of using vanilla gradient descent, TRPO solves a constrained optimization problem:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"399\" height=\"52\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq1.webp\" alt=\"Trust Region Policy Optimization (TRPO)\u00a0\" class=\"wp-image-221851\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq1.webp 399w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq1-300x39.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq1-150x20.webp 150w\" sizes=\"auto, (max-width: 399px) 100vw, 399px\"\/><\/figure>\n<p>This KL-divergence constraint ensures that the new policy is not too far from the previous policy, leading to more stable updates.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-trpo-algorithm-amp-key-mathematical-concepts\">TRPO Algorithm &amp; Key Mathematical Concepts<\/h2>\n<p>TRPO optimizes the policy using Generalized Advantage Estimation (GAE) and Conjugate Gradient Descent.<\/p>\n<p><b>1. Generalized Advantage Estimation (GAE):<\/b> Computes an advantage function to estimate how much better an action is compared to the expected return.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"191\" height=\"89\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq2.webp\" alt=\"Generalized Advantage Estimation (GAE)\" class=\"wp-image-221850\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq2.webp 191w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq2-150x70.webp 150w\" sizes=\"auto, (max-width: 191px) 100vw, 191px\"\/><\/figure>\n<p>\u00a0 where <b><i>\u03b4_t<\/i><\/b>\u00a0is the TD error:\u00a0\u00a0<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"299\" height=\"59\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq3.webp\" alt=\"Generalized Advantage Estimation (GAE)\" class=\"wp-image-221849\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq3.webp 299w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq3-150x30.webp 150w\" sizes=\"auto, (max-width: 299px) 100vw, 299px\"\/><\/figure>\n<p><b>2.\u00a0Trust Region Constraint:<\/b> Ensures updates stay within a safe region using KL-divergence.<\/p>\n<p>\u00a0 where <b><i>\u03b4\u00a0\u00a0<\/i><\/b>is the maximum step size.\u00a0\u00a0<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"215\" height=\"68\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq4.webp\" alt=\"Trust Region Constraint\" class=\"wp-image-221848\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq4.webp 215w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_Eq4-150x47.webp 150w\" sizes=\"auto, (max-width: 215px) 100vw, 215px\"\/><\/figure>\n<p><b>3.\u00a0 Conjugate Gradient Optimization:<\/b> Instead of directly computing the inverse Hessian, TRPO uses a conjugate gradient to find the optimal update direction efficiently.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-code-example-trpo-training-loop\">Code Example: TRPO Training Loop<\/h2>\n<p>Below is the main TRPO training function, where we apply trust region updates and compute the discounted rewards and advantages. (Only the key function is shown; the full notebook <a href=\"https:\/\/colab.research.google.com\/drive\/15g81iXXCtVvAkSXuVcd9zoqobyrl6kCW?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">link<\/a>.)<\/p>\n<pre class=\"wp-block-code\"><code>def train_trpo(env, policy, num_episodes=500, gamma=0.99):\n    reward_history = []\n\n    for episode in range(num_episodes):\n        state = env.reset()\n        if isinstance(state, tuple):\n            state = state[0]  # Handle Gym versions that return (state, info)\n\n        log_probs = []\n        states = []\n        actions = []\n        rewards = []\n\n        done = False\n        while not done:\n            state_tensor = torch.tensor(state, dtype=torch.float32).unsqueeze(0)\n            probs = policy(state_tensor)\n            action_dist = torch.distributions.Categorical(probs)\n            action = action_dist.sample()\n\n            step_result = env.step(action.item())\n\n            if len(step_result) == 5:\n                next_state, reward, terminated, truncated, _ = step_result\n                done = terminated or truncated  # New Gym API\n            else:\n                next_state, reward, done, _ = step_result  # Old Gym API\n\n            log_probs.append(action_dist.log_prob(action))\n            states.append(state_tensor)\n            actions.append(action)\n            rewards.append(reward)\n\n            state = next_state\n\n        # Compute discounted rewards and advantages\n        discounted_rewards = compute_discounted_rewards(rewards, gamma)\n        discounted_rewards = (discounted_rewards - discounted_rewards.mean()) \n        \/ (discounted_rewards.std() + 1e-9)\n\n        # Convert lists to tensors\n        states = torch.cat(states)\n        actions = torch.tensor(actions)\n        advantages = discounted_rewards\n\n        # Copy old policy before updating\n        old_policy = PolicyNetwork(env.observation_space.shape[0], \n        env.action_space.n)\n        old_policy.load_state_dict(policy.state_dict())\n\n        # Apply TRPO update\n        trpo_step(policy, old_policy, states, actions, advantages)\n\n        total_episode_reward = sum(rewards)\n        reward_history.append(total_episode_reward)\n\n        if (episode + 1) % 50 == 0:\n            print(f\"Episode {episode+1}, Total Reward: {total_episode_reward}\")\n\n    return reward_history<\/code><\/pre>\n<p>\u00a0 \ud83d\udd17\u00a0<a href=\"https:\/\/colab.research.google.com\/drive\/15g81iXXCtVvAkSXuVcd9zoqobyrl6kCW?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">Full implementation available\u00a0here\u00a0<\/a><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-code-explanation-0\">Code Explanation<\/h3>\n<p>The <b><i>train_trpo <\/i><\/b>function implements the Trust Region Policy Optimization update. The training loop initializes the environment and runs 500 episodes, collecting states, actions, and rewards for each step. The key difference from Policy Gradient (PG) is that TRPO maintains an old policy copy and updates the new policy while ensuring the update remains within a KL-divergence bound.<\/p>\n<p>The advantages are computed using discounted rewards and normalized to reduce variance. Finally, conjugate gradient descent is used to determine the optimal policy step direction. Unlike standard gradient updates, TRPO restricts step size to prevent drastic policy changes, leading to more stable performance.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-expected-outcomes-amp-justification-0\">Expected Outcomes &amp; Justification<\/h3>\n<p>The training curve for TRPO exhibits significant reward fluctuations, and the numerical results indicate that the policy does not consistently improve over time as shown below.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"562\" height=\"455\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OP1.webp\" alt=\"Expected Outcomes &amp; Justification\" class=\"wp-image-221847\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OP1.webp 562w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OP1-300x243.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OP1-150x121.webp 150w\" sizes=\"auto, (max-width: 562px) 100vw, 562px\"\/><\/figure>\n<figure class=\"wp-block-image size-full is-resized figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"319\" height=\"231\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OpP2.webp\" alt=\"Expected Outcomes &amp; Justification\" class=\"wp-image-221846\" style=\"width:319px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OpP2.webp 319w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OpP2-300x217.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/TRPO_OpP2-150x109.webp 150w\" sizes=\"auto, (max-width: 319px) 100vw, 319px\"\/><\/figure>\n<p>Unlike Policy Gradient (PG), which showed steady learning progress, TRPO struggles to maintain consistent improvements. Despite its theoretical advantages (trust region constraint preventing catastrophic updates), the actual results show high instability. The total rewards oscillate between low values (9-20), indicating that the agent fails to learn an optimal strategy efficiently.<\/p>\n<p>This is a known issue with TRPO\u2014it requires careful tuning of KL divergence constraints, and in many cases, the update process is computationally expensive and prone to suboptimal convergence. The reward fluctuations suggest that the agent isn\u2019t exploiting learned knowledge effectively, reinforcing the need for a more practical and robust policy optimization method.\u00a0PPO simplifies TRPO by approximating the trust region constraint using a clipped objective function, leading to faster and more efficient training.\u00a0<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-proximal-policy-optimization-ppo\">Proximal Policy Optimization (PPO)<\/h2>\n<p>TRPO ensures stable policy updates but is computationally expensive due to solving a constrained optimization problem at each step. PPO (Proximal Policy Optimization) simplifies this process by using a clipped objective function to restrict updates without requiring second-order optimization.<\/p>\n<p>Instead of solving:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"409\" height=\"74\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq1.webp\" alt=\"Proximal Policy Optimization (PPO)\" class=\"wp-image-221844\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq1.webp 409w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq1-300x54.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq1-150x27.webp 150w\" sizes=\"auto, (max-width: 409px) 100vw, 409px\"\/><\/figure>\n<p>PPO modifies the objective function by introducing a clipped surrogate loss:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"544\" height=\"66\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq2.webp\" alt=\"Proximal Policy Optimization (PPO)\" class=\"wp-image-221843\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq2.webp 544w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq2-300x36.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq2-150x18.webp 150w\" sizes=\"auto, (max-width: 544px) 100vw, 544px\"\/><\/figure>\n<p>where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b><i>r_t\u200b(\u03b8<\/i><\/b><b><i>)<\/i><\/b> is the probability ratio between new and old policies.<\/li>\n<li><b><i>A_<\/i><\/b><b><i>t\u200b<\/i><\/b> is the advantage estimate.<\/li>\n<li><b><i>\u03f5<\/i><\/b> is a small constant (e.g., 0.2) that limits excessive policy updates.<\/li>\n<\/ul>\n<p>This prevents overshooting updates, making PPO more computationally efficient while retaining TRPO\u2019s stability.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-ppo-algorithm-amp-key-mathematical-concept\">PPO Algorithm &amp; Key Mathematical Concept<\/h2>\n<p><b>1. Advantage Estimation using GAE:<\/b> PPO improves TRPO by using Generalized Advantage Estimation (GAE) to compute stable gradients:\u00a0\u00a0<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"213\" height=\"94\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq3.webp\" alt=\"PPO Algorithm &amp; Key Mathematical Concept\" class=\"wp-image-221840\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq3.webp 213w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq3-150x66.webp 150w\" sizes=\"auto, (max-width: 213px) 100vw, 213px\"\/><\/figure>\n<p>\u00a0 where <b><i>\u03b4_t <\/i><\/b>= <b><i>r_t\u00a0<\/i><\/b>+\u00a0<b><i>\u03b3V(s_(t+1))\u00a0<\/i><\/b>\u2212<b><i>V(s_t)<\/i><\/b>.\u00a0\u00a0<\/p>\n<p><b>2.\u00a0<\/b><strong>Clipped Objective Function:<\/strong> Unlike TRPO, which enforces a strict KL constraint, PPO approximates the constraint using clipping:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"264\" height=\"61\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq4.webp\" alt=\"PPO Algorithm &amp; Key Mathematical Concept\" class=\"wp-image-221839\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq4.webp 264w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_Eq4-150x35.webp 150w\" sizes=\"auto, (max-width: 264px) 100vw, 264px\"\/><\/figure>\n<p>This ensures that the update does not move too far, preventing policy collapse.<\/p>\n<p><b>3.\u00a0<\/b><strong>Mini-Batch Training:<\/strong> Instead of updating the policy after each episode, PPO trains using mini-batches over multiple epochs, improving sample efficiency.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-code-example-ppo-training-loop\">Code Example: PPO Training Loop<\/h2>\n<p>Below is the main PPO training function, where we compute advantages, apply clipped policy updates, and use mini-batches for stable learning. (Only the key function is shown; full notebook <a href=\"https:\/\/colab.research.google.com\/drive\/1Qphha6EKPR12s2Yt71qd1UPFXR2KFiw3?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">link<\/a>.)<\/p>\n<pre class=\"wp-block-code\"><code>def train_ppo(env, policy, optimizer, num_episodes=500, gamma=0.99, lambda_=0.95, epsilon=0.2, batch_size=32, epochs=5):\n    reward_history = []\n\n    for episode in range(num_episodes):\n        state = env.reset()\n        if isinstance(state, tuple):\n            state = state[0]  # Handle Gym versions returning (state, info)\n\n        log_probs = []\n        values = []\n        states = []\n        actions = []\n        rewards = []\n\n        done = False\n        while not done:\n            state_tensor = torch.tensor(state, dtype=torch.float32).unsqueeze(0)\n            probs = policy(state_tensor)\n            action_dist = torch.distributions.Categorical(probs)\n            action = action_dist.sample()\n\n            step_result = env.step(action.item())\n\n            # Handle different Gym API versions\n            if len(step_result) == 5:\n                next_state, reward, terminated, truncated, _ = step_result\n                done = terminated or truncated  # New API\n            else:\n                next_state, reward, done, _ = step_result  # Old API\n\n            log_probs.append(action_dist.log_prob(action))\n            states.append(state_tensor)\n            actions.append(action)\n            rewards.append(reward)\n\n            state = next_state\n\n        # Compute advantages\n        values = [0] * len(rewards)  # Placeholder for value estimates (since we use policy-only PPO)\n        advantages = compute_advantages(rewards, values, gamma, lambda_)\n        advantages = (advantages - advantages.mean()) \/ (advantages.std() + 1e-9)  # Normalize advantages\n\n        # Convert lists to tensors\n        states = torch.cat(states)\n        actions = torch.tensor(actions)\n        old_log_probs = torch.tensor(log_probs)\n\n        # PPO Training Loop\n        for _ in range(epochs):\n            for i in range(0, len(states), batch_size):\n                batch_indices = slice(i, i + batch_size)\n\n                new_probs = policy(states[batch_indices])\n                new_action_dist = torch.distributions.Categorical(new_probs)\n                new_log_probs = new_action_dist.log_prob(actions[batch_indices])\n\n                loss = ppo_loss(old_log_probs[batch_indices], new_log_probs, advantages[batch_indices], epsilon)\n\n                optimizer.zero_grad()\n                loss.backward()\n                optimizer.step()\n\n        total_episode_reward = sum(rewards)\n        reward_history.append(total_episode_reward)\n\n        if (episode + 1) % 50 == 0:\n            print(f\"Episode {episode+1}, Total Reward: {total_episode_reward}\")\n\n    return reward_history<\/code><\/pre>\n<p>\ud83d\udd17\u00a0<a href=\"https:\/\/colab.research.google.com\/drive\/1Qphha6EKPR12s2Yt71qd1UPFXR2KFiw3?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">Full implementation available here\u00a0<\/a><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-code-explanation-1\">Code Explanation<\/h3>\n<p>The <b><i>train_ppo <\/i><\/b>function implements Proximal Policy Optimization (PPO) using a clipped surrogate loss and mini-batch updates. Unlike TRPO, which computes trust region constraints, PPO approximates them by clipping policy updates, making it much more efficient.<\/p>\n<ul class=\"wp-block-list\">\n<li>The function begins by collecting episode trajectories (states, actions, log probabilities, and rewards).<\/li>\n<li>Advantage estimation is computed using Generalized Advantage Estimation (GAE).<\/li>\n<li>Mini-batches are used to update the policy over multiple epochs, improving sample efficiency.<\/li>\n<li>Instead of a strict KL divergence constraint, PPO applies a clipped loss function to prevent destructive updates.<\/li>\n<\/ul>\n<h3 class=\"wp-block-heading\" id=\"h-expected-outcomes-for-ppo\">Expected Outcomes for PPO<\/h3>\n<p>The PPO training curve and numerical results show a clear improvement in policy learning over time:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"571\" height=\"455\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP1.webp\" alt=\"PPO training curve\" class=\"wp-image-221837\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP1.webp 571w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP1-300x239.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP1-150x120.webp 150w\" sizes=\"auto, (max-width: 571px) 100vw, 571px\"\/><\/figure>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"336\" height=\"227\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP2.webp\" alt=\"PPO training curve\" class=\"wp-image-221836\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP2.webp 336w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP2-300x203.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/PPO_OP2-150x101.webp 150w\" sizes=\"auto, (max-width: 336px) 100vw, 336px\"\/><\/figure>\n<h4 class=\"wp-block-heading\" id=\"h-key-observations\">Key Observations:<\/h4>\n<ul class=\"wp-block-list\">\n<li><b>Stable Improvement: <\/b>The early rewards (Ep 50-100) are low, indicating the agent is still exploring.<\/li>\n<li><b>Steady Progress: <\/b>By Episode 200, the total reward surpasses 200, showing the agent is learning a structured policy.<\/li>\n<li><b>Fluctuations Exist, But Recovery is Fast:<\/b> Between Ep 300-400, rewards drop, but PPO stabilizes and quickly rebounds to peak performance (500).<\/li>\n<li>\u00a0<b>Final Convergence:<\/b> The model reaches 500 rewards (max score) by Ep 500, confirming PPO effectively learns an optimal strategy.<\/li>\n<\/ul>\n<p><strong>Compared to TRPO, PPO exhibits:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>Less noisy training<\/li>\n<li>Faster convergence<\/li>\n<li>More efficient sample utilization: These improvements validate PPO\u2019s clipped updates and mini-batch training as a superior approach to policy learning.<\/li>\n<\/ul>\n<p>PPO is excellent for reward-based learning, but it struggles with preference-based fine-tuning in applications like LLMs (e.g., ChatGPT, DeepSeek, Claude, Gemini). DPO (Direct Preference Optimization) improves upon PPO by directly learning from human preference data instead of optimizing pure rewards.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-direct-preference-optimization-dpo-preference-learning-for-llms\">Direct Preference Optimization (DPO) \u2013 Preference Learning for LLMs<\/h2>\n<p>Traditional reinforcement learning (RL) techniques are designed to optimize numerical reward-based objectives. However, Large Language Models (LLMs) like ChatGPT, DeepSeek, Claude, and Gemini require fine-tuning that aligns with human preferences rather than just maximizing a reward function. This is where Direct Preference Optimization (DPO) plays a crucial role. Unlike RL-based methods like PPO, which rely on an explicitly trained reward model, DPO optimizes models directly using human feedback. By leveraging preference pairs (where one response is preferred over another), DPO enables models to learn human-like responses efficiently.<\/p>\n<p>DPO eliminates the need for a separate reward model, making it a simpler and more data-driven approach compared to Reinforcement Learning from Human Feedback (RLHF). Instead of reward-based fine-tuning, DPO updates the model parameters to increase the probability of preferred responses while decreasing the probability of rejected responses. This makes the training process more stable and avoids the complexities of RL algorithms like PPO, which involve constrained policy updates and KL penalties.<\/p>\n<p>The significance of DPO lies in its ability to fine-tune LLMs in a way that ensures better response alignment with human expectations. By removing explicit reward models, it prevents the instability often associated with RL-based fine-tuning. Moreover, DPO reduces the risk of harmful, misleading, or biased outputs, making LLMs safer and more reliable. This streamlined optimization process makes it a practical alternative to RL-based fine-tuning, especially when human preference data is available at scale.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-the-dpo-training-dataset\">The DPO Training Dataset<\/h3>\n<p>For DPO, we use human preference data, where each prompt has a preferred response and a rejected response.<\/p>\n<p>Example Preference Dataset (Used for Fine-Tuning)<\/p>\n<pre class=\"wp-block-code\"><code>preference_data = [\n    {\"prompt\": \"What is the capital of France?\",\n     \"preferred\": \"The capital of France is Paris.\",\n     \"rejected\": \"France is a country in Europe.\"},\n\n    {\"prompt\": \"Who wrote Hamlet?\",\n     \"preferred\": \"Hamlet was written by William Shakespeare.\",\n     \"rejected\": \"Hamlet is an old book.\"},\n\n    {\"prompt\": \"Tell me a joke.\",\n     \"preferred\": \"Why did the scarecrow win an award? Because he was outstanding in his field!\",\n     \"rejected\": \"I don\u2019t know any jokes.\"},\n\n    {\"prompt\": \"What is artificial intelligence?\",\n     \"preferred\": \"Artificial intelligence is the simulation of human intelligence in machines.\",\n     \"rejected\": \"AI is just robots.\"},\n\n    {\"prompt\": \"How to stay motivated?\",\n     \"preferred\": \"Set clear goals, track progress, and reward yourself for achievements.\",\n     \"rejected\": \"Just be motivated.\"},\n]\n<\/code><\/pre>\n<p>The preferred responses are accurate, informative, and well-structured, while the rejected responses are vague, incorrect, or unhelpful.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-the-dpo-loss-function\">The DPO Loss Function<\/h3>\n<p>DPO is formulated as a pairwise ranking problem between a preferred response and a rejected response for the same prompt. The goal is to increase the log probability of preferred responses while decreasing the probability of rejected ones.<\/p>\n<p>Mathematically, the DPO objective is:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"610\" height=\"62\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_Eq1.webp\" alt=\"The DPO Loss Function\" class=\"wp-image-221833\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_Eq1.webp 610w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_Eq1-300x30.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_Eq1-150x15.webp 150w\" sizes=\"auto, (max-width: 610px) 100vw, 610px\"\/><\/figure>\n<p>Where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b>y^+<\/b> is the preferred response<\/li>\n<li><b><i>y^- <\/i><\/b>is the rejected response<\/li>\n<li><i><b>\u03b2<\/b><\/i> is a scaling hyperparameter controlling preference strength<\/li>\n<li><b><i>P_\u03b8(y\u2223x)<\/i><\/b> is the log probability of generating a response given input <b><i>x<\/i><\/b><\/li>\n<\/ul>\n<p>This is similar to logistic regression, where the model maximizes separation between preferred and rejected responses.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-code-example-direct-preference-optimization-dpo\">Code Example: Direct Preference Optimization (DPO)<\/h2>\n<p>DPO fine-tunes LLMs by training on human-labeled preference pairs. The core logic of DPO training involves optimizing model weights based on preferred vs. rejected responses. The function below trains a transformer-based model to increase the likelihood of preferred responses while decreasing the likelihood of rejected ones. Below is the key function for computing the DPO loss and updating the model (only the main function is shown for scope; full notebook is <a href=\"https:\/\/colab.research.google.com\/drive\/1MPEYbpJRjocNOCTPouIKOXlPWPVXv753?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">linked<\/a>).<\/p>\n<pre class=\"wp-block-code\"><code>def dpo_loss(preferred_log_probs, rejected_log_probs, beta=0.1):\n    \"\"\"Computes the DPO loss function to optimize based on preferences\"\"\"\n    return -torch.mean(torch.sigmoid(beta * (preferred_log_probs - \n    rejected_log_probs)))\n\ndef encode_text(prompt, response):\n    \"\"\"Encodes the prompt + response into tokenized format with proper padding\"\"\"\n    tokenizer.pad_token = tokenizer.eos_token  # Fix padding issue\n    input_text = f\"User: {prompt}\\nAssistant: {response}\"\n\n    inputs = tokenizer(\n        input_text,\n        return_tensors=\"pt\",\n        padding=True,         # Enable padding\n        truncation=True,      # Truncate if too long\n        max_length=512        # Set max length for safety\n    )\n\n    return inputs[\"input_ids\"], inputs[\"attention_mask\"]\n\nloss_history = []  # Store loss values\n\noptimizer = optim.AdamW(model.parameters(), lr=5e-5)\n\nfor epoch in range(10):  # Train for 10 epochs\n    total_loss = 0\n\n    for data in preference_data:\n        prompt, preferred, rejected = data[\"prompt\"], data[\"preferred\"], \n        data[\"rejected\"]\n\n        # Encode preferred and rejected responses\n        pref_input_ids, pref_attention_mask = encode_text(prompt, preferred)\n        rej_input_ids, rej_attention_mask = encode_text(prompt, rejected)\n\n        # Get log probabilities from the model\n        preferred_logits = model(pref_input_ids, attention_mask=\n        pref_attention_mask).logits[:, -1, :]\n        rejected_logits = model(rej_input_ids, attention_mask=rej_attention_mask)\n        .logits[:, -1, :]\n\n        preferred_log_probs = preferred_logits.log_softmax(dim=-1)\n        rejected_log_probs = rejected_logits.log_softmax(dim=-1)\n\n        # Compute DPO loss\n        loss = dpo_loss(preferred_log_probs, rejected_log_probs, beta=0.5)\n\n        # Optimize the model\n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n\n        total_loss += loss.item()\n\n    loss_history.append(total_loss)  # Store loss for visualization\n    print(f\"Epoch {epoch + 1}, Loss: {total_loss:.4f}\")\n<\/code><\/pre>\n<p>\ud83d\udd17\u00a0<a href=\"https:\/\/colab.research.google.com\/drive\/1MPEYbpJRjocNOCTPouIKOXlPWPVXv753?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">Full implementation available here\u00a0<\/a><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-expected-output-amp-analysis\">Expected Output &amp; Analysis<\/h3>\n<p>The outcomes of Direct Preference Optimization (DPO) can be analyzed from multiple angles: loss convergence, probability shifts, and qualitative response improvements. The training loss curve shows a sharp drop in the initial epochs, followed by stabilization, indicating that the model quickly learns to align with human preferences. The plateau in loss suggests that further optimization yields diminishing improvements, confirming effective preference-based fine-tuning.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"574\" height=\"455\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP1.webp\" alt=\"Direct Preference Optimization (DPO)\" class=\"wp-image-221831\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP1.webp 574w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP1-300x238.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP1-150x119.webp 150w\" sizes=\"auto, (max-width: 574px) 100vw, 574px\"\/><\/figure>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"254\" height=\"225\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP3.webp\" alt=\"Direct Preference Optimization (DPO)\" class=\"wp-image-221828\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP3.webp 254w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP3-150x133.webp 150w\" sizes=\"auto, (max-width: 254px) 100vw, 254px\"\/><\/figure>\n<p>The probability shift visualization reveals that preferred responses consistently achieve higher log probabilities than rejected ones. This confirms that DPO successfully adjusts the model\u2019s behaviour, reinforcing the correct responses while suppressing undesired ones. Some variance in probability shifts suggests that certain prompts may still require fine-tuning for optimal alignment.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"562\" height=\"455\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP2.webp\" alt=\"DPO probability shift visualization\" class=\"wp-image-221826\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP2.webp 562w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP2-300x243.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP2-150x121.webp 150w\" sizes=\"auto, (max-width: 562px) 100vw, 562px\"\/><\/figure>\n<p>A direct comparison of model responses before and after DPO fine-tuning highlights clear improvements. Initially, the model fails to generate a joke, instead providing an irrelevant response. After fine-tuning, it attempts humor but still lacks coherence. This demonstrates that while DPO enhances preference alignment, additional refinements or complementary techniques may be required to generate high-quality, structured responses.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"1609\" height=\"126\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP4.webp\" alt=\"DPO\" class=\"wp-image-221825\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP4.webp 1609w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP4-300x23.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP4-768x60.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP4-1536x120.webp 1536w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/DPO_OP4-150x12.webp 150w\" sizes=\"auto, (max-width: 1609px) 100vw, 1609px\"\/><\/figure>\n<p>Although DPO effectively tunes LLMs without an explicit reward function, it lacks the structured policy learning of reinforcement learning-based methods. This is where General Reinforcement Pretraining Optimization (GRPO) by DeepSeek comes in, combining the strengths of DPO and PPO to enhance LLM fine-tuning further. The next section will explore how GRPO refines policy optimization for large-scale models.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-grpo-group-relative-policy-optimization-deepseek-s-approach\">GRPO \u2013 Group Relative Policy Optimization (DeepSeek\u2019s Approach)<\/h2>\n<p>DeepSeek\u2019s Group Relative Policy Optimization (GRPO) is an advanced preference optimization technique that extends Direct Preference Optimization (DPO) while incorporating elements from Proximal Policy Optimization (PPO). Unlike traditional policy optimization methods that operate on single preference pairs, GRPO leverages group-wise preference ranking, enabling better alignment with human feedback in large-scale LLM fine-tuning.<\/p>\n<p>Traditional preference-based optimization methods, such as DPO (Direct Preference Optimization), operate on pairwise comparisons\u2014one preferred and one rejected response. However, this approach fails to scale efficiently when optimizing on large datasets where multiple responses per prompt are ranked in order of preference. To address this limitation, DeepSeek introduced Group Relative Policy Optimization (GRPO), which allows group-based preference ranking rather than just single-pair preference updates. Instead of comparing two responses at a time, GRPO compares all ranked responses within a batch and optimizes the policy accordingly.<\/p>\n<p>Mathematically, GRPO extends DPO\u2019s reward-free optimization by defining an ordered preference ranking among multiple completions and optimizing their relative likelihoods accordingly.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-mathematical-foundation-of-grpo\">Mathematical Foundation of GRPO<\/h2>\n<p>Since this is the main intent behind the blog, we will dive deep into the mathematics of this.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-expected-return-in-preference-optimization\">1. Expected Return in Preference Optimization<\/h3>\n<p>In standard reinforcement learning, the expected return of a policy <b><i>\u03c0_\u03b8\u00a0<\/i><\/b>is:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"324\" height=\"101\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq1.webp\" alt=\"Expected Return in Preference Optimization\" class=\"wp-image-221824\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq1.webp 324w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq1-300x94.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq1-150x47.webp 150w\" sizes=\"auto, (max-width: 324px) 100vw, 324px\"\/><\/figure>\n<p>where <b><i>R(s_t,a_t)<\/i><\/b>\u00a0is the reward at timestep <b><i>t<\/i><\/b>.<\/p>\n<p>However, LLM fine-tuning does not operate in traditional reward-based RL. Instead, we optimize over human preferences, meaning that reward models are unnecessary.<\/p>\n<p>Instead of learning a reward function, GRPO directly optimizes the model parameters to increase the likelihood of higher-ranked responses over lower-ranked ones.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-ranking-based-probability-optimization\">2. Ranking-Based Probability Optimization<\/h3>\n<p>Given a set of responses <b><i>r_1,r_2, \u2026, r_n<\/i><\/b>\u00a0ranked in order of preference, we define a likelihood ratio:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"119\" height=\"82\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq2.webp\" alt=\"Ranking-Based Probability Optimization\" class=\"wp-image-221823\"\/><\/figure>\n<p>where <b><i>x\u00a0<\/i><\/b>is the input prompt, and <b><i>\u03c0_\u03b8\u00a0<\/i><\/b>represents the policy (LLM) parameterized by <i><b>\u03b8<\/b><\/i>. The key objective is to maximize the probability of higher-ranked responses while suppressing the probability of lower-ranked ones.<\/p>\n<p>To enforce relative preference constraints, GRPO optimizes the following pairwise ranking loss across all response pairs:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"649\" height=\"96\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq3.webp\" alt=\"Ranking-Based Probability Optimization\" class=\"wp-image-221821\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq3.webp 649w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq3-300x44.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq3-640x96.webp 640w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq3-150x22.webp 150w\" sizes=\"auto, (max-width: 649px) 100vw, 649px\"\/><\/figure>\n<p>where:<\/p>\n<ul class=\"wp-block-list\">\n<li><b><i>\u03c3(x)<\/i><\/b> is the sigmoid function ensuring probability normalization<\/li>\n<li><b><i>\u03b2<\/i><\/b> is a temperature scaling parameter controlling gradient magnitude.<\/li>\n<li><b><i>\u03c0_<\/i><\/b><b><i>\u03b8\u200b<\/i><\/b> is the policy (LLM).<\/li>\n<li>The sum iterates over all pairs (i, j) where <b><i>r_i<\/i><\/b>, is ranked higher than <b><i>r_j<\/i><\/b>.<\/li>\n<\/ul>\n<p>The KL-regularized version of GRPO adds a penalty term to prevent drastic shifts in model behaviour:<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"820\" height=\"73\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq4.webp\" alt=\"KL-regularized version of GRPO\" class=\"wp-image-221820\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq4.webp 820w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq4-300x27.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq4-768x68.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_Eq4-150x13.webp 150w\" sizes=\"auto, (max-width: 820px) 100vw, 820px\"\/><\/figure>\n<p>where <b><i>D_KL<\/i><\/b>\u200b ensures conservative updates to prevent overfitting.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-data-for-grpo-fine-tuning\">Data for GRPO Fine-Tuning<\/h2>\n<p>Below is an example dataset used to fine-tune an LLM using ranked preferences:<\/p>\n<pre class=\"wp-block-code\"><code>grpo_preference_data = [\n    {\"prompt\": \"What is the capital of France?\",\n     \"responses\": [\n         {\"text\": \"The capital of France is Paris.\", \"rank\": 1},\n         {\"text\": \"Paris is the largest city in France.\", \"rank\": 2},\n         {\"text\": \"Paris is in France.\", \"rank\": 3},\n         {\"text\": \"France is a country in Europe.\", \"rank\": 4}\n     ]},\n\n    {\"prompt\": \"Tell me a joke.\",\n     \"responses\": [\n         {\"text\": \"Why did the scarecrow win an award? Because he was outstanding \n         in his field!\", \"rank\": 1},\n         {\"text\": \"Why did the chicken cross the road? To get to the other side.\",\n          \"rank\": 2},\n         {\"text\": \"Jokes are funny.\", \"rank\": 3},\n         {\"text\": \"I don\u2019t know any jokes.\", \"rank\": 4}\n     ]}\n]<\/code><\/pre>\n<p>Each prompt has multiple responses with assigned ranks. The model learns to increase the probability of higher-ranked responses while reducing the probability of lower-ranked ones.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-code-implementation-group-based-preference-optimization\">Code Implementation: Group-Based Preference Optimization<\/h2>\n<p>Below is the key function for computing the DPO loss and updating the model (only the main function is shown for scope; the full notebook is\u00a0<a href=\"https:\/\/colab.research.google.com\/drive\/1zamZF7qm4qx-tjysJPn1STZbtJFcQaqe?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">linked<\/a>). The GRPO training function processes multiple ranked responses per prompt, optimizing log-likelihood differences while enforcing KL constraints.<\/p>\n<pre class=\"wp-block-code\"><code>def deepseek_grpo_loss(log_probs, rankings, input_ids, beta=1.0, kl_penalty=0.02, epsilon=1e-6):\n    \"\"\"Computes DeepSeek GRPO loss with pairwise ranking and KL regularization.\"\"\"\n    loss_terms = []\n    num_pairs = 0\n\n    log_probs = torch.clamp(log_probs, min=-10, max=10)  # Prevent extreme values\n\n    for i in range(len(rankings)):\n        for j in range(i + 1, len(rankings)):\n            if rankings[i]  0 else torch.tensor(0.0, device=log_probs.device)\n\n    # KL regularization to prevent policy divergence\n    old_logits = base_model(input_ids).logits[:, -1, :]\n    old_log_probs = old_logits.log_softmax(dim=-1)\n\n    kl_div = torch.nn.functional.kl_div(log_probs, old_log_probs.clamp(min=epsilon), reduction=\"batchmean\")\n\n    return loss + (kl_penalty * kl_div.mean())  # Ensure single scalar<\/code><\/pre>\n<h2 class=\"wp-block-heading\" id=\"h-training-loop-for-grpo\">Training Loop for GRPO<\/h2>\n<p>The training loop processes ranked responses, computes loss, and updates the model while enforcing stability constraints.<\/p>\n<pre class=\"wp-block-code\"><code>loss_history = []\nnum_epochs = 15\n\nfor epoch in range(num_epochs):\n    total_loss = 0\n\n    for data in grpo_preference_data:\n        prompt, responses = data[\"prompt\"], data[\"responses\"]\n\n        input_ids, rankings = encode_text(prompt, responses)\n\n        logits = model(input_ids).logits[:, -1, :]\n        log_probs = logits.log_softmax(dim=-1)\n\n        loss = deepseek_grpo_loss(log_probs, rankings, input_ids)\n\n        if torch.isnan(loss):\n            print(f\"Skipping update at epoch {epoch} due to NaN loss.\")\n            continue\n\n        optimizer.zero_grad()\n        loss.backward()\n\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n\n        optimizer.step()\n        total_loss += loss.item()\n\n    loss_history.append(total_loss)\n    scheduler.step()\n    print(f\"Epoch {epoch + 1}, Loss: {total_loss:.4f}\")\n<\/code><\/pre>\n<p>\ud83d\udd17 <a href=\"https:\/\/colab.research.google.com\/drive\/1zamZF7qm4qx-tjysJPn1STZbtJFcQaqe?usp=sharing\" target=\"_blank\" rel=\"nofollow noopener\">Full implementation available here\u00a0<\/a><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-expected-outcome-and-results\">Expected Outcome and Results<\/h2>\n<p>The expected outcomes of GRPO fine-tuning on the LLM, based on the provided outputs, highlight improvements in model optimization and preference-based ranking.<\/p>\n<p>The training loss curve shows a gradual and stable decline over 15 epochs, indicating that the model is learning effectively. Unlike conventional policy optimization methods, GRPO ensures that ranked responses improve without drastic fluctuations, suggesting smooth convergence.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"567\" height=\"455\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP1.webp\" alt=\"DeepSeek GRPO Training Loss Curve\" class=\"wp-image-221818\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP1.webp 567w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP1-300x241.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP1-150x120.webp 150w\" sizes=\"auto, (max-width: 567px) 100vw, 567px\"\/><\/figure>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"227\" height=\"333\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP4.webp\" alt=\"DeepSeek GRPO Training Loss Curve\" class=\"wp-image-221817\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP4.webp 227w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP4-205x300.webp 205w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP4-150x220.webp 150w\" sizes=\"auto, (max-width: 227px) 100vw, 227px\"\/><\/figure>\n<p>The loss value distribution over epochs presents a histogram where most values concentrate around a decreasing trend, showing that GRPO efficiently optimizes the model while maintaining stable loss updates. This distribution further indicates that loss values do not exhibit large variations, preventing instability in preference ranking.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"691\" height=\"470\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP2.webp\" alt=\"distribution of Loss Values over Epochs\" class=\"wp-image-221815\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP2.webp 691w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP2-300x204.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP2-150x102.webp 150w\" sizes=\"auto, (max-width: 691px) 100vw, 691px\"\/><\/figure>\n<p>The log probability distribution before vs. after fine-tuning provides crucial insights into the model\u2019s response generation. The shift in probability distribution suggests that after fine-tuning, the model assigns higher confidence to preferred responses. This shift results in responses that align better with human expectations and rankings.<\/p>\n<figure class=\"wp-block-image size-full figure  mt-2 mb-2 d-table mx-auto\"><img loading=\"lazy\" decoding=\"async\" width=\"700\" height=\"470\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP3.webp\" alt=\"Log Probability Distribution before vs. after fine-tuning\" class=\"wp-image-221813\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP3.webp 700w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP3-300x201.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/02\/GRPO_OP3-150x101.webp 150w\" sizes=\"auto, (max-width: 700px) 100vw, 700px\"\/><\/figure>\n<p>Overall, the expected outcome of GRPO fine-tuning is a well-optimized model capable of generating high-quality responses ranked effectively based on preference learning. This demonstrates why GRPO is an effective alternative to traditional RL methods like PPO or DPO, offering a structured approach to optimizing LLMs without explicit reward models.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-final-model-insights-why-grpo-excels-in-llm-fine-tuning\">Final Model Insights: Why GRPO Excels in LLM Fine-Tuning<\/h2>\n<p>Unlike pairwise DPO and trust-region PPO, GRPO allows LLMs to learn from multiple ranked completions per prompt, significantly improving response quality, stability, and human alignment.<\/p>\n<ul class=\"wp-block-list\">\n<li>More scalable than pairwise methods \u2192 Learns from multiple ranked completions rather than just binary comparisons.<\/li>\n<li>No explicit reward modeling \u2192 Unlike RLHF, GRPO fine-tunes without requiring a trained reward model.<\/li>\n<li>KL regularization stabilizes updates \u2192 Prevents catastrophic shifts in response distribution.<\/li>\n<li>Better generalization across prompts \u2192 Ensures the LLM produces high-quality, human-aligned responses.<\/li>\n<\/ul>\n<p>With reinforcement learning playing an increasingly central role in fine-tuning LLMs, GRPO stands out as the next step in AI preference learning, setting a new standard for human-aligned language modeling.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>Policy optimization techniques play a critical role in reinforcement learning and LLM fine-tuning. Each method\u2014Policy Gradient (PG), Trust Region Policy Optimization (TRPO), Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO)\u2014offers unique advantages and trade-offs. PG serves as the foundation but suffers from high variance, while TRPO provides stability at the cost of computational complexity. PPO, being a refined version of TRPO, balances efficiency and robustness, making it widely used in RL applications. DPO, on the other hand, optimizes LLMs directly using preference data, eliminating the need for a reward model. Finally, GRPO, as introduced by DeepSeek, enhances preference-based fine-tuning by leveraging relative ranking in a structured manner.<\/p>\n<p>Below is a comparison of these LLM Optimization methods based on key aspects such as variance, stability, sample efficiency, and suitability for reinforcement learning versus LLM fine-tuning:<\/p>\n<table>\n<tr>\n<th>Method<\/th>\n<th>Variance<\/th>\n<th>Stability<\/th>\n<th>Sample Efficiency<\/th>\n<th>Best for<\/th>\n<th>Limitations<\/th>\n<\/tr>\n<tr>\n<td>PG (REINFORCE)<\/td>\n<td>High<\/td>\n<td>Low<\/td>\n<td>Inefficient<\/td>\n<td>Simple RL problems<\/td>\n<td>High variance, slow convergence<\/td>\n<\/tr>\n<tr>\n<td>TRPO<\/td>\n<td>Low<\/td>\n<td>High<\/td>\n<td>Moderate<\/td>\n<td>High-stability RL tasks<\/td>\n<td>Complex second-order updates, expensive<\/td>\n<\/tr>\n<tr>\n<td>PPO<\/td>\n<td>Medium<\/td>\n<td>High<\/td>\n<td>Efficient<\/td>\n<td>General RL tasks, Robotics, Games<\/td>\n<td>May require careful hyperparameter tuning<\/td>\n<\/tr>\n<tr>\n<td>DPO<\/td>\n<td>Low<\/td>\n<td>High<\/td>\n<td>High<\/td>\n<td>LLM fine-tuning with human preferences<\/td>\n<td>Lacks explicit reinforcement learning framework<\/td>\n<\/tr>\n<tr>\n<td>GRPO<\/td>\n<td>Low<\/td>\n<td>High<\/td>\n<td>High<\/td>\n<td>Preference-based LLM fine-tuning<\/td>\n<td>Newer method, requires further empirical validation<\/td>\n<\/tr>\n<\/table>\n<p>For practitioners, the choice depends on the task at hand. If optimizing reinforcement learning agents in games or robotics, PPO is the best choice due to its balance of efficiency and performance. If high-stability optimization is required, TRPO is preferred despite its computational cost. DPO and GRPO, however, are better suited for LLM fine-tuning, with GRPO providing an even stronger optimization framework based on relative preference ranking rather than just binary preference signals.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-key-takeaways\">Key Takeaways<\/h3>\n<p>Reinforcement learning (RL) plays a crucial role in both game-playing agents and LLM fine-tuning, but the optimization techniques vary significantly.<\/p>\n<ul class=\"wp-block-list\">\n<li>PG, TRPO, and PPO are fundamental in RL, with PPO being the most practical choice for its efficiency and performance balance.<\/li>\n<li>DPO introduced a major shift in LLM fine-tuning by eliminating explicit reward models, making human preference alignment easier and more efficient.<\/li>\n<li>GRPO, pioneered by DeepSeek, further refines LLM fine-tuning by optimizing for relative ranking rather than just binary comparisons, improving preference-based alignment.<\/li>\n<li>For RL tasks, PPO remains the dominant method, while for LLM fine-tuning, DPO and GRPO are superior choices due to their ability to fine-tune models using direct preference data without RL instability.<\/li>\n<\/ul>\n<p>This blog highlights how reinforcement learning and preference-based fine-tuning are converging, with new techniques like GRPO bridging the gap between structured optimization and real-world deployment of large-scale AI systems.<\/p>\n<p><strong>The media shown in this article is not owned by Analytics Vidhya and is used at the Author\u2019s discretion.<\/strong><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/karthik3852845\/\"\/><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/mimi6\/\"\/><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1739779236761\"><strong class=\"schema-faq-question\">Q1. What is the difference between PPO and DPO?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. PPO (Proximal Policy Optimization) is an RL-based optimization method that improves policies while maintaining stability using a clipping mechanism. It is widely used in reinforcement learning tasks such as robotics and game-playing AI. DPO (Direct Preference Optimization), on the other hand, is designed specifically for LLM fine-tuning, directly optimizing the model based on human preferences without requiring an explicit reward model. DPO is simpler and more efficient for aligning language models with human intent.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739779253497\"><strong class=\"schema-faq-question\">Q2. Why is GRPO better than DPO for preference-based fine-tuning?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. GRPO (Group Relative Policy Optimization) improves upon DPO by optimizing preferences in a ranked manner instead of binary preference signals. While DPO only differentiates between \u201cpreferred\u201d and \u201crejected\u201d responses, GRPO assigns relative rankings across multiple responses, capturing nuanced differences in preference. This allows LLMs to learn more refined distinctions and align better with human feedback.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739779271109\"><strong class=\"schema-faq-question\">Q3. When should I use TRPO over PPO?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. TRPO (Trust Region Policy Optimization) should be used when strict stability constraints are required, such as in high-stakes RL environments (e.g., robotics, autonomous driving). However, it is computationally expensive due to second-order optimization. PPO (Proximal Policy Optimization) provides a more efficient and scalable alternative by approximating TRPO\u2019s constraints using a clipping mechanism, making it the preferred choice in most RL scenarios.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1739779285188\"><strong class=\"schema-faq-question\">Q4. Why do LLMs need preference optimization techniques like DPO and GRPO?<\/strong> <\/p>\n<p class=\"schema-faq-answer\">Ans. Traditional RL methods focus on maximizing numerical rewards, which do not always align with human expectations in language models. DPO and GRPO fine-tune LLMs based on human preference data, ensuring responses are helpful, honest, and harmless. Unlike Reinforcement Learning with Human Feedback (RLHF), these methods eliminate the need for a separate reward model, making fine-tuning more efficient and reducing potential biases from reward misalignment.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/akashdas\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_sT5MuqV.webp\" width=\"48\" height=\"48\" alt=\"Neil D\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Neil is a research professional currently working on the development of AI agents. He has successfully contributed to various AI projects across different domains, with his works published in several high-impact, peer-reviewed journals. His research focuses on advancing the boundaries of artificial intelligence, and he is deeply committed to sharing knowledge through writing. Through his blogs, Neil strives to make complex AI concepts more accessible to professionals and enthusiasts alike.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>For decades,\u00a0Reinforcement Learning (RL)\u00a0has been the driving force behind breakthroughs in\u00a0robotics, game-playing AI (AlphaGo, OpenAI Five), and control systems. RL\u2019s strength lies in its ability to optimize decision-making by\u00a0maximizing long-term rewards, making it ideal for problems requiring sequential reasoning. However,\u00a0large language models (LLMs)\u00a0initially relied on\u00a0supervised learning, where models were fine-tuned on static datasets. This approach [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":94665,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5815,44995,44996,11572],"dealstore":[],"offerexpiration":[],"class_list":["post-94664","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-blogathon","tag-gradient","tag-grpo","tag-policy"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>From Policy Gradient to GRPO - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=94664\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"From Policy Gradient to GRPO - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"For decades,\u00a0Reinforcement Learning (RL)\u00a0has been the driving force behind breakthroughs in\u00a0robotics, game-playing AI (AlphaGo, OpenAI Five), and control systems. RL\u2019s strength lies in its ability to optimize decision-making by\u00a0maximizing long-term rewards, making it ideal for problems requiring sequential reasoning. However,\u00a0large language models (LLMs)\u00a0initially relied on\u00a0supervised learning, where models were fine-tuned on static datasets. This approach [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=94664\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-18T00:54:17+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"29 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=94664#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=94664\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"From Policy Gradient to GRPO\",\"datePublished\":\"2025-02-18T00:54:17+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=94664\"},\"wordCount\":4525,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=94664#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp\",\"keywords\":[\"Blogathon\",\"Gradient\",\"GRPO\",\"Policy\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=94664#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=94664\",\"url\":\"https:\/\/fivemor.com\/?p=94664\",\"name\":\"From Policy Gradient to GRPO - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=94664#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=94664#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp\",\"datePublished\":\"2025-02-18T00:54:17+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=94664#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=94664\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=94664#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=94664#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"From Policy Gradient to GRPO\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"From Policy Gradient to GRPO - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=94664","og_locale":"en_US","og_type":"article","og_title":"From Policy Gradient to GRPO - Som2ny Network","og_description":"For decades,\u00a0Reinforcement Learning (RL)\u00a0has been the driving force behind breakthroughs in\u00a0robotics, game-playing AI (AlphaGo, OpenAI Five), and control systems. RL\u2019s strength lies in its ability to optimize decision-making by\u00a0maximizing long-term rewards, making it ideal for problems requiring sequential reasoning. However,\u00a0large language models (LLMs)\u00a0initially relied on\u00a0supervised learning, where models were fine-tuned on static datasets. This approach [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=94664","og_site_name":"Som2ny Network","article_published_time":"2025-02-18T00:54:17+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"29 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=94664#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=94664"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"From Policy Gradient to GRPO","datePublished":"2025-02-18T00:54:17+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=94664"},"wordCount":4525,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=94664#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp","keywords":["Blogathon","Gradient","GRPO","Policy"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=94664#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=94664","url":"https:\/\/fivemor.com\/?p=94664","name":"From Policy Gradient to GRPO - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=94664#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=94664#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp","datePublished":"2025-02-18T00:54:17+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=94664#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=94664"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=94664#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/From-Policy-Gradient-to-GRPO-A-Deep-Dive-into-LLM-Optimization-with-Code-Math-1.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=94664#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"From Policy Gradient to GRPO"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/94664","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=94664"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/94664\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/94665"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=94664"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=94664"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=94664"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=94664"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=94664"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}