{"id":22566,"date":"2025-01-11T09:46:15","date_gmt":"2025-01-11T09:46:15","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/a-comprehensive-guide-to-observed-power-post-hoc-power\/"},"modified":"2025-01-11T09:46:15","modified_gmt":"2025-01-11T09:46:15","slug":"a-comprehensive-guide-to-observed-power-post-hoc-power","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=22566","title":{"rendered":"A Comprehensive Guide to Observed Power (Post Hoc Power)"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>\u201cObserved power\u201d, \u201cPost hoc Power\u201d, and \u201cRetrospective power\u201d all refer to the statistical power of a statistical significance test to detect a true effect equal to the observed effect. In a broader sense these terms may also describe any power analysis performed after an experiment has completed. Importantly, it is the first, narrower sense that will be used throughout the article.<\/p>\n<p>Here are three <strong>typical scenarios of using observed power<\/strong> in online A\/B testing and in more general applications of statistical tests:<\/p>\n<ol class=\"wp-block-list\">\n<li>An experiment concludes with a statistically non-significant outcome. Observed power is compared to the target power to determine if the test was underpowered. If that is the case it is argued that there is a true effect which remains undetected. This can lead to a larger follow-up experiment or in some cases to extending the experiment until it achieves whatever level of observed power level is deemed sufficient;<\/li>\n<li>An experiment concludes with a statistically significant outcome, but the observed power is smaller than the target power level of, say, 80% or 90%. This is seen as the test being underpowered and therefore questionable or outright unreliable;<\/li>\n<li>An experiment concludes with a statistically significant outcome, but the observed effect size is lower than the target MDE (minimum detectable effect), and that is seen as cause for concern for its trustworthiness;<\/li>\n<\/ol>\n<p>While these might seem reasonable, there are unsurmountable issues in using post hoc power. These have been covered with increasing frequency in the scientific literature <sup>[1][2][3][4][5]<\/sup>, yet the problem persists in statistical practice.<\/p>\n<p>This article takes certain novel approaches to debunking observed power, including using a post hoc power calculator as an educational device. Through a comprehensive examination of the rationale behind observed power <strong>all possible use cases are shown to rest on valid, but unsound chains of reasoning<\/strong>, which might explain their surface appeal.<\/p>\n<p>This and the two companion articles further explore the <strong>different harms caused by using observed power<\/strong> in scenarios readers should find familiar, in a uniquely comprehensive manner.<\/p>\n<p><strong>Table of Contents:<\/strong><\/p>\n<ol class=\"wp-block-list\">\n<li><a href=\"#observedandtruepower\">Observed power as a proxy for true power<\/a><br \/>1.1. <a href=\"#hightruepower\">High true power is not a requirement for a good test<\/a><br \/>1.2. <a href=\"#observedpowerasestimate\">Observed power as an estimate of true power<\/a><\/li>\n<li><a href=\"#powerasthreshold\">Viewing the power level as a threshold to be met<\/a><br \/>2.1. <a href=\"#lowobservedunderpowered\">Low observed power does not mean a test is underpowered<\/a><\/li>\n<li><a href=\"#observedpowerandpvalues\">Observed power is a direct product of the p-value<\/a><br \/>3.1. <a href=\"#posthocpowercalculator\">Post hoc power calculator<\/a><\/li>\n<li><a href=\"#opnonsig\">Observed power with a non-significant outcome<\/a><br \/>4.1. <a href=\"#obpeeking\">Does peeking through observed power inflate the false positive rate?<\/a><\/li>\n<li><a href=\"#opstatsig\">Observed power with a statistically significant outcome?<\/a><br \/>5.1. <a href=\"#mdethreshold\">MDE as a threshold<\/a><br \/>5.2. <a href=\"#effectsizeconfusion\">Viewing the observed effect size as the one a test\u2019s conclusion is about<\/a><\/li>\n<li><a href=\"#propterhocpower\">All power is propter hoc power<\/a><\/li>\n<li><a href=\"#takeaways\">Takeaways<\/a><\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"observedandtruepower\">Observed power as a proxy for true power<\/h2>\n<p>A major line of reasoning behind the use of observed power is what may be called <em><strong>\u201cthe search for true power\u201d<\/strong><\/em>.<\/p>\n<p>The logic of searching for \u201ctrue power\u201d<strong>*<\/strong> goes like this:<\/p>\n<ol class=\"wp-block-list\">\n<li>\u201cGood\u201d or \u201ctrustworthy\u201d statistical tests should have high \u201ctrue power\u201d: high probability of detecting the true effect, whatever it may be;<\/li>\n<li>The true effect size is unknown and there is often little data to produce good pre-test predictions about where it lies;<\/li>\n<li>Post data, the observed effect (or an appropriate point estimate), is our best guess for the true effect size;<\/li>\n<li>Substituting the true effect for the observed effect in a power calculation should result in a good estimate of the \u201ctrue power\u201d of a test.<\/li>\n<\/ol>\n<p>These premises lead to the following conclusion:<\/p>\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>To ensure a test is \u201cgood\u201d or \u201ctrustworthy\u201d, aim to achieve high power versus the observed effect as the best proxy for the true statistical power of the test.<\/p>\n<\/blockquote>\n<p>While the above is a valid argument for using observed power, it is not a sound one since premises one and four can be shown to be false.<\/p>\n<details class=\"wp-block-details is-layout-flow wp-block-details-is-layout-flow\">\n<summary>* side remark on \u201ctrue power\u201d<\/summary>\n<p>The focus some people have on \u201ctrue power\u201d may stem from poor and incomplete definitions of power found in multiple textbooks and articles. Namely, one can often see power described as \u201cthe probability of finding an effect, if there is one\u201d. However, this is a misleading definition since it skips the all important \u201d of a given size\u201d instead of the easily assumed \u201cof any size\u201d. Such a confusion result may lead to searches of \u201ctrue power\u201d instead of acceptance of its hypothetical nature. Compare the above shorthand to the definition given in \u201cStatistical Methods in Online A\/B Testing\u201d: \u201cThe statistical power of a statistical test is defined as the probability of observing a p-value statistically significant at a certain threshold \u03b1 if a true effect of a certain magnitude \u03bc1 is in fact present.[\u2026]\u201d<\/p>\n<\/details>\n<h3 class=\"wp-block-heading\" id=\"hightruepower\">High true power is not a requirement for a good test<\/h3>\n<p>When planning an A\/B test or another type of experiment, one often does not know the true effect size and may have very limited information to make predictions about where it may lie. Planning a test with a primary motivation of achieving a high true power results in committing to as large a sample size as practically achievable and in choosing as low a significance threshold as possible. This would maximize the chance of detecting even tiny true effects.<\/p>\n<p>However, such an approach would ignore the <strong>trade-offs<\/strong> between running a test for longer \/ with more users, and the potential benefits that may be realized (for details see <a href=\"https:\/\/blog.analytics-toolkit.com\/2022\/statistical-power-mde-and-designing-statistical-tests\/\">\u201cStatistical Power, MDE, and Designing Statistical Tests\u201d<\/a>, <a href=\"https:\/\/www.analytics-toolkit.com\/glossary\/risk-reward-analysis\/\" target=\"_blank\" rel=\"noreferrer noopener\">risk-reward analysis<\/a> and its references). True effects below a given value are simply not worth pursuing due to the <a href=\"https:\/\/blog.analytics-toolkit.com\/2017\/costs-benefits-ab-testing-comprehensive-guide\/\" target=\"_blank\" rel=\"noreferrer noopener\">risks of running the test<\/a> no longer being justified by the potential benefits.<\/p>\n<p>The pivotal value where that happens is called a minimum detectable effect (MDE) or a minimum effect of interest (MEI). It is the smallest effect size one wants to detect with high probability, where \u201cdetect\u201d means a statistically significant outcome at a specified significance threshold. It is the smallest effect size one would not want to miss. Aiming for high power for any value below the MEI <strong>would, by definition, achieve poorer business outcomes from A\/B testing<\/strong>.<\/p>\n<p>Even if it is <em>almost sure<\/em> what the true effect size is, it would rarely match the minimum effect of interest. The framework for choosing optimal sample sizes and significance thresholds deployed at <a href=\"https:\/\/www.analytics-toolkit.com\/ab-testing-hub\/\" target=\"_blank\" rel=\"noreferrer noopener\">Analytics Toolkit\u2019s A\/B testing hub<\/a> often arrives at optimal parameters under which the MDE (at a high power level such as 80% or 90%) is larger than the expected effect size. This decision framework is most fully explained in the relevant chapter of <a href=\"https:\/\/www.abtestingstats.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u201cStatistical Methods in Online A\/B Testing\u201d<\/a>.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;67823de786bf3&quot;}\" data-wp-interactive=\"core\/image\" class=\"aligncenter size-large wp-lightbox-container\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"611\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-09-30-AB-Power-MDE-Risk-Reward-Analysis-1024x611.png\" alt=\"\" class=\"wp-image-1819\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-09-30-AB-Power-MDE-Risk-Reward-Analysis-1024x611.png 1024w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-09-30-AB-Power-MDE-Risk-Reward-Analysis-500x299.png 500w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-09-30-AB-Power-MDE-Risk-Reward-Analysis-768x459.png 768w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-09-30-AB-Power-MDE-Risk-Reward-Analysis.png 1117w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><button class=\"lightbox-trigger\" type=\"button\" aria-haspopup=\"dialog\" aria-label=\"Enlarge image\" data-wp-init=\"callbacks.initTriggerButton\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-style--right=\"state.imageButtonRight\" data-wp-style--top=\"state.imageButtonTop\"><br \/>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewbox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\"><\/path>\n\t\t\t<\/svg><br \/>\n\t\t<\/button><figcaption class=\"wp-element-caption\"><strong>Figure 1<\/strong>: The result of a risk-reward analysis in which the minimum effect of interest at 80% is notably higher than the effect considered highly likely (click to see full size image)<\/figcaption><\/figure>\n<\/div>\n<p>Finally, it should be noted that taking premise one as given and aiming for high true power would turn the planning phase of an experiment into an exercise of guessing what the true effect of an intervention is. <strong>The logical conclusion of <em>\u201cthe search for true power\u201d<\/em> renders all pre-test power analyses and test planning futile.<\/strong> However, good planning is what ensures an experiment, be it an online A\/B test or a scientific experiment, offers a good balance between the expected business risks and rewards.<\/p>\n<p>Rewriting premise one to make it true would result in:<\/p>\n<ol class=\"wp-block-list\">\n<li>\u201cGood\u201d statistical tests should have high power to detect the minimum effect of interest.<\/li>\n<\/ol>\n<p>Since the minimum effect of interest is rarely the true effect size, the above statement breaks the rest of the argument and makes it unsound. <strong>The search for true power is therefore <em>illogical <\/em>and it does not justify the use of observed power.<\/strong> This holds even if observed power was a good estimate of true power. Whether that is the case is explored in the following section.<\/p>\n<h3 class=\"wp-block-heading\" id=\"observedpowerasestimate\">Observed power as an estimate of true power<\/h3>\n<p>Another reason why searching for true power is not a logical justification for using observed power is that the fourth premise does not hold either. Namely, <strong>the observed power of a test is a poor estimate of its true power.<\/strong> To be a proper estimator of true power, observed power should ideally be all of the following:<\/p>\n<p>Some of these conditions can be relaxed in complex modeling problems where an ideal estimator either does not exist or has not yet been found. However, observed power performs really poorly on all accounts:<\/p>\n<ul class=\"wp-block-list\">\n<li>It does not hone in on the true value with larger amounts of information; (is not consistent)<\/li>\n<li>Observed power is almost always heavily biased vis-\u00e0-vis true power;<\/li>\n<li>There is no reason to suspect it is a sufficient statistic of true power;<\/li>\n<li>Its variance is far too large to be useful (even if it somehow turns out to be efficient though biased)<\/li>\n<\/ul>\n<p>Yuan &amp; Maxwell end their work titled \u201cOn the Post Hoc Power in Testing Mean Differences\u201d<sup>[6]<\/sup> with the following assessment:<\/p>\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>\u201cUsing analytical, numerical, and Monte Carlo approaches, our results show that the estimated power does not provide useful information when the true power is small. It is almost always a biased estimator of the true power. The bias can be negative or positive. Large sample size alone does not guarantee the post hoc power to be a good estimator of the true power.\u201d<\/p>\n<\/blockquote>\n<p>I\u2019ve conducted some R simulations of my own to visualize how poorly observed power behaves as an estimate of true power:<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;67823de7874ab&quot;}\" data-wp-interactive=\"core\/image\" class=\"aligncenter size-large wp-lightbox-container\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"619\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-03-Observed-Power-versus-MLE-1024x619.png\" alt=\"\" class=\"wp-image-1821\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-03-Observed-Power-versus-MLE-1024x619.png 1024w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-03-Observed-Power-versus-MLE-500x302.png 500w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-03-Observed-Power-versus-MLE-768x464.png 768w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-03-Observed-Power-versus-MLE.png 1208w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><button class=\"lightbox-trigger\" type=\"button\" aria-haspopup=\"dialog\" aria-label=\"Enlarge image\" data-wp-init=\"callbacks.initTriggerButton\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-style--right=\"state.imageButtonRight\" data-wp-style--top=\"state.imageButtonTop\"><br \/>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewbox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\"><\/path>\n\t\t\t<\/svg><br \/>\n\t\t<\/button><figcaption class=\"wp-element-caption\"><strong>Figure 2<\/strong>: Simulations of Observed Power versus MLE as estimators for true power and the true effect size, respectively (click to see full size image)<\/figcaption><\/figure>\n<\/div>\n<p>All three simulations have 100,000 runs and share the same properties except for the different true effect and hence true power. Histograms in each column come from the same simulation and offer a direct comparison which reveals just how poorly observed power does as an estimate compared to the observed effect size which in this case is the <a href=\"https:\/\/www.analytics-toolkit.com\/glossary\/maximum-likelihood-estimate\/\" target=\"_blank\" rel=\"noreferrer noopener\">MLE<\/a>. In short, it has inconsistent variance, the bias is obvious, and consistency is obviously non-existent.<\/p>\n<h2 class=\"wp-block-heading\" id=\"powerasthreshold\">Viewing the power level as a threshold to be met<\/h2>\n<p>Another error shared by different use-cases of observed power is to view the choice of statistical power level (1 \u2013 \u03b2) and its complement, the type II error rate \u03b2, as being in the same category as the choice of the significance threshold \u03b1 and the resulting type I error rate. The logic is as follows:<\/p>\n<ol class=\"wp-block-list\">\n<li>The test\u2019s chosen power level is 80% (or 90%, etc.) against a given effect size of interest;<\/li>\n<li>The pre-test power level is a threshold which should be met for the outcome to be considered trustworthy;<\/li>\n<li>The observed power is what needs to be compared to the target power level to ensure said threshold is met;<\/li>\n<\/ol>\n<p>Naturally, the conclusion is that:<\/p>\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>If the observed power is lower than the power level , then the test is underpowered as it has failed to achieve its statistical power threshold \/ target type II error rate.<\/p>\n<\/blockquote>\n<p>The above logic is unsound due to issues with premises two and three. Regarding premise two, the test\u2019s target power level is not a cut-off that needs to be met by any statistic of the observed data in order to declare or conclude something. The power level and the type II error rate are inherently pre-data probabilities. The level of statistical power against a chosen MDE (MEI) is achieved by the experiment reaching (or exceeding) the sample size determined to be necessary by way of a pre-test power analysis *.<\/p>\n<details class=\"wp-block-details is-layout-flow wp-block-details-is-layout-flow\">\n<summary>* the obvious exception<\/summary>\n<p>The only exception to this may be when the observed variance of the outcome measure is meaningfully different from the variance used in the power analysis and sample size calculations. For example, the base rate of a binomial variable may turn out to be significantly lower than expected, in which case the analysis may need to be rerun and a new sample size may need to be calculated. In e-commerce this may happen if seasonality or a significant change in the advertising strategy cause a sudden influx of prospective customers who are less qualified or less willing to purchase than previous ones, among other reasons. Even when that happens to be the case, only the observed variance may play a role in this sample size recalculation, not the observed effect size, so one cannot speak of observed power in the sense defined at the beginning paragraph of this article.<\/p>\n<\/details>\n<p><strong>Once an experiment reaches its target sample size it has achieved its desired statistical power<\/strong> level against the minimum effect of interest regardless of the outcome and any other statistic.<\/p>\n<p>Regarding premise three, it would be useful to point out that the significance threshold \u03b1 is a cut-off set for the <a href=\"https:\/\/www.analytics-toolkit.com\/glossary\/p-value\/\" target=\"_blank\" rel=\"noreferrer noopener\">p-value<\/a>. The p-value is a post-data summary of the data showing how likely (or unlikely) it would have been to observe said data under a specified null hypothesis. The significance threshold applies directly to the observed p-value and leads to different conclusions like \u201creject the null hypothesis\u201d or \u201cfail to reject the null hypothesis\u201d.<\/p>\n<p>The target level of statistical power is not a cut-off in the same way and there is no equivalent to the p-value for statistical power. Observed power surely cannot perform the function of the p-value since power calculations do not use any of the data obtained from an experiment. Making the hypothetical true effect equal to the observed effect does not make the power calculation any more informative.<\/p>\n<h3 class=\"wp-block-heading\" id=\"lowobservedunderpowered\">Low observed power does not mean a test is underpowered<\/h3>\n<p>As a consequence of the above, a test cannot be said to be underpowered based on the observed power being lower than the target power (see <a href=\"https:\/\/blog.analytics-toolkit.com\/2020\/underpowered-a-b-tests-confusions-myths-reality\/\" target=\"_blank\" rel=\"noreferrer noopener\">Underpowered A\/B Tests \u2013 Confusions, Myths, and Reality<\/a>). A test has either reached its target sample size or not and observed power cannot tell us anything about that. In this light both \u201cobserved power\u201d and \u201cpost-hoc power\u201d are bordering on being misnomers. More on this point in the last section of this article.<\/p>\n<h2 class=\"wp-block-heading\" id=\"observedpowerandpvalues\">Observed power is a direct product of the p-value<\/h2>\n<p>To some, the fact that observed power can be computed as a direct transformation of the p-value is the nail in the coffin of the concept. However, this has been known since at least Hoenig &amp; Heisey\u2019s 2001<sup>[1]<\/sup> article in The American Statistician. Yet, fallacies related to observed power persist. This is why I have tackled some of the logical issues with its application before turning to this obvious point.<\/p>\n<p>What is meant by a direct product? Given any p-value <em>p<\/em> from a given test of size <em>\u03b1<\/em>, the observed power can be computed analytically using the following equation<sup>[1]<\/sup>:<\/p>\n<p class=\"has-text-align-center\"><strong>G<sub>Zp<\/sub>(\u03b1) = 1 \u2212 \u03a6(Z<sub>\u03b1<\/sub> \u2212 Z<sub>p<\/sub>)<\/strong><\/p>\n<p>Where \u03a6 is the <a href=\"https:\/\/www.gigacalculator.com\/calculators\/normal-distribution-calculator.php\" target=\"_blank\" rel=\"noreferrer noopener\">standard normal cumulative distribution function<\/a>, Z<sub>\u03b1<\/sub> and Z<sub>p<\/sub> being the z-scores corresponding to the significance threshold and the p-value, respectively. Observed power is therefore determined completely by the p value and the chosen significance threshold. In other words, <strong>once the p-value is known, there is nothing that calculating observed power can add to the picture.<\/strong><\/p>\n<h3 class=\"wp-block-heading\" id=\"posthocpowercalculator\">Post hoc power calculator<\/h3>\n<p>To illustrate this, I\u2019ve built a simple post hoc power calculator which you can use by entering your significance threshold and the observed p-value. The level of statistical power against the observed effect will be computed instantaneously, and you can explore the relationship of any p-value and the post hoc power it results in by hovering over the graph.<\/p>\n<p><iframe id=\"calcframe\" src=\"https:\/\/www.analytics-toolkit.com\/statistical-power-calculator\/post-hoc-power-calculator.php\" sandbox=\"allow-same-origin allow-scripts allow-popups allow-forms\" style=\"display: block; border: none; width: 100%; max-width: 500px; margin: 0px auto 20px auto; padding: 0px; height: 666px; visibility: visible; opacity: 1;\"><\/iframe><\/p>\n<p>The green vertical line represents the significance threshold, whereas the violet line is the observed p-value. Feel free to test it against some of your own, independent calculations of statistical power.<\/p>\n<p>The interesting question is what follows from the fact that the observed power is a simple transformation of the p-value? The answer is broken down below by whether the outcome is <a href=\"https:\/\/www.analytics-toolkit.com\/glossary\/statistical-significance\/\" target=\"_blank\" rel=\"noreferrer noopener\">statistically significant<\/a> or not.<\/p>\n<h2 class=\"wp-block-heading\" id=\"opnonsig\">Observed power with a non-significant outcome<\/h2>\n<p>When using observed power with non-significant test outcomes, it is virtually guaranteed to label such tests as \u201cunderpowered\u201d. This is due to the conflation between low true power and a test being underpowered, coupled with the misguided belief that \u201cobserved power\u201d is a good estimate for \u201ctrue power\u201d. These have been debunked in earlier sections.<\/p>\n<p>To compound the issue, the maximum observed power possible with a non-significant outcome is below 0.5 (50%) at any given significance threshold. Examine the fraction of the observed power curve which spans the non-significant outcomes:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"504\" height=\"501\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Observed-Power-for-Non-significant-Outcomes.png\" alt=\"Observed power with non-significant outcomes\" class=\"wp-image-1828\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Observed-Power-for-Non-significant-Outcomes.png 504w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Observed-Power-for-Non-significant-Outcomes-352x350.png 352w\" sizes=\"auto, (max-width: 504px) 100vw, 504px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 3<\/strong>: Observed power with non-significant outcomes<\/figcaption><\/figure>\n<\/div>\n<p>For all non-significant values the power is lower than 50%. In this example a threshold of significance of 0.05 is used, but you can use the post-hoc power calculator above to check that it is the case by entering different thresholds. <\/p>\n<p>In a best case scenario, an observed power of 50% or lower would be interpreted as the test being underpowered which can be used to justify a follow-up test with significantly larger sample size \/ longer duration.<\/p>\n<p>There is also a contradiction with the p-value. As evident from <em>Figure 3<\/em>, the smaller the p-value, the higher the observed power is. Interpreting higher observed power as offering stronger evidence for the null hypothesis <strong>directly contradicts the logic of p-values<\/strong> in which a smaller p-value offers stronger grounds for rejecting the null hypothesis, not weaker.<\/p>\n<p><strong>In the worst possible scenario, low observed power is taken as a grant to continue the test<\/strong> until a given higher sample size is achieved:<\/p>\n<h3 class=\"wp-block-heading\" id=\"obpeeking\">Does peeking through observed power inflate the false positive rate?<\/h3>\n<p>Peeking through observed power, a.k.a. \u201cfishing for high observer power\u201d, is bound to lead to a higher false positive rate than what\u2019s nominally targeted and reported. As shown in the simulations in <a href=\"https:\/\/blog.analytics-toolkit.com\/2024\/observed-power-in-online-a-b-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u201cUsing observed power in online A\/B tests\u201d<\/a>, continuing an experiment after seeing a non-significant outcome coupled (inevitably) with low power until a significant outcome is observed or for as long as practically possible leads to around <strong>2.4x inflation of the type I error rate<\/strong>.<\/p>\n<p>Changing some of the necessary assumptions may lower the actual type I error rate to \u201cjust\u201d being two-fold higher than the nominal one. The actual rate can be threefold the nominal under less favorable assumptions. In all cases the increase is far outside what is typically tolerable.<\/p>\n<h2 class=\"wp-block-heading\" id=\"opstatsig\">Observed power with a statistically significant outcome?<\/h2>\n<p>Given that statistical power relates to type II errors and the false negative rate, it might seem contradictory to even look at statistical power after a test has produced a statistically significant outcome. With a significant result in hand there remains only one error possible: that of a false positive and not of a false negative.<\/p>\n<p>Further, power is a pre-test probability, just like the odds of any particular person winning a lottery draw are. When someone wins the lottery it is illogical to point to their low odds of winning as a reason to reject giving them the prize. These odds are irrelevant once the outcome has been realized. It is similarly illogical to point to a low probability of observing a statistically significant outcome as an argument against the outcome being valid after it has already been observed. <em>This is regardless if the power analysis  focuses on observed power or not.<\/em><\/p>\n<p>As a practical matter, requiring an outcome to be both statistically significant and to have observed power higher than the target level results in an actual error control which is multiple times more stringent than the nominally targeted level.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"507\" height=\"510\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Actual-Significance-Threshold.png\" alt=\"Actual vs nominal threshold if requiring both high observed power and significance\" class=\"wp-image-1836\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Actual-Significance-Threshold.png 507w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Actual-Significance-Threshold-348x350.png 348w\" sizes=\"auto, (max-width: 507px) 100vw, 507px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 4<\/strong>: Actual vs nominal threshold if requiring both high observed power and significance<\/figcaption><\/figure>\n<\/div>\n<p><strong><em>Figure 4<\/em> demonstrates visually that the high observed power requirement completely overrides the significance threshold, making it redundant. <\/strong>That is because p-values below 0.006 would be disregarded by the requirement for low post hoc power, whereas the significance threshold is at 0.05.<\/p>\n<p>Knowing the above it may seem unfathomable that someone may insist on the utility of observed power even when presented with a statistically significant outcome. The already discussed \u201csearch for true power\u201d and viewing the chosen power level as a threshold can explain a lot of the motivation behind the misguided use of observed power in ensuring the trustworthiness of even statistically significant tests.<\/p>\n<p>However, there are two more to consider:<\/p>\n<h3 class=\"wp-block-heading\" id=\"mdethreshold\">MDE as a threshold<\/h3>\n<p>Another angle through which misunderstandings about statistical power make it into A\/B testing practice is when the question of a test being underpowered is framed in terms of the minimum detectable effect (MDE). Namely, it is seen as an issue if the observed effect size is smaller than the MDE, especially if the outcome is statistically significant.<\/p>\n<p>The numerous misunderstandings regarding the relationship (or lack thereof), as well as the harms which follow are explored in <a href=\"https:\/\/blog.analytics-toolkit.com\/2024\/what-if-the-observed-effect-is-smaller-than-the-mde\/\" target=\"_blank\" rel=\"noreferrer noopener\">What if the observed effect is smaller than the MDE?<\/a>. In short, the outcome of requiring significant results to also have an observed effect greater than the MDE are exactly the same as when requiring high observed power.<\/p>\n<h3 class=\"wp-block-heading\" id=\"effectsizeconfusion\">Viewing the observed effect size as the one a test\u2019s conclusion is about<\/h3>\n<p>From a slightly peculiar angle, observed power may be invoked with significant outcomes due to a warranted skepticism regarding an observed effect that seems \u201ctoo good to be true\u201d especially if it is coupled with a low p-value. The misunderstanding here is a common one related to p-values. Through its lens, <strong>p-values are viewed as giving the probability that the observed effect is due to chance alone<\/strong>. p-values, however, reflect the probability of obtaining an effect as large as the observed, or larger, assuming a data-generating mechanism in which the null hypothesis is true. Given the role of the null hypothesis in hypothesis testing, it is obvious that the p-value attaches to the test procedure and not to any given hypothesis, be it the null or alternative. P-values do not speak directly of the probability of the observed effect being genuine or not.<\/p>\n<p>If influenced by the above misconception, some might be skeptical upon seeing an effect that seems too good to be true accompanied by a p-value which seems too reassuring. Given the incorrect interpretation of what the statistics mean, such skepticism is warranted, but observed power is not the correct way to address it.<\/p>\n<p>A post hoc power calculation makes no use of the test data and is therefore unable to add anything to the observed p-value. Instead, the uncertainty surrounding a point estimate can be conveyed through confidence intervals at different levels, or in terms of severity (as in <a href=\"https:\/\/errorstatistics.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">D. Mayo\u2019s works<\/a>) for any effect of interest.<\/p>\n<h2 class=\"wp-block-heading\" id=\"propterhocpower\">All power is propter hoc power<\/h2>\n<p>In some sense, it can be argued that it is a misnomer to speak of \u201cObserved power\u201d, \u201cPost hoc Power\u201d, or \u201cRetrospective power\u201d. No formula or code for computing statistical power and has a reference to anything observed in the experiment. The reason is that statistical power calculations are inherently pre-data probabilities and so the statistical power function \/ power curve is computed using pre-data information.<\/p>\n<p>Examining any particular value of that function after an experiment is the same as if done beforehand since the entire power curve is constructed without any reference to observed test data.<\/p>\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p><strong>Statistical power is inherently a pre-data probability. It does not use any of the data from a controlled experiment. As result, the power level for any hypothetical true effect can be computed before any data has been gathered.<\/strong><\/p>\n<\/blockquote>\n<p>Maybe the term \u201cpost hoc power analysis\u201d has some merit in that it is a power analysis performed after the fact on the timeline, but given the other possible meanings just calling it a \u201cpower analysis\u201d would avoid many of the problems and misunderstandings associated with attaching a \u201cpost hoc\u201d or \u201cretrospective\u201d qualifier to \u201cpower\u201d.<\/p>\n<h2 class=\"wp-block-heading\" id=\"takeaways\">Takeaways<\/h2>\n<p>All possible justifications for the use of observed power rest on unsound arguments and stem from misunderstanding of one or more statistical concepts. To recap:<\/p>\n<ul class=\"wp-block-list\">\n<li>High true power is not necessary for a test to be well-powered, and even if it were, observed power is not a good estimate of true power<\/li>\n<li>Neither the target power level nor the target MDE are thresholds to be met by any observed statistic (\u201cobserved\u201d power or observed effect size respectively being the common choices)<\/li>\n<li>Non-significant outcomes always appear underpowered if the observed power level is compared to the target power level<\/li>\n<li>Requiring high post hoc power for significant outcomes is equivalent to drastically increasing the significance threshold of the test, and so is requiring the observed effect size to be greater than or equal to the target MDE (as the two are equivalent)<\/li>\n<li>Statistical power is a pre-test probability, regardless of when it is computed, and it does not use the test data<\/li>\n<\/ul>\n<p>At best, observed power is unnecessary given a p-value has been calculated as it has direct functional relationship to it. Enter a p-value and a threshold and the post-hoc power calculator will output the observed power.<\/p>\n<p>At worst, using observed power may lead to all kinds of issues, including:<\/p>\n<ul class=\"wp-block-list\">\n<li>mislabeling well-powered tests as either underpowered or overpowered ones<\/li>\n<li>identifying trustworthy results as untrustworthy<\/li>\n<li>peeking through observed power, resulting in a significantly inflated type I error rate<\/li>\n<li>greatly raising the bar for significance in a non-obvious manner (when required with significant outcomes)<\/li>\n<li>unjustifiably extending the duration and increasing the sample sizes of A\/B tests by orders of magnitude<\/li>\n<\/ul>\n<p>Given the above, Deborah Mayo\u2019s term \u201cshpower\u201d is spot on when describing observed power, defined as the level of statistical power to detect a true effect equal to the observed one. One should be able to guess what two words \u201cshpower\u201d is derived from, even if English is not your native language.<\/p>\n<h3 class=\"wp-block-heading\">References<\/h3>\n<p><span class=\"reference\">1<\/span> Hoenig, J. M., Heisey, D. M. (2001) \u201cThe abuse of power: The pervasive fallacy of power calculations for data analysis.\u201d\u00a0<em>The American Statistician, 55<\/em>, 19-24.<br \/><span class=\"reference\">2<\/span> Lakens, D. (2014) \u201cObserved power, and what to do if your editor asks for post-hoc power analyses\u201d, online at https:\/\/daniellakens.blogspot.com\/2014\/12\/observed-power-and-what-to-do-if-your.html<br \/><span class=\"reference\">3<\/span> Zhang et al. (2019) \u201cPost hoc power analysis: is it an informative and meaningful analysis?\u201d <em>General Psychiatry<\/em> 32(4):e100069<br \/><span class=\"reference\">4<\/span> Christogiannis et al. (2022) \u201cThe self-fulfilling prophecy of post-hoc power calculations\u201d <em>American Journal of Orthodontics and Dentofacial Orthopedics<\/em> 161:315-7<br \/><span class=\"reference\">5<\/span> Heckman, M.G., Davis III, J.M., Crowson, C.S. (2022) \u201cPost Hoc Power Calculations: An Inappropriate Method for Interpreting the Findings of a Research Study\u201d <em>The Journal of Rheumatology<\/em> 49(8), 867-870<br \/><span class=\"reference\">6<\/span> Yuan, K.-H., Maxwell, S. (2005) \u201cOn the Post Hoc Power in Testing Mean Differences\u201d <em>Journal of Educational and Behavioral Statistics<\/em>, <em>30<\/em>(2), 141\u2013167.<\/p>\n<h4 id=\"authorStart\" class=\"tc-darkerGrey\">About the author<\/h4>\n<div id=\"authorBlock\" class=\"bgLighterGrey\">\n\t\t<img decoding=\"async\" src=\"https:\/\/www.analytics-toolkit.com\/img\/Georgi-Georgiev-150x150.png\" width=\"150\" height=\"150\" loading=\"lazy\" \/><\/p>\n<p class=\"authorName\">Georgi Georgiev <span class=\"float-end fs-5\"><a href=\"https:\/\/www.linkedin.com\/in\/geoprofi\/\" rel=\"nofollow\" target=\"_blank\" title=\"Follow me on LinkedIn\"><i class=\"fa-brands fa-linkedin\"><\/i><\/a><a href=\"https:\/\/twitter.com\/georgizgeorgiev\" rel=\"nofollow\" target=\"_blank\" class=\"mx-2\" title=\"Follow me on Twitter\"><i class=\"fa-brands fa-twitter\"><\/i><\/a><a href=\"http:\/\/www.facebook.com\/AnalyticsToolkit\" rel=\"nofollow\" target=\"_blank\" title=\"Follow me on Facebook\"><i class=\"fa-brands fa-facebook-square\"><\/i><\/a><\/span><\/p>\n<p class=\"fs-7\">Managing owner of Web Focus and creator of Analytics-toolkit.com, Georgi has over twenty years of experience in online marketing, web analytics, statistics, and design of business experiments for hundreds of websites.<\/p>\n<p class=\"fs-7\">He is the author of the book &#8220;Statistical Methods in Online A\/B Testing&#8221;, of white papers on statistical analysis of A\/B tests, and has been a speaker at conferences, seminars, and courses. Georgi has been distinguished as a winner in the Data &amp; Analytics category of the 2024 Experimentation Thought Leadership Awards.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><script async src=\"\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><br \/>\n<br \/><\/p>\n","protected":false},"excerpt":{"rendered":"<p>\u201cObserved power\u201d, \u201cPost hoc Power\u201d, and \u201cRetrospective power\u201d all refer to the statistical power of a statistical significance test to detect a true effect equal to the observed effect. In a broader sense these terms may also describe any power analysis performed after an experiment has completed. Importantly, it is the first, narrower sense that [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":22567,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[10921,2059,15178,15177,2649,1783],"dealstore":[],"offerexpiration":[],"class_list":["post-22566","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-comprehensive","tag-guide","tag-hoc","tag-observed","tag-post","tag-power"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>A Comprehensive Guide to Observed Power (Post Hoc Power) - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=22566\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"A Comprehensive Guide to Observed Power (Post Hoc Power) - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"\u201cObserved power\u201d, \u201cPost hoc Power\u201d, and \u201cRetrospective power\u201d all refer to the statistical power of a statistical significance test to detect a true effect equal to the observed effect. In a broader sense these terms may also describe any power analysis performed after an experiment has completed. Importantly, it is the first, narrower sense that [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=22566\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-01-11T09:46:15+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"675\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"22 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=22566#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=22566\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"A Comprehensive Guide to Observed Power (Post Hoc Power)\",\"datePublished\":\"2025-01-11T09:46:15+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=22566\"},\"wordCount\":4457,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=22566#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png\",\"keywords\":[\"Comprehensive\",\"Guide\",\"Hoc\",\"Observed\",\"Post\",\"Power\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=22566#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=22566\",\"url\":\"https:\/\/fivemor.com\/?p=22566\",\"name\":\"A Comprehensive Guide to Observed Power (Post Hoc Power) - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=22566#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=22566#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png\",\"datePublished\":\"2025-01-11T09:46:15+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=22566#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=22566\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=22566#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png\",\"width\":1200,\"height\":675},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=22566#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"A Comprehensive Guide to Observed Power (Post Hoc Power)\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"A Comprehensive Guide to Observed Power (Post Hoc Power) - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=22566","og_locale":"en_US","og_type":"article","og_title":"A Comprehensive Guide to Observed Power (Post Hoc Power) - Som2ny Network","og_description":"\u201cObserved power\u201d, \u201cPost hoc Power\u201d, and \u201cRetrospective power\u201d all refer to the statistical power of a statistical significance test to detect a true effect equal to the observed effect. In a broader sense these terms may also describe any power analysis performed after an experiment has completed. Importantly, it is the first, narrower sense that [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=22566","og_site_name":"Som2ny Network","article_published_time":"2025-01-11T09:46:15+00:00","og_image":[{"width":1200,"height":675,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png","type":"image\/png"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"22 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=22566#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=22566"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"A Comprehensive Guide to Observed Power (Post Hoc Power)","datePublished":"2025-01-11T09:46:15+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=22566"},"wordCount":4457,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=22566#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png","keywords":["Comprehensive","Guide","Hoc","Observed","Post","Power"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=22566#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=22566","url":"https:\/\/fivemor.com\/?p=22566","name":"A Comprehensive Guide to Observed Power (Post Hoc Power) - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=22566#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=22566#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png","datePublished":"2025-01-11T09:46:15+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=22566#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=22566"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=22566#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-07-Comprehensive-Guide-Observed-Power.png","width":1200,"height":675},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=22566#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"A Comprehensive Guide to Observed Power (Post Hoc Power)"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/22566","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=22566"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/22566\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/22567"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=22566"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=22566"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=22566"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=22566"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=22566"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}