{"id":27625,"date":"2025-01-15T03:02:01","date_gmt":"2025-01-15T03:02:01","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/using-observed-power-in-online-a-b-tests\/"},"modified":"2025-01-15T03:02:01","modified_gmt":"2025-01-15T03:02:01","slug":"using-observed-power-in-online-a-b-tests","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=27625","title":{"rendered":"Using Observed Power in Online A\/B Tests"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>Observed power, often referred to as \u201cpost hoc power\u201d and \u201cretrospective power\u201d is the statistical power of a test to detect a true effect equal to the observed effect size. \u201cDetect\u201d in the context of a statistical hypothesis test means to result in a statistically significant outcome.<\/p>\n<p>Some calculators aimed at A\/B testing practitioners <strong>use observed power to guide users as to what the sample size of their test should be<\/strong>. When performing statistical significance calculations, these tools present something like \u201cRequired sample size\u201d or \u201cHow long to run the test for\u201d next to their results. Those numbers are apparently supposed to tell if the test needs to run for longer or not, or at the very least to indicate if a test is underpowered or overpowered.<\/p>\n<p>Despite several attempts over the years to nudge the authors of said tools to remove this harmful \u201cfunctionality\u201d from their calculators, it continues to exist. Given the relative popularity of these tools among less experienced practitioners, the use of observed power may be doing noticeable harm to the online experimentation field, hence this article.<\/p>\n<p><strong>The following topics are covered below:<\/strong><\/p>\n<ol class=\"wp-block-list\">\n<li><a href=\"#observedpower\">Using observed power in practice<\/a><\/li>\n<li><a href=\"#sigpower\">Requiring statistical significance AND high observed power<\/a><\/li>\n<li><a href=\"#fishing\">Fishing for high observed power<\/a><br \/>3.1 <a href=\"#thesim\">The simulation<\/a><br \/>3.2 <a href=\"#simresults\">Results from peeking through observed power<\/a><\/li>\n<li><a href=\"#takeaways\">Takeaways<\/a><\/li>\n<\/ol>\n<h2 class=\"wp-block-heading\" id=\"observedpower\">Using observed power in practice<\/h2>\n<p>There are three possible scenarios for using observed power in any one of these calculators:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>The outcome of an A\/B test is calculated as being statistically significant<\/strong>, and the observed power is higher than whatever threshold was chosen (80%, 90%, etc.). In this case the calculator will almost always say that the test was <strong>overpowered<\/strong>. In other words, one has \u201cwaited for too long\u201d and has exposed too many users to the test. For example, the sample size may have been 10,000 users per variant, but the \u201cRequired sample size\u201d field may say only 8,000 were needed.<\/li>\n<li><strong>The outcome may be statistically significant, but the observed power is lower<\/strong> than the chosen threshold, which will (wrongly) suggest the test was <strong>underpowered<\/strong>. One might interpret this to mean the test is untrustworthy and cause them to ignore a perfectly valid statistically significant outcome. A different interpretation may be to treat the low observed power as a license to continue the test at least until the \u201cRequired sample size\u201d is reached and perhaps even for longer if it happens to need even more data after that first extension. Both of these routes lead to very different outcomes.<\/li>\n<li>Finally, <strong>the outcome may not be statistically significant<\/strong>, in which case the calculator will <em>always <\/em>suggest the A\/B test is <strong>significantly underpowered<\/strong>. For example, a user may have tested with 10,000 users per variant, but the \u201cRequired sample size\u201d output will inform them they need multiple times larger sample size, with numbers 5x or 10x not being uncommon. In the same example that would mean the tools suggest exposing 40,000 or 90,000 more users per variant than had initially been planned (assuming there was a plan) and taking five to ten times longer.<\/li>\n<\/ol>\n<p><strong>Notably, one is almost never going to see that their test is well-powered based on observed power calculations.<\/strong> The math is bound to almost always label an A\/B test as either underpowered (more often), or overpowered (typically less often). My <a href=\"https:\/\/blog.analytics-toolkit.com\/2024\/comprehensive-guide-to-observed-power-post-hoc-power\/\" target=\"_blank\" rel=\"noreferrer noopener\">comprehensive guide on observed power<\/a> would be a great read for anyone keen on understanding the matter in detail, and it explains why the above scenarios encapsulate all that can happen. It also includes a calculator which lets you thoroughly explore observed power.<\/p>\n<p>In the first scenario many experimenters would typically just walk away with a statistically significant finding. What happens in each of the other two scenarios is far more interesting. Let us deal with them one by one.<\/p>\n<h2 class=\"wp-block-heading\" id=\"sigpower\">Requiring statistical significance AND high observed power<\/h2>\n<p>In scenario two some practitioners will just ignore the warning about their test \u201crequiring more users\u201d or \u201cmore time\u201d since the result is statistically significant. However, some might be swayed into accepting a test\u2019s significant outcome only if the outcome is both statistically significant and \u201cwell-powered\u201d or \u201coverpowered\u201d (based on observed power). This kind of behavior leads to multiple times more stringent type I error control than the nominal threshold sets.<\/p>\n<p>The exact type I error which would be achieved can be estimated easily by plotting p-values and observed power on a graph, and checking what p-value corresponds to the chosen power level. This can also be done with the <a href=\"https:\/\/blog.analytics-toolkit.com\/2024\/comprehensive-guide-to-observed-power-post-hoc-power\/#posthocpowercalculator\" target=\"_blank\" rel=\"noreferrer noopener\">post hoc power calculator<\/a> I\u2019ve created for educational purposes.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"507\" height=\"510\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Actual-Significance-Threshold-1.png\" alt=\"Actual vs nominal threshold if requiring both high observed power and significance\" class=\"wp-image-1837\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Actual-Significance-Threshold-1.png 507w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/10\/2024-10-04-Actual-Significance-Threshold-1-348x350.png 348w\" sizes=\"auto, (max-width: 507px) 100vw, 507px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 1<\/strong>: Actual vs nominal threshold if requiring both high observed power and significance<\/figcaption><\/figure>\n<\/div>\n<p>For example, using \u201ctextbook\u201d values of 0.05 and 0.80 as the cut-off for significance and for observed power (80% power of the design), one would in fact achieve a type I error rate of about 0.6%, or <strong>more than eight times smaller<\/strong> than the target 5%. Note that the same is true if one requires the observed effect to be greater than the MDE.<\/p>\n<p>If using 0.05 significance threshold and 90% power, the actual type I error rate would be 0.17%, or about <strong>thirty times smaller<\/strong> than the target of 5%. If using a stricter significance threshold of 0.01 with 80% power, the actual type I rate is 0.076%, or about thirteen times smaller, and so on. The stricter each of the thresholds is, the smaller the actual type I error attained through this approach.<\/p>\n<p>Additionally I have confirmed all of the above independently through simulations.<\/p>\n<p>Such a practice has a crippling effect on the achieved statistical power, which instead of being 80% or 90%, is reduced significantly. If it is not obvious why that is a <em>big <\/em>issue, consider <a href=\"https:\/\/blog.analytics-toolkit.com\/2017\/importance-statistical-power-online-ab-tests\/\" target=\"_blank\" rel=\"noreferrer noopener\">The importance of statistical power in online A\/B testing<\/a>.<\/p>\n<p>A different problem faces those who might choose to continue the initially significant but seemingly underpowered test until it is eventually both statistically significant and well-powered. Interestingly, my simulations show that such A\/B tests will not suffer from an inflation of their type I error. However, they should expect to commit to an <strong>eightfold increase in sample size<\/strong>, on average, at least in the cases where there is no true effect or it is of minimal size.<\/p>\n<h2 class=\"wp-block-heading\" id=\"fishing\">Fishing for high observed power<\/h2>\n<p>In scenario three, some may interpret the fact that their sample size is many times lower than the \u201crequired sample size\u201d as a license to continue testing until the recommended sample size is reached. After all, if a statistically insignificant test seems \u201cunderpowered\u201d, it sounds like there is a true effect waiting to be found, if one just extends their sample size to what\u2019s needed to make it \u201cpowerful enough\u201d. Why then not to continue testing for however long it takes to reach the \u201crequired sample size\u201d suggested by the calculator?<\/p>\n<p>However, examining A\/B test data with an intent to stop based on the outcome constitutes <a href=\"https:\/\/www.analytics-toolkit.com\/glossary\/peeking\/\" target=\"_blank\" rel=\"noreferrer noopener\">peeking<\/a> and is a violation of crucial assumptions behind the tests used in these calculators. Classic peeking is sometimes called \u201cfishing for significance\u201d so by analogy we may call this approach \u201cfishing for observed power\u201d. Just like its cousin, peeking through observed power results in an inflated false positive rate (type I error rate). In a typical scenario of a <a href=\"https:\/\/blog.analytics-toolkit.com\/2017\/one-tailed-two-tailed-tests-significance-ab-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">one-sided test<\/a>, if the true effect is zero or negative, statistically significant outcomes would be \u201cfound\u201d multiple times more often than the nominal p-value or confidence level would suggest.<\/p>\n<p>Unlike classic peeking, how much type I error inflation one is bound to get by fishing for high observed power is unknown, at least as far as I am aware. Given this, a simulation is a straightforward way to emulate the behavior of peeking through observed power and to estimate what it leads to.<\/p>\n<h3 class=\"wp-block-heading\" id=\"thesim\">The simulation<\/h3>\n<p>The simulation below was performed under the following assumptions:<\/p>\n<ul class=\"wp-block-list\">\n<li>one starts with some reasonable sample size before their first calculation of observed power<\/li>\n<li>one stops a test immediately if it is nominally statistically significant at any evaluation (regardless of post hoc power)<\/li>\n<li>if the observed effect is zero or negative, the test stops (it will be non-significant)<\/li>\n<li>if the observed effect is positive, the test is extended with the recommended sample size calculated using the observed effect size as an MDE, but the extension is truncated to no more than 10 times the initial sample size. Changing this parameter or removing it altogether does not meaningfully alter the outcomes, but makes the simulations more tractable and somewhat closer to reality.<\/li>\n<li>if after one or more extensions the sample size reaches 49 times the initial sample size, the test stops regardless of outcome. This is to keep the simulation somewhat closer to real life where few if any would be willing to continue past such large multiples of the original sample size.<\/li>\n<\/ul>\n<p>The above describes scenario 3 with some limits added to keep things realistic. What happens without some of those limits is shared in Appendix I as a sanity check.<\/p>\n<p>100,000 simulation runs were performed with the following parameters:<\/p>\n<ul class=\"wp-block-list\">\n<li>a true effect of exactly zero<\/li>\n<li>an initial sample size of 50,000 users per variant<\/li>\n<li>a significance threshold of 0.05 (textbook value, equivalent to requiring at least 95% confidence), resulting in a target type I error rate of 5%.<\/li>\n<li>a power threshold of 0.8 (80%) to which observed power was compared and which was used to perform the \u201crequired sample size\u201d calculation<\/li>\n<\/ul>\n<p>The summary results are presented below.<\/p>\n<h3 class=\"wp-block-heading\" id=\"simresults\">What happens when peeking through observed power<\/h3>\n<p>Since observed power and p-values have a direct functional relationship, it is unsurprising that fishing for high observed power is in a way fishing for statistical significance. Expectedly, it results in an inflation of the <a href=\"https:\/\/www.analytics-toolkit.com\/glossary\/type-i-error\/\" target=\"_blank\" rel=\"noreferrer noopener\">type  I error<\/a>.<\/p>\n<p>11,685 out of 100,000 simulations resulted in a statistically significant outcome, meaning the type I error rate was 11.685% instead of the target 5%, representing a <strong>nearly 2.4x increase of the false positive rate<\/strong>. Additionally, the average sample size these tests ended up with was close to 540,000 users per variant which is more than 10x the initial target of 50,000.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"710\" height=\"529\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-P-value-Distribution-When-Peeking-With-Observed-Power.png\" alt=\"P value Distribution When Peeking With Observed Power\" class=\"wp-image-1786\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-P-value-Distribution-When-Peeking-With-Observed-Power.png 710w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-P-value-Distribution-When-Peeking-With-Observed-Power-470x350.png 470w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 2<\/strong>: p-value distribution from peeking with observed power<\/figcaption><\/figure>\n<\/div>\n<p>The distribution of p-values is decidedly non-uniform, and highly skewed towards low values (a proper p-value should have a uniform distribution under a true null). The fact that there are any values between 0.05 and 0.5 at all is mostly due to the maximum sample size limit of 49x the initial sample size.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"710\" height=\"532\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Observed-Power-When-Peeking.png\" alt=\"Observed power under a true null when peeking with observed power\" class=\"wp-image-1787\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Observed-Power-When-Peeking.png 710w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Observed-Power-When-Peeking-467x350.png 467w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 3<\/strong>: Observed power under a true null when peeking with observed power<\/figcaption><\/figure>\n<\/div>\n<p>The distribution of observed power at the time when a test was stopped is also somewhat interesting. The skewed nature of the distribution is entirely expected even without peeking, but the bump just above 50% is not. Peeking is what formed the \u201cvalley\u201d to the left of 50%. Since values above 50% are those of tests with statistically significant outcomes, it was values that were initially just below 50% which got extended and since the true effect is zero, these tests ended up with a smaller observed power as the observed effect drifts towards the true effects with increasing sample size.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"710\" height=\"530\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Sample-Size-Distribution-of-Peeking-With-Power.png\" alt=\"Sample size distribution of peeking with observed power\" class=\"wp-image-1789\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Sample-Size-Distribution-of-Peeking-With-Power.png 710w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Sample-Size-Distribution-of-Peeking-With-Power-469x350.png 469w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 4<\/strong>: Sample size distribution of the simulation runs<\/figcaption><\/figure>\n<\/div>\n<p>The peculiar nature of the sample size distribution is due to the caps on how much the \u201crequired sample size\u201d is allowed to increase the sample size at each step. The higher proportion of achieved sample sizes above 2mln users per variant compared to the earlier segment is due to the maximum cap on the sample size.<\/p>\n<h2 class=\"wp-block-heading\" id=\"takeaways\">Takeaways<\/h2>\n<p><strong>Using observed power in any way is harmful to the error control and trustworthiness of any experiment.<\/strong> At the very least, observed power will almost surely label a test as either underpowered or overpowered, without any reference to it\u2019s actual qualities. However, different scenarios of using observed power result in additional problems of severe consequences.<\/p>\n<p>Requiring observed power in conjunction with a statistically significant outcome in order to reject a null hypothesis results in a nominal false positive rate multiple times higher than the one actually achieved. It raises the bar of evidence required to declare a \u201cwinning\u201d experiment to levels far above the initially targeted. If one ignores significant outcomes from \u201cunderpowered\u201d tests and continues testing until the outcome is both significant and the observed power is above the target power level, then it is just a massive (more than eightfold) waste of time and user exposure.<\/p>\n<p>When observed power is used with non-significant outcomes to guide how long to run a test for and with what sample size, the effect is similar to classic peeking. This practice leads to multiple times higher actual false positive rate versus the nominal one (an inflated type I error rate) with realistic simulations showing this increase to be at least 2.4x. Altering the assumptions more or less generously regarding what a practitioner does results in actual error rates <strong>between two times and three times the target error rate<\/strong>. A test with such a discrepancy between its target (or reported) error rate and its actual one cannot be seen as valid or trustworthy.<\/p>\n<p>If you have found the above interesting, you might consider checking out <a href=\"https:\/\/www.analytics-toolkit.com\/ab-testing-hub\/\" target=\"_blank\" rel=\"noreferrer noopener\">the statistically-rigorous A\/B testing calculators at Analytics Toolkit<\/a>.<\/p>\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<h3 class=\"wp-block-heading\">Appendix I. Simulation results with different caps<\/h3>\n<p>1) This first set of results are without the hard limit on the total sample size being no more than 49x the initial sample size after all possible sample size increases.<\/p>\n<p>14,169 out of 100,000 simulations resulted in a statistically significant outcome, meaning the type I error rate was 14.169% instead of the target 5%, representing a more than 2.8x increase of the false positive rate. The average sample size these tests ended up with was close to 1,120,000 users per variant which is more than 22x the initial target of 50,000.<\/p>\n<p>The resulting distributions are:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"710\" height=\"530\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-P-value-Distribution-When-Peeking-With-Observed-Power-Unbounded.png\" alt=\"P value Distribution When Peeking With Observed Power (Unbounded)\" class=\"wp-image-1790\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-P-value-Distribution-When-Peeking-With-Observed-Power-Unbounded.png 710w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-P-value-Distribution-When-Peeking-With-Observed-Power-Unbounded-469x350.png 469w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" \/><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"710\" height=\"530\" src=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Observed-Power-When-Peeking-Unbounded.png\" alt=\"Distribution of Observed Power When Peeking through Observed Power (Unbounded)\" class=\"wp-image-1791\" srcset=\"https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Observed-Power-When-Peeking-Unbounded.png 710w, https:\/\/blog.analytics-toolkit.com\/wp-content\/uploads\/2024\/09\/2024-09-30-Observed-Power-When-Peeking-Unbounded-469x350.png 469w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" \/><\/figure>\n<\/div>\n<p>The sample size distribution is too heavy tailed to feature here in a meaningful manner.<\/p>\n<p>2) Removing both the hard 49x limit and the limit on each increment to be no more than 10x the initial sample size results in a false positive rate of around 12.5% (with a target type I error of 5%).<\/p>\n<p>3) Lowering the limit on each increment from 10x the initial sample size to 2x the initial sample size while retaining the hard 49x limit results in a false positive rate of around 14.5% (again with the nominal level being 5%). The reason is that it results in more peeks than the 10x limit.<\/p>\n<h4 id=\"authorStart\" class=\"tc-darkerGrey\">About the author<\/h4>\n<div id=\"authorBlock\" class=\"bgLighterGrey\">\n\t\t<img decoding=\"async\" src=\"https:\/\/www.analytics-toolkit.com\/img\/Georgi-Georgiev-150x150.png\" width=\"150\" height=\"150\" loading=\"lazy\" \/><\/p>\n<p class=\"authorName\">Georgi Georgiev <span class=\"float-end fs-5\"><a href=\"https:\/\/www.linkedin.com\/in\/geoprofi\/\" rel=\"nofollow\" target=\"_blank\" title=\"Follow me on LinkedIn\"><i class=\"fa-brands fa-linkedin\"><\/i><\/a><a href=\"https:\/\/twitter.com\/georgizgeorgiev\" rel=\"nofollow\" target=\"_blank\" class=\"mx-2\" title=\"Follow me on Twitter\"><i class=\"fa-brands fa-twitter\"><\/i><\/a><a href=\"http:\/\/www.facebook.com\/AnalyticsToolkit\" rel=\"nofollow\" target=\"_blank\" title=\"Follow me on Facebook\"><i class=\"fa-brands fa-facebook-square\"><\/i><\/a><\/span><\/p>\n<p class=\"fs-7\">Managing owner of Web Focus and creator of Analytics-toolkit.com, Georgi has over twenty years of experience in online marketing, web analytics, statistics, and design of business experiments for hundreds of websites.<\/p>\n<p class=\"fs-7\">He is the author of the book &#8220;Statistical Methods in Online A\/B Testing&#8221;, of white papers on statistical analysis of A\/B tests, and has been a speaker at conferences, seminars, and courses. Georgi has been distinguished as a winner in the Data &amp; Analytics category of the 2024 Experimentation Thought Leadership Awards.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><script async src=\"\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><br \/>\n<br \/><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Observed power, often referred to as \u201cpost hoc power\u201d and \u201cretrospective power\u201d is the statistical power of a test to detect a true effect equal to the observed effect size. \u201cDetect\u201d in the context of a statistical hypothesis test means to result in a statistically significant outcome. Some calculators aimed at A\/B testing practitioners use [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":27626,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[15177,2485,1783,11326],"dealstore":[],"offerexpiration":[],"class_list":["post-27625","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-observed","tag-online","tag-power","tag-tests"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Using Observed Power in Online A\/B Tests - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=27625\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Using Observed Power in Online A\/B Tests - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Observed power, often referred to as \u201cpost hoc power\u201d and \u201cretrospective power\u201d is the statistical power of a test to detect a true effect equal to the observed effect size. \u201cDetect\u201d in the context of a statistical hypothesis test means to result in a statistically significant outcome. Some calculators aimed at A\/B testing practitioners use [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=27625\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-01-15T03:02:01+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"675\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=27625#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=27625\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Using Observed Power in Online A\/B Tests\",\"datePublished\":\"2025-01-15T03:02:01+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=27625\"},\"wordCount\":2458,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=27625#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png\",\"keywords\":[\"Observed\",\"ONLINE\",\"Power\",\"tests\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=27625#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=27625\",\"url\":\"https:\/\/fivemor.com\/?p=27625\",\"name\":\"Using Observed Power in Online A\/B Tests - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=27625#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=27625#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png\",\"datePublished\":\"2025-01-15T03:02:01+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=27625#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=27625\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=27625#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png\",\"width\":1200,\"height\":675},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=27625#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Using Observed Power in Online A\/B Tests\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Using Observed Power in Online A\/B Tests - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=27625","og_locale":"en_US","og_type":"article","og_title":"Using Observed Power in Online A\/B Tests - Som2ny Network","og_description":"Observed power, often referred to as \u201cpost hoc power\u201d and \u201cretrospective power\u201d is the statistical power of a test to detect a true effect equal to the observed effect size. \u201cDetect\u201d in the context of a statistical hypothesis test means to result in a statistically significant outcome. Some calculators aimed at A\/B testing practitioners use [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=27625","og_site_name":"Som2ny Network","article_published_time":"2025-01-15T03:02:01+00:00","og_image":[{"width":1200,"height":675,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png","type":"image\/png"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=27625#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=27625"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Using Observed Power in Online A\/B Tests","datePublished":"2025-01-15T03:02:01+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=27625"},"wordCount":2458,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=27625#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png","keywords":["Observed","ONLINE","Power","tests"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=27625#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=27625","url":"https:\/\/fivemor.com\/?p=27625","name":"Using Observed Power in Online A\/B Tests - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=27625#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=27625#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png","datePublished":"2025-01-15T03:02:01+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=27625#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=27625"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=27625#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/01\/2024-10-11-Observed-Power-in-AB-Testing.png","width":1200,"height":675},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=27625#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Using Observed Power in Online A\/B Tests"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/27625","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=27625"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/27625\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/27626"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=27625"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=27625"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=27625"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=27625"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=27625"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}