It has happened to you, hasn't it? You submit your article to a journal, and Reviewer 2 asks you to calculate the statistical power of your analysis... after you have already conducted it. It sounds reasonable, but there is a problem: post-hoc power calculation is mathematically circular. And it is not just me saying this: it is a growing consensus among statisticians and methodologists. Yet many reviewers continue to request it. So let us debunk this myth and see what to do instead.
What is statistical power and why does it matter?
Before explaining why post-hoc calculation is problematic, let us review the basics. Statistical power is the probability of detecting an effect when it truly exists. It is calculated before collecting data: hence the term a priori: and depends on three things: the effect size you expect to find, your sample size, and the significance level (alpha). Choosing the right statistical test is the first step for this calculation to be meaningful.
A study with 80% power means that, if the effect exists, you have an 80% probability of detecting it. It is a fundamental component of research design. So far, so good.
If you want to go from reading the curve to picking your n (the sample size that buys you the power you are after, given your test and effect size), the sample size calculator works it out instantly, no G*Power install required.
The problem with post-hoc calculation: a tautology in disguise
The trouble begins when a reviewer asks you to calculate power after you have obtained your results, using the effect size you observed in your sample. Why is this problematic? Because post-hoc power calculated with the observed effect is a direct function of the p-value.
Think of it this way: if your result was not significant (p > .05), the post-hoc power will inevitably be low. And if it was significant (p < .05), it will be high. Always. It is pure mathematics: it provides no new information. It is like saying "we did not find the effect because we did not have enough power, and we know we did not have enough power because we did not find the effect." A perfect tautology that tells you absolutely nothing useful.
Why do so many reviewers still request it?
Good question. The short answer: because it sounds intuitive. It seems logical to want to know whether your study "had enough power" to detect something. But the problem is conceptual: power is a property of the design, not of the results. Once you have the data, the relevant question is no longer "did I have power?" but "how precise is my estimate?"
Several landmark methodological articles have argued firmly against this practice. And an increasing number of journal guidelines advise against reporting post-hoc power. Yet the myth persists, especially in clinical and educational psychology.
What to do when the reviewer asks for post-hoc power
Here is the practical advice. You have two main options, and you can use both:
Option 1: Sensitivity analysis
Instead of calculating power for the effect you found, calculate the power you had for different theoretically relevant effect sizes. For example: "With our sample of 120 participants, we had 95% power to detect large effects (d = 0.8), 70% for medium effects (d = 0.5), and 25% for small effects (d = 0.2)."
This is infinitely more informative because it tells the reader which effects your study was sensitive to and which it was "blind" to. It is an honest acknowledgment of limitations that reviewers respect far more than a circular number.
Option 2: Confidence intervals for the effect size
This is, in my opinion, the superior alternative. Report the 95% confidence interval for your effect size (you can calculate the effect size with our tool). If your Cohen's d is 0.15 with a 95% CI of [-0.10, 0.40], you are communicating that your data are compatible with both the absence of an effect and a moderate effect. That says much more than any power calculation.
A narrow interval around zero suggests with considerable confidence that the effect is small or nonexistent. A wide interval spanning from zero to large values indicates that your study was simply not precise enough to draw firm conclusions.
How to respond to the reviewer with tact
I know that saying "no" to a reviewer can be daunting. But there are diplomatic ways to do so. You could write something like:
"We appreciate the reviewer's suggestion. However, post-hoc power calculation using the observed effect size has been advised against by several methodologists [cite references], given that it is a direct function of the p-value and does not provide independent information. Instead, we have included a sensitivity analysis and report confidence intervals for the effect size, which we consider more informative for evaluating the precision of our estimates."
With this response, you are not only declining the request politely: you are demonstrating that you know more about methodology than the reviewer themselves. And believe me, editors notice.
So this does not happen again: paste your next manuscript into the free AI paper reviewer before you submit it. A post-hoc power request is exactly the kind of objection it flags in advance, along with the missing effect sizes and confidence intervals that usually trigger it, each one anchored to a line of your own text.
So, when DOES it make sense to calculate power?
A priori power: before collecting data: remains absolutely essential. Every study should include a power analysis during the design phase to justify the sample size. Tools such as G*Power, the pwr package in R, or our sample size calculator make this very accessible.
What does not make sense is calculating it afterward with the results you already have. It is like looking at the final score of a match and calculating the "probability of winning" for the losing team. It adds nothing.
If you are dealing with reviewers who request analyses you know are problematic, do not hesitate to contact us. We help you prepare solid, well-grounded responses that protect the methodological integrity of your work: and that convince even the most stubborn Reviewer 2.