The data are collected. The participants are gone. And a reviewer has written that the study is underpowered, that no a priori power analysis is reported, and that the null findings are therefore uninterpretable. It reads like a demand for a time machine, and the standard reply, computing observed power from the effect you found, is the one move that will make things worse.
The good news is that “underpowered” is not one objection, it is three, and only one of them requires new participants. The reviewer may mean that you never justified your sample size, which is a reporting problem you can fix in an afternoon. They may mean that your N cannot detect effects of the size the literature reports, which is a real limitation that you can quantify precisely and honestly. Or they may mean that you interpreted a non-significant result as evidence of no effect, which is a claim problem, and the fix is to change the claim or to test absence properly.
This page separates the three, explains why post hoc power is circular, and walks through what you can legitimately run and report after data collection has ended: sensitivity power analysis, interval estimates, equivalence testing, Bayes factors, and the design-analysis reasoning about exaggeration that reviewers in 2026 increasingly expect to see.
Before you write anything, work out which complaint you received, because the three have almost nothing in common except the vocabulary.
This is a reporting gap and it is the most common of the three. The reviewer is not disputing your N, they are noting that you never explained where it came from. It is now a required element in the reporting guidelines most journals in psychology and health sciences endorse, and it is trivially fixable.
You can justify a sample size in more than one legitimate way, and only one of them is an a priori power analysis. Resource constraint is acceptable if stated honestly: the sample is the full cohort of patients treated in a defined window, or the entire staff of the participating units, or what the funded recruitment period allowed. A precision rationale is acceptable: the N gives a confidence interval of a stated width around the parameter of interest. What is not acceptable is silence, or the sentence “a sample of 120 was considered sufficient”, which asserts the thing under dispute.
This is a design complaint about the match between N and model complexity, and it is usually correct. It bites hardest in three places: interaction and moderation tests, latent variable models, and multilevel models with too few clusters.
The arithmetic is unforgiving. Detecting a between-groups d of 0.50 with 80% power in a two-group design needs roughly 64 per group; a d of 0.35, which is closer to the average published effect once you correct for publication bias, needs about 130 per group. Detecting a correlation of r = .30 at 80% power needs about 85 participants; r = .20 needs about 195. A two-way interaction of the same standardized size as a main effect requires roughly four times the sample, because the interaction contrast has twice the standard error, and if the interaction is the theoretically interesting half effect, it is sixteen times. A study powered for its main effect is essentially never powered for its moderation.
For structural equation models, the familiar rules of thumb (ten cases per parameter, N above 200) have poor support and depend heavily on loading size, number of indicators and model complexity. Reviewers with a psychometric background will accept a Monte Carlo power simulation for your specific model far more readily than a citation to a rule of thumb. For multilevel models, the constraint that matters is usually the number of level-2 units, not the total N: estimates of level-2 variance components and cross-level interactions are unstable with fewer than about 30 to 50 clusters, and small-sample corrections such as Kenward-Roger become necessary rather than optional.
This is a claim complaint, and it is the one where authors most often dig in and lose. A non-significant test is not evidence for the null hypothesis; it is a failure to reject it, and with a small sample it is what you would expect even if a substantial effect exists. Writing “there was no difference between groups” after p = .18 with 30 per group is the sentence reviewers are reacting to.
There are two honest ways out and they are both available after data collection. Either weaken the claim to what the data support, which is that the study did not detect a difference and that the interval is wide enough to be consistent with effects ranging from meaningful benefit to meaningful harm, or test absence properly with an equivalence test or a Bayes factor. The second is a genuine analysis, not a workaround, and it converts an embarrassing null into an informative one.
The instinct after this comment is to open G*Power, enter the effect size you observed, your sample size and your alpha, and report the resulting number as the study’s power. This is observed power, and it is not a diagnostic of anything. It is a deterministic transformation of your p value: any non-significant result yields observed power below 50%, always, by construction. It cannot tell you whether your non-significant result reflects a small effect or a small sample, because it is computed from the very estimate whose reliability is in question.
This is a well-established point in the methodological literature and it has fully reached reviewer culture. Reporting observed power in response to a power criticism now signals to a statistically literate reviewer that the authors do not understand the criticism, and it tends to escalate rather than close the comment. Some journals list it explicitly among the analyses they will not accept.
The legitimate replacement is sensitivity power analysis, which asks a different and answerable question: given the sample size I actually have, my alpha and a target power of 80%, what is the smallest effect I could have detected? That number is fixed by the design rather than by the result, so it is not circular, and it converts the discussion into something concrete. If your sensitivity analysis says the minimum detectable effect was d = 0.62 and the meta-analytic estimate for your effect is d = 0.30, you have just told the reader exactly what your study could and could not have seen. That is a limitation stated with precision, and reviewers accept it.
Four analyses are available to you without collecting a single additional participant, and between them they answer most versions of this comment.
Design analysis reasoning is worth understanding because it cuts both ways, and reviewers increasingly raise it. When a study is underpowered, the significant results it does produce are systematically exaggerated: to reach significance with a small N, the observed effect has to be large, so the published estimate overstates the true effect. That is the exaggeration ratio, sometimes called the Type M error. In severely underpowered designs there is also a non-trivial probability that a significant result has the wrong sign, the Type S error.
The practical consequence is that if your paper reports a significant effect from a small sample, the correct move is not to celebrate but to say plainly that the point estimate is likely to be inflated and that the effect size should not be used to power future studies. Authors who write that sentence themselves almost never receive it from a reviewer, and it materially improves how the discussion reads.
Three places in the paper need to change, and changing only one of them is what produces a second round of the same comment.
In the methods, add a sample size paragraph. It should state how the N was arrived at (a priori calculation with its parameters and source, a resource constraint, or a full enumeration of an available population), and, when no a priori calculation was performed, the sensitivity result: “With N = 84 and alpha = .05, the study had 80% power to detect a between-groups difference of d = 0.62 or larger. Effects smaller than this could not be reliably detected.” Cite the software and version.
In the results, report intervals alongside every primary estimate, and, where the primary result is null and you have run one, the equivalence test or Bayes factor. Do not bury these in supplementary material; the reviewer’s objection is about the primary claim, so the answer belongs where the primary claim is.
In the discussion, make the claim match the evidence. Replace “no effect was found” with the specific statement of what was and was not ruled out. If the study is genuinely exploratory or a pilot, say so in the title or the first line of the abstract rather than in the last paragraph, because a reviewer who discovers on page 22 that they have been reading a pilot study will be annoyed about the previous 21 pages.
The response letter for this comment has a reliable shape. Open by agreeing with the substance rather than the framing, because the substance is usually right and arguing about it costs you credibility for the rest of the letter. Then state what you have done, which analyses you added and where they now appear. Then state, without apology, the limit you are accepting and how the manuscript’s claims were changed to respect it.
A version that works: “We agree that the sample does not support strong conclusions about the moderation effect. We have added a sensitivity power analysis to the Methods (p. 9, lines 201 to 208) showing that the design could detect interaction effects of f² = 0.06 or larger at 80% power, and we now state in the Results that the interaction test was underpowered for the effect sizes reported in prior work. We have removed the moderation claim from the abstract and the title, and the interaction analysis is presented as exploratory throughout. We have also added a TOST equivalence test for the primary comparison (Table 3), which was non-significant, so we do not claim equivalence and instead report the interval.”
Notice what that response does not do. It does not report observed power, it does not promise future research as its main concession, and it does not defend the moderation claim. It gives ground on the claim, which is cheap, and keeps the paper, which is the point. If the reviewer instead demands more data, that is a different objection with different options, and it is worth reading the dedicated page on it before you answer.
You can compute it, but it will not help and will usually hurt. Observed power is a one-to-one function of your p value, so it adds no information beyond the result you already reported, and any non-significant test necessarily yields power below 50%. Statistically literate reviewers treat it as a signal that the criticism was not understood. Run a sensitivity power analysis instead, which asks what effect size your design could have detected and does not depend on what you found.
It solves the power equation for effect size instead of for N. You fix the sample size you have, the alpha level and a target power, usually 80% or 90%, and the analysis returns the smallest effect that design could detect at that power. Because it uses only design parameters, it is not circular. Report it for the primary test and interpret it against the effect sizes reported in the relevant literature.
Say what actually determined it. A resource constraint, a recruitment window, the size of the available population, or a precision target are all defensible justifications when stated plainly, and reporting guidelines accept them. Add a sensitivity power analysis so the reader knows what the resulting N could detect. The unacceptable option is a sentence claiming the sample was adequate without saying against what criterion.
A non-significant test never shows that. To make a claim about absence you need either an equivalence test, where you specify a smallest effect of interest in advance and show your effect is significantly inside that bound, or a Bayes factor quantifying evidence for the null against a stated alternative. Both are legitimate after data collection. Both require you to defend a choice, the equivalence bound or the prior, so make that choice on substantive grounds and report a robustness check.
Roughly four times what you need for a main effect of the same standardized magnitude, because the interaction contrast carries about twice the standard error. In practice, interaction effects are also usually smaller than the main effects they qualify, and when the interaction is half the size the requirement grows to about sixteen times. This is why most published moderation tests in samples of a few hundred are badly underpowered, and why reviewers treat significant interactions in small samples with suspicion rather than enthusiasm.
No. It is fatal when it is paired with a confident claim, and it is manageable when it is paired with a matching claim. Small samples are appropriate and publishable in rare conditions, in intensive longitudinal and single-case designs, in qualitative work with a different logic of sampling, and in studies estimating a large effect. What reviewers will not accept is a small sample analysed and discussed as if it were a large one.
Usually you should demote them rather than delete them. Move the underpowered subgroup or moderation analyses out of the abstract and out of the conclusions, label them explicitly as exploratory in the results, and report intervals rather than significance. Deleting them entirely raises a different problem, since selective removal of analyses after seeing them is itself a reporting concern. Transparency about what was run and how it should be read is the safer path.