“Please report effect sizes for all analyses” is one of the shortest comments a reviewer can write and one of the most work it can generate. It looks like a formatting request. It is not: it asks you to decide what quantity your study is actually estimating, and different choices give different numbers from the same data, sometimes very different ones.
It is also close to non-negotiable. APA style has required effect sizes and confidence intervals for years, most psychology and health journals repeat the requirement in their author guidelines, and reporting guidelines for trials and observational studies ask for the magnitude of the effect with a measure of precision. A results section built entirely from p values is now unusual enough to attract attention on its own.
This page is a practical map: which effect size goes with which analysis, the denominator choices that quietly make Cohen’s d wrong, why partial eta squared is not comparable across designs, how to obtain confidence intervals for effect sizes that do not have a closed-form standard error, and how to write about magnitude without falling back on small, medium and large labels from 1988 that no longer describe most fields.
The first decision is between the standardized mean difference family, the correlational family and the ratio family. It is decided by your design and your outcome, not by preference, and choosing across families makes results incommensurable with the literature you want to be compared to.
The standardized mean difference is the default. Cohen’s d divides the raw difference by a standard deviation; Hedges’ g applies a small-sample correction that matters below roughly 20 per group and is negligible above it. Report g when your groups are small and say which you used, because reviewers cannot tell from a bare number.
The choice that actually changes the result is the denominator. The pooled standard deviation of both groups is standard. Glass’s delta uses the control group standard deviation only, which is preferable when the intervention plausibly changed variability as well as the mean, a common situation in treatment studies. Using one group’s standard deviation when the two differ substantially can change d by a third or more. Whatever you choose, name it in the analysis section in one clause.
This is where most reporting errors happen, and reviewers with methods training look for it specifically. There is more than one Cohen’s d for a paired design. Standardizing the mean change by the standard deviation of the difference scores, sometimes written d subscript z, produces larger numbers when pre and post are strongly correlated, because a high correlation shrinks the denominator. Standardizing by the average of the pre and post standard deviations, or by the pretest standard deviation alone, produces a value comparable with between-groups effect sizes and with meta-analyses.
The consequence is real: the same pre-post change can be reported as d = 0.45 or d = 1.10 depending on the choice, and neither is wrong, they answer different questions. For meta-analytic comparability, the pretest or averaged standard deviation is what you want. State the choice explicitly, and if your outcome is a change score in an intervention study, consider reporting the raw change with its interval as well, because the raw units are what a clinician can interpret.
Eta squared gives the proportion of total variance explained by an effect, so all effects in a design sum to something interpretable. Partial eta squared removes the variance explained by the other effects from the denominator, so it is larger, often much larger, and it does not sum to anything. Statistical software reports partial eta squared by default, which is why it dominates published tables.
The problem reviewers raise is comparability. Because the denominator of partial eta squared depends on which other factors are in the design, the same underlying effect yields different values in a one-way and a three-way design, and in a within-subjects and a between-subjects version of the same study. Values are therefore not comparable across studies with different designs, and using them to build a meta-analysis or to power a new study is a known trap. Generalized eta squared was developed precisely to be comparable across designs and is the better choice when your effect will be compared with others; it is available in standard R packages and increasingly requested. Omega squared is the less biased alternative to eta squared and is worth using with small samples, where eta squared is noticeably upward biased.
For a bivariate association, r is itself an effect size, and R squared for the model. In multiple regression, report the standardized coefficient with its confidence interval, and the semipartial correlation squared for the unique variance attributable to each predictor, which is what readers usually want when they ask how much a predictor contributes. Note that standardized coefficients depend on the variances in your sample, so they are not comparable across samples with different degrees of range restriction; unstandardized coefficients with intervals are preferable when the units mean something, as they do for age, dose or days.
For hierarchical models, the change in R squared for each block with its confidence interval is the effect size for the increment, and f squared converts it into the metric power analyses use.
For a contingency table, Cramér’s V for association and the odds ratio, risk ratio or risk difference for two-by-two designs. Which of these you choose matters more than authors assume: odds ratios exaggerate relative to risk ratios when the outcome is common, above roughly 10%, and are routinely misread as risk ratios by readers and sometimes by authors. In clinical papers the risk difference and the number needed to treat are the quantities practitioners can use, and health journals increasingly ask for them.
For logistic regression, report the odds ratio with its interval, and consider adding predicted probabilities at representative covariate values, since a coefficient on the log-odds scale is uninterpretable without them. For rank-based tests, the rank-biserial correlation for Mann-Whitney and Wilcoxon, and epsilon squared for Kruskal-Wallis, are the matching effect sizes, and reporting a Cohen’s d next to a non-parametric test is a mismatch reviewers notice.
Reviewers who ask for effect sizes almost always mean effect sizes with intervals, and a table of point estimates alone often produces the same comment a second time. The interval is the part that carries information about precision, and in small studies it is the part that changes how the result should be read.
For d and g, the interval is not symmetric and is not obtained by adding and subtracting a multiple of a standard error, because the sampling distribution is a noncentral t. Standard software now computes these properly: the effectsize and MBESS packages in R, the built-in options in jamovi and JASP, and the confidence interval settings in recent SPSS versions. For eta squared and its relatives, the interval comes from the noncentral F distribution, and a one-sided lower bound is often reported because the quantity is bounded at zero.
For anything without a clean analytic interval, or for effect sizes computed on transformed or trimmed data, bootstrap percentile intervals with a few thousand resamples are acceptable and easy to justify. Say how many resamples and which interval type you used.
Two habits make intervals more useful in the write-up. First, interpret the width, not only whether the interval contains zero, because an interval on d from -0.05 to 0.98 and one from 0.44 to 0.49 tell you entirely different things while both being “significant” or “not” in the usual reading. Second, report intervals for the effects that did not reach significance too. Selective reporting of intervals only for significant results is a pattern reviewers pick up quickly.
The second half of a good effect size report is the interpretation, and this is where the small, medium, large labels cause trouble. Those benchmarks were offered as a rough fallback for situations where no field-specific information existed, and their author said as much. Applied mechanically, they mislead in both directions.
In individual differences and social psychology, the typical published correlation sits near r = .20, and correcting for publication bias pushes the realistic figure lower. Calling such an effect “small” implies it is unimportant, when an effect of that size accumulating across a population or across repeated exposures can matter a great deal. In psychotherapy outcome research, controlled comparisons against active treatments rarely exceed d = 0.3, while comparisons against waitlist inflate to d = 0.8 or above, so the same label attached to two studies can describe completely different evidential situations.
The interpretations that persuade reviewers are comparative and consequential. Comparative: place your effect against the distribution of effects in your own literature, ideally citing a meta-analysis, and say whether yours is at the low end, typical or unusually large. Consequential: translate into a scale readers understand. A change of 3.2 points on the PHQ-9 against a minimal clinically important difference of around 5, a shift of 0.4 standard deviations expressed as the probability that a randomly chosen treated participant scores better than a randomly chosen control, a risk difference expressed as events avoided per hundred patients. One sentence of translation is worth a paragraph of labels.
Do not, however, over-claim in the other direction. Reviewers are alert to effect sizes described as “meaningful” or “practically significant” without a stated benchmark. If you assert clinical relevance, name the threshold you are using and where it comes from.
Three places, and skipping the first is what makes reviewers ask twice.
This comes up with reanalysis of archival data, with collaborators who have moved on, and with software output you no longer have. Most effect sizes can be recovered from summary statistics. Cohen’s d can be reconstructed from group means, standard deviations and sample sizes, or from a t statistic and the group sizes. Eta squared can be recovered from an F statistic and its degrees of freedom. An odds ratio comes straight from a two-by-two table, and r from t and df. Conversion formulas between families are standard in meta-analysis texts and implemented in packages such as esc and metafor.
Two cautions. Conversions between families, for example d to r, assume things about the underlying distributions and about how any dichotomization was done, and they propagate any error in the original reporting. And a conversion from a rounded F statistic inherits that rounding, so give one decimal fewer than you would from raw data. If you reconstruct rather than compute, say so in a footnote; reviewers respect it and it costs nothing.
This is one of the easiest reviewer comments to close completely, and it is worth closing completely rather than partially, because a half-answer here invites a closer look at the rest of the analysis.
A response that works: “We have added effect sizes with 95% confidence intervals throughout. For the between-groups comparisons we report Hedges’ g with the pooled standard deviation; for the pre-post comparisons we report g standardized on the pretest standard deviation, so that the two sets of values are comparable with each other and with the meta-analytic estimates we cite in the Discussion. Intervals were obtained from the noncentral t distribution. Tables 2 and 3 now include an effect size column, the analysis section describes the choices on p. 10, and we have added a sentence to the Discussion placing our primary effect against the pooled estimate reported by [reference].”
Two additions raise the quality of the answer. If replacing partial eta squared with generalized eta squared changed the apparent size of an effect, say so rather than letting the reviewer discover it. And if adding intervals revealed that an effect you had described as robust is estimated imprecisely, change the description in the same revision. Reviewers who see authors following the numbers where they lead tend to become considerably easier to deal with in round two.
Cohen’s d, or Hedges’ g when groups are small, standardized on the pooled standard deviation for a between-groups comparison. State the denominator explicitly. If the intervention plausibly changed the variability as well as the mean, Glass’s delta using the control group standard deviation is defensible and sometimes preferable. Always report the interval alongside the point estimate.
Eta squared expresses an effect as a proportion of the total variance in the design, so the values across effects are on a common scale. Partial eta squared removes variance attributable to the other effects from the denominator, which makes it larger and makes it depend on what else is in your design. That dependence is why partial eta squared is not comparable across studies with different designs, and why generalized eta squared is the better choice when comparability matters.
Not by adding and subtracting a standard error, because the sampling distribution of d is a noncentral t and the interval is asymmetric. Use software that computes it from the noncentral distribution: the effectsize or MBESS packages in R, jamovi, JASP, or the effect size interval options in recent SPSS versions. Bootstrap percentile intervals are an acceptable alternative, particularly for effect sizes computed on trimmed or transformed data.
They are widely used but weakly justified, and they were offered as a fallback for fields with no better information. Reviewers increasingly ask for interpretation against the effect sizes actually observed in your literature, ideally from a meta-analysis, or against a clinically meaningful threshold for your outcome measure. Using the labels is not an error; relying on them alone as the interpretation is what draws the comment.
Yes, and they are arguably more important there. A non-significant test with an effect size and a wide interval tells the reader the study was uninformative; the same test with a small effect size and a narrow interval tells them the effect is probably close to zero. Those are opposite conclusions and only the effect size distinguishes them. Reporting intervals only for significant results is a pattern reviewers notice and question.
Usually yes. Cohen’s d can be reconstructed from means, standard deviations and group sizes, or from t and the degrees of freedom; eta squared from F and its degrees of freedom; odds ratios from a two-by-two table. Meta-analysis packages implement the conversions. Note in a footnote that the values were derived from summary statistics, and be careful with cross-family conversions, which rest on distributional assumptions and inherit any rounding in the original report.
The rank-biserial correlation, which expresses the probability that a randomly chosen observation from one group exceeds one from the other and is directly interpretable. For Kruskal-Wallis, epsilon squared. Reporting Cohen’s d next to a rank-based test is a mismatch, since d refers to means while the test refers to ranks, and reviewers with methods training will flag it.