How to Report Non-Parametric Tests in APA 7 (Mann-Whitney, Wilcoxon, Kruskal-Wallis)

In short

A non-parametric test needs four things in APA 7: the test statistic (U for Mann-Whitney, T or z for Wilcoxon, H for Kruskal-Wallis, χ²F for Friedman), the p-value (exact rather than asymptotic when the groups are small), an effect size your software probably did not print, and descriptives based on the median and the interquartile range, never on the mean and SD. It reads like this: "Group A (Mdn = 24.00, n = 32) scored higher than group B (Mdn = 19.50, n = 30), U = 312.50, z = −2.87, p = .004, r = .36."

Your data violated normality, you picked the right non-parametric test, SPSS (or JASP, or jamovi) handed you a U, an H or an asymptotic p... and then you freeze in front of the blank page. Do I report the mean or the median? Where is the effect size if the program does not give it to me? Is it Mann-Whitney's U or the z? This guide gives you a paragraph template and a table for each test, with the values you swap for your own.

Software is irrelevant: a non-parametric test is written up the same way in APA 7 whether you ran it in SPSS, JASP, jamovi or R. The output looks different in each program, but the results section always needs the same thing: the test statistic, the p-value (exact if the sample is small), an effect size, and descriptives based on the median, not the mean.

What APA 7 requires when reporting a non-parametric test

The 7th edition of the APA manual does not dedicate a section to non-parametric tests, but established practice in psychology journals sets a clear standard. A complete report includes: (1) the test's own statistic (U, T/W, H, χ²F or rs; (2) the p-value, exact when the sample is small; (3) an effect size, which almost no program reports by default and which the reviewer does expect; and (4) descriptives of central tendency and dispersion based on the median (Mdn) and the interquartile range (IQR), not on the mean and standard deviation.

The most common foundational error is precisely that last one: reporting M ± SD after choosing a rank-based test. If you opted for Mann-Whitney or Kruskal-Wallis it is because the mean does not represent your data well (skew, extreme values, ordinal scale); reporting the mean afterwards is contradictory, and seasoned reviewers spot it instantly. Pair each group with its Mdn and its IQR, and always state each group's sample size.

If the formatting is what's blocking you, the APA 7 results formatter gives you the full string from a loose statistic. And if what's missing is precisely the effect size almost no program computes for you (Mann-Whitney's r, Kruskal-Wallis's ε²…), effect size from text pulls it straight from your already-written result, no hand division required.

Reference table Each non-parametric test, its parametric analogue and what to report Test Parametric analogue Statistic Effect size Mann-Whitney independent-samples t U r = z/√N Wilcoxon (signed-rank) paired-samples t T / z r = z/√N Kruskal-Wallis one-way ANOVA H ε² Friedman repeated-measures ANOVA χ²F W Spearman Pearson correlation rs the rs itself Note. N = total sample size. For Mann-Whitney and Wilcoxon, r = |z|/√N (Rosenthal, 1991); Cohen's benchmarks: .10 small, .30 medium, .50 large. ε² (epsilon squared) = H/(N − 1) for Kruskal-Wallis (Tomczak & Tomczak, 2014). W = Kendall's W for Friedman (0 to 1). Descriptives in all cases: median (Mdn) and IQR, never mean ± SD.
Find your test, identify the statistic that gets reported and the effect size you must compute by hand.

Which statistic, which p and which effect size

The big omission is the effect size. SPSS and friends give you the statistic and the p, but rarely the effect, and since APA 6, reinforced in APA 7, reporting it is mandatory. For Mann-Whitney and Wilcoxon, the standard option is r = |z|/√N (Rosenthal, 1991), interpreted with the same cutoffs as a correlation. For Kruskal-Wallis, use epsilon squared, ε² = H/(N − 1), a proportion of variance explained between 0 and 1. For Friedman, Kendall's W. And for Spearman, the effect size is the coefficient rs itself.

On the p-value: with small samples (as a rough guide, any group with n < 20) report the exact significance, not the asymptotic one. The asymptotic approximation the program gives by default can be imprecise when there are few cases or many ties; SPSS, JASP and R offer the exact computation and that is what the reviewer expects to see under those conditions.

Results paragraph: template with real values

These are the templates for the four most common tests. Statistical symbols go in italics and the bracketed values are placeholders for your own:

Mann-Whitney (two independent groups). "Because [variable] was not normally distributed, a Mann-Whitney U test was used to compare [group A] and [group B]. [Group A] (Mdn = [x.x], n = [n₁]) scored [higher / lower] than [group B] (Mdn = [x.x], n = [n₂]), U = [value], z = [−x.xx], p = [.xxx], r = [.xx]."

Wilcoxon (two related measures). "A Wilcoxon signed-rank test was used to compare [pretest] and [posttest]. Scores at [posttest] (Mdn = [x.x]) were significantly [higher / lower] than at [pretest] (Mdn = [x.x]), z = [−x.xx], p = [.xxx], r = [.xx]."

Kruskal-Wallis (three or more groups). "A Kruskal-Wallis test revealed significant differences in [variable] across the [k] groups, H([df]) = [value], p = [.xxx], ε² = [.xx]. Pairwise comparisons (Dunn's test with [Holm / Bonferroni] correction) showed that [group X] (Mdn = [x.x]) differed from [group Y] (Mdn = [x.x]), p = [.xxx]."

Spearman (correlation). "The Spearman correlation between [variable 1] and [variable 2] was [positive / negative] and [significant], rs([df]) = [.xx], p = [.xxx], with [N] participants."

Style points reviewers check: the symbols (U, H, z, p, r, rs, Mdn, n) are italicized; the test names (Mann-Whitney, Kruskal-Wallis) are not. Coefficients that cannot exceed 1 (p, r, rs, ε²) are written without a leading zero: .032, not 0.032. For Kruskal-Wallis, the degrees of freedom of H are k − 1 (number of groups minus one); for Spearman, N − 2. And when your output does not match the template and you cannot find the z or the effect size, upload it to the output interpreter and it shows you where each value lives.

The descriptives table in APA 7 format

When you compare several groups, a table with each group's median and IQR communicates far better than burying the descriptives in the text. This is the structure APA 7 expects, with the test and its effect size in the note:

Table 1 Anxiety symptoms by treatment group (Kruskal-Wallis) Group n Mdn IQR Cognitive-behavioral therapy 38 9.0 6.0 – 12.0 Mindfulness 41 11.0 8.0 – 14.0 Waitlist (control) 40 15.0 12.0 – 18.0 Note. N = 119. IQR = interquartile range (25th – 75th percentile). Higher scores indicate more anxiety symptoms. The Kruskal-Wallis test indicated differences between groups, H(2) = 18.74, p < .001, ε² = .16. Dunn's test (Holm correction) showed that the control group differed from CBT (p < .001) and Mindfulness (p = .004); CBT and Mindfulness did not differ (p = .21). Table golden rules • Median and IQR, not mean ± SD. • n per group. • Test and effect in the note, not the header. • Clarify the scale direction (what scoring high means).
Example non-parametric descriptives table in APA 7. One row per group, median and IQR, and the test with its effect size in the note.

Your data don't meet the assumptions and you're not sure how to report it?

I am a PhD in psychology and I handle the full analysis of your thesis or paper: I pick the right test, compute the effect size and hand you the tables and the APA 7 text ready to paste. Fixed quote in 24 h.

Get a free assessment →

The five most common errors reviewers send back

After reviewing dozens of manuscripts with non-parametric tests, these are the problems that show up again and again:

1. Reporting the mean instead of the median. This is the error that fastest gives away an incoherent results section: you chose a rank-based test because the mean did not represent the data well, and then you report M ± SD. Always use Mdn and IQR.

2. Omitting the effect size. "U = 210, p = .01" is incomplete. Without r, ε² or W, the reader cannot tell whether the difference is trivial or substantial. Since almost no program computes it by default, do it by hand: r = |z|/√N is a single division.

3. Confusing Mann-Whitney with Wilcoxon. Mann-Whitney compares two independent groups; the Wilcoxon signed-rank test compares two related measures (the same subject, before and after). Naming them wrong, or applying the one that does not fit, is a design flaw, not a wording slip.

4. Using the asymptotic p with small samples. With groups of few cases or many ties, the asymptotic approximation the program gives by default is imprecise. Report the exact significance; it is one click away in SPSS, JASP and R.

5. Running Kruskal-Wallis and stopping there. A significant Kruskal-Wallis only says that some group differs, not which one. You need post-hoc pairwise comparisons (Dunn's test) with a correction for multiple comparisons (Holm or Bonferroni). Without that step, the conclusion about which groups differ is not justified.

Before you decide: do you really need the non-parametric test?

An important caveat before closing. A significant Shapiro-Wilk test does not automatically force you into the non-parametric version: with large samples, Shapiro-Wilk detects departures from normality that are irrelevant for the t-test or ANOVA, which are fairly robust. Before giving up the parametric version, check the sample size, the shape of the distribution and the statistical assumptions as a whole. The non-parametric test is the right tool when the scale is ordinal, the sample is small or there are extreme values distorting the mean, not as a reflex to any p < .05 in a normality test.

Checklist before submitting the manuscript

Before hitting submit, verify that your results section includes: the correct test statistic (U, T/z, H, χ²F or rs); the p-value, exact if the sample is small; the effect size (r, ε², W) interpreted with its benchmarks; the per-group descriptives with median and IQR; the total and per-group sample size; and, after a significant Kruskal-Wallis or Friedman, the pairwise comparisons with their correction.

Frequently asked questions

How do you report a Mann-Whitney U test in APA 7?

With the median and sample size of each group, the U, the z, the p and the effect size: "Group A (Mdn = 24.00, n = 32) scored higher than group B (Mdn = 19.50, n = 30), U = 312.50, z = −2.87, p = .004, r = .36." The z matters because it is what lets you compute the effect size; report both, not just the U.

Do I report the mean or the median with a non-parametric test?

The median, with the interquartile range. If you chose a rank-based test it is because the mean does not represent your data well, so reporting M ± SD afterwards contradicts the decision you just made. Give each group its Mdn, its IQR and its n.

Where do I get the effect size if my software does not give it?

You compute it by hand from what the output does give you. For Mann-Whitney and Wilcoxon, r = |z|/√N, with N the total sample size (Rosenthal, 1991), read with the same benchmarks as a correlation: .10 small, .30 medium, .50 large. For Kruskal-Wallis, epsilon squared = H/(N − 1). For Friedman, Kendall's W, which already runs from 0 to 1. For Spearman, the coefficient itself is the effect size.

Exact or asymptotic p-value?

Exact when the samples are small (as a rough guide, any group under 20 cases) or when there are many ties, because the asymptotic approximation is unreliable there. SPSS, JASP and R all offer the exact computation. With comfortable sample sizes the asymptotic p is fine, and it is worth stating which one you report.

What are the degrees of freedom for Kruskal-Wallis and Friedman?

For Kruskal-Wallis, k − 1, where k is the number of groups: with four groups, H(3). For Friedman, also k − 1, with k the number of repeated measures. Mann-Whitney and Wilcoxon do not carry degrees of freedom, so do not invent any; what you report there is the sample size.

Does a significant Shapiro-Wilk force me into a non-parametric test?

No. With large samples Shapiro-Wilk flags departures from normality that are irrelevant for the test, and what actually matters for a t-test or an ANOVA is the sampling distribution of the mean, not the distribution of the raw data. Look at the histogram and the Q-Q plot, consider the sample size, and remember that the non-parametric route also costs you power and changes the null hypothesis you are testing.

Which correction do I use for non-parametric post hoc comparisons?

After a significant Kruskal-Wallis, Dunn's test with Holm or Bonferroni adjustment, reporting the median of each group involved and the adjusted p. After a significant Friedman, pairwise Wilcoxon tests with the same logic. The correction is not optional and it must be named in the text.

Non-parametric tests are just one piece of the puzzle. If you want the equivalent template for their parametric analogues, it is in how to report a t-test and chi-square in APA 7 and, for correlation, in how to report correlations in APA 7 (where I distinguish Pearson from Spearman). And if what you need is the master guide with every template in one place, it is in how to report results in APA 7. If you would rather delegate the analysis and the write-up of your results, my statistical consulting service walks you through it end to end.

Keep reading

All blog articles