In short
A violated assumption does not sink a paper; an unexplained one does. Every paragraph that survives review makes the same three moves in the same order: what you checked and what came out, with the statistic and its p; the rule you applied, written down before you looked; and the analysis you actually ran, with its effect size and confidence interval. The default alternatives are stable enough to memorise: normality goes to a bootstrap interval or a rank test run as a sensitivity analysis; homoscedasticity goes to Welch (t or F) or to HC3 robust standard errors; sphericity goes to Greenhouse-Geisser when ε < .75 and Huynh-Feldt above it; linearity goes to a transformed predictor, a centred quadratic term or a model with the right link function; independence goes to a mixed model, with the ICC reported in the text. What gets flagged is never the correction itself, it is the sentence “the data were not normally distributed, so non-parametric tests were used”, with no statistic, no criterion and no word on whether the conclusion changed.
Nobody writes to a forum to ask what Levene's test is. They write because the residual plot has come out as a fan, because the software has just printed a significant Shapiro-Wilk, because the deadline is Friday, and because every answer they get is “use a non-parametric test” with not one word about how the sentence should read. This article is the missing half: one model paragraph for each of the five assumptions that actually come up in review, with the numbers your reader needs, the alternative that fits, and wording you can paste into your draft and swap the values in. If what you still need is the checking rather than the writing, it is laid out step by step in how to verify statistical assumptions.
The three moves that make one of these paragraphs publishable
Read enough review reports and you notice that the objection is almost never “your data violated an assumption”. It is “I cannot tell what you decided, when you decided it, or what it cost you”. Someone reading “normality was not met, so non-parametric tests were used” has been handed two claims and no way to connect them: no statistic, no threshold, no sample size and no way of knowing whether the parametric analysis would have said something different. The repair is three sentences and about sixty words, and it is the same three sentences every time.
Move one. What you checked and what came out, with the statistic and its p, not with an adjective. Move two. The rule you applied, phrased so it is clear you had it before the data: “as specified in the analysis plan”, “Welch's test was used for all comparisons”. Move three. The analysis you actually ran, with its effect size and confidence interval, plus the sensitivity analysis if you ran one and what it showed. Nothing else belongs in the paragraph. The hand-wringing goes in the limitations.
There is a trap in the middle of all this, and it is precisely the habit the software encourages: run the assumption test and let its p decide which analysis you report. That two-step procedure has been studied and it does not behave: conditioning the choice of test on a preliminary test raises the false positive rate of the pair above the nominal 5%, because you have added a data-dependent decision to the chain (Zimmerman, 2004; Rasch, Kubinger and Moder, 2011). The way out is not to stop checking, it is to stop letting the check decide. Pick the option that stays valid in both worlds and declare it in advance: Welch instead of Student, HC3 instead of classical standard errors, a bootstrap interval instead of the normal-theory one. When the assumption holds you lose almost nothing, and when it fails you keep your error rate.
Three sentences, three homes. The rule, the software and the version go in the Method, in the voice of a plan that was already made: “all group comparisons used Welch's t test; heteroscedasticity-consistent HC3 standard errors were used in all regression models”. The numbers, the decision and the sensitivity analysis go in the Results, next to the test they qualify and not in a block of checks at the top of the section. And if the violation genuinely limits what you can claim — a dependency you could not model, a curvature you could only flag — that sentence goes in the Discussion, once, and says which way the bias runs.
Three sentences get flagged with almost comic regularity, and it helps to recognise them before you write them. The first: “the data were normally distributed”, with not one number behind it. The second: “non-parametric tests were used given the violation of assumptions”, without naming the assumption or the test. And the third, the subtle one: “the assumption of homogeneity of variances was met (p = .21)”, written as though a non-significant test proved equality, when all it indicates is that at this sample size there was no power to detect the difference.
One warning before the five paragraphs: assumptions belong to the model, not to your variables. Regression and analysis of variance do not ask that X or Y be normally distributed, they ask it of the residuals; the chi-square test asks for nothing of the sort, it asks for adequate expected frequencies; and the Mann-Whitney test, which many people treat as a shelter, carries its own assumption of comparable shapes if you want to read it as a difference between medians. Checking the wrong assumption before the wrong test is the fastest way to spend half a page saying nothing.
Normality: rarely the real problem, almost always badly written up
What the reviewer writes. “Shapiro-Wilk was significant and a t test was nevertheless applied.” Or the mirror image, just as common: “the authors switched to Mann-Whitney and then interpret means”. Both ask the same thing: on what basis did you choose, and does the answer change depending on which route you take?
What sits underneath. Three facts that rarely make it into the manuscript. The first is that normality is asked of the residuals, not of the raw variable: a variable that is bimodal by design can produce impeccable residuals. The second is that the power of Shapiro-Wilk grows with sample size, so at n = 300 it flags irrelevant departures and at n = 20 it flags nothing; the statistic that carries the magnitude is the skewness, not the p. The third is that real psychological data are rarely normal (Micceri, 1989), and that the F test holds its nominal error rate under the degrees of skew that turn up in real samples (Blanca, Alarcón, Arnau, Bono and Bendayan, 2017). The practical consequence: non-normality damages the coverage of your confidence interval long before it damages the point estimate.
What the alternatives are. Four of them, each with a price. The bootstrap (percentile, or better, bias-corrected and accelerated, with 5,000 resamples) keeps the analysis and the interpretation on the original scale and changes only how precision is estimated; it is the default when you still want to talk about means. A rank test is legitimate, but it answers a different question: Mann-Whitney asks whether a case drawn at random from one group exceeds one drawn from the other, and it is equivalent to a comparison of medians only if you assume similar shapes (Fagerland and Sandvik, 2009). A transformation fixes the distribution and commits you to interpreting on the transformed scale, which has to be said out loud. And robust estimators — 20% trimmed means, Yuen's t — keep the logic of the comparison and remain the great unknown of the list.
Model paragraph (normality). Anxiety scores were positively skewed in both groups (skewness = 1.42, SE = 0.22; kurtosis = 2.11, SE = 0.44), and Shapiro-Wilk rejected normality of the residuals, W = .94, p < .001. As specified in the analysis plan, the planned comparison was retained and its precision estimated with a bias-corrected and accelerated bootstrap (5,000 resamples), given that with n = 61 and n = 57 the sampling distribution of the mean difference is approximately normal and the t test holds its nominal error rate. Participants in the guided condition scored above those in the self-help condition, Mdiff = 3.18, 95% CI [0.55, 5.81], t(116) = 2.41, p = .018, d = 0.44, 95% CI [0.09, 0.79]. A Mann-Whitney U test, pre-specified as a sensitivity analysis, led to the same conclusion, U = 1,315, z = −2.28, p = .023, r = .21.
What makes that paragraph strong is not the bootstrap, it is that both routes are visible and they agree. If they disagreed, the next sentence would be just as short and considerably more valuable: “the two tests diverge (p = .018 versus p = .071), so the result is treated as preliminary and depends on three cases with extreme scores in the guided condition”. No reviewer rejects a paper for that kind of honesty. If you want the detail on why the Shapiro-Wilk p is so widely misread, it is in the normality test almost everyone reads backwards.
Homoscedasticity: Welch by default, HC3 in regression
What the reviewer writes. “Levene's test was significant (p = .002) and the authors report Student's t.” In regression the same objection arrives with a different face: “the plot of residuals against fitted values shows a clear fan; the standard errors are not credible”.
Group comparisons: Welch. Welch's correction does not compare variances: it stops pooling them into a single estimate and adjusts the degrees of freedom so the test stays valid when they differ. It costs very little power when variances are equal and it saves the analysis when they are not, especially in the dangerous combination: unequal group sizes with the larger variance in the smaller group, which makes Student's test liberal. That is why it is recommended as the default whether or not the assumption holds (Delacre, Lakens and Leys, 2017). The visible signature that you used it is the degrees of freedom, which stop being whole numbers: t(89.14) instead of t(116). For follow-up comparisons, Games-Howell instead of Tukey, for the same reason.
Model paragraph (unequal variances between groups). Burnout scores showed unequal variances across the three settings, Levene's F(2, 211) = 6.41, p = .002, with the widest dispersion in the community teams (SD = 1.18) and the narrowest in the hospital teams (SD = 0.71). Welch's analysis of variance, used for all comparisons as stated in the analysis plan, indicated differences between settings, F(2, 121.47) = 5.83, p = .004. Games-Howell comparisons, which likewise do not assume equal variances, located the difference between hospital teams (M = 3.42, SD = 0.71) and community teams (M = 2.94, SD = 1.18), mean difference = 0.48, 95% CI [0.06, 0.90], p = .021, and between hospital and primary care teams (M = 3.08, SD = 0.94), mean difference = 0.34, 95% CI [0.02, 0.66], p = .034.
Regression: robust standard errors. Here heteroscedasticity does not touch the coefficients, it touches their precision: the estimate of B stays unbiased and what stops being trustworthy is the SE, and with it the p value and the interval. The standard fix is heteroscedasticity-consistent standard errors; in samples below roughly 250 cases the recommended variant is HC3 (Long and Ervin, 2000), available through sandwich::vcovHC(model, type = "HC3") in R, the robust option in Stata, and the complex samples menu or a macro in SPSS. Before you reach for it, one uncomfortable question is worth asking: a fan in the residuals is sometimes not a standard error problem at all, it is a misspecified model missing an interaction or a non-linear term. If that is what it is, fix it there and the heteroscedasticity leaves on its own.
Model paragraph (heteroscedasticity in regression). Residuals from the model of wellbeing on workload, tenure and perceived support showed increasing spread across fitted values, Breusch-Pagan χ²(3, N = 204) = 11.86, p = .008. All coefficients are therefore reported with heteroscedasticity-consistent HC3 standard errors (Long and Ervin, 2000); the point estimates are unaffected by this choice. Perceived support predicted wellbeing, B = 0.38, SEHC3 = 0.14, t(200) = 2.71, p = .007, 95% CI [0.10, 0.66]; the classical standard error for the same coefficient was 0.11, so the conventional analysis would have made the estimate look roughly a quarter more precise than it is. The full breakdown of Levene's test, including how to write it up when it is not significant, is in how to report Levene's test in APA 7.
Sphericity: epsilon and the degrees of freedom with decimals
What the reviewer writes. “Please report the value of epsilon and the corrected degrees of freedom.” And the variant that gives away that you applied nothing: “the degrees of freedom for the within-subjects effect are whole numbers, so no correction has been used”. It is the easiest of the five to spot, because it shows in the number without anyone reading the text.
What sits underneath. Sphericity asks that the variances of the differences between every pair of measurements be equivalent, and with three or more repeated measurements it almost never holds, because assessments close in time correlate more strongly than distant ones. When it fails, the within-subjects F test becomes liberal. The correction does not recompute the F: it multiplies both degrees of freedom by epsilon, and the p value moves with them. The effect size does not move either, because the sums of squares are identical. With two measurements there is nothing to correct — only one difference exists — and the software prints W = 1.000.
Model paragraph (sphericity). Mauchly's test indicated that the assumption of sphericity was violated for the effect of time, W = .52, χ²(5, N = 52) = 32.51, p < .001. Degrees of freedom were therefore corrected using the Greenhouse-Geisser estimate, ε = .68, following the convention of applying Greenhouse-Geisser below .75 and Huynh-Feldt above it (Girden, 1992); the Huynh-Feldt estimate was ε = .72. Emotion regulation scores changed across the four assessments, F(2.04, 104.04) = 9.47, p < .001, η2p = .16. Pairwise comparisons used a Bonferroni correction for six contrasts, and their confidence intervals carry the same correction.
The elegant way out. If you have missing measurements — and at a six-month follow-up you do — repeated measures analysis of variance deletes the whole case as soon as one assessment is missing, and that is a far higher price than sphericity. A mixed model with an unstructured or first-order autoregressive covariance structure does not need the assumption, uses the incomplete cases, and is reported with Satterthwaite degrees of freedom, which also come out with decimals. When it is worth changing model, and how it is written up, is in mixed models versus repeated measures ANOVA.
Linearity: transform it, bend it, or change the model
What the reviewer writes. “The relationship looks curvilinear and a straight line has been fitted.” Or, in a dose design: “the authors treat an obviously saturating dose-response relationship as linear”. It is the least checked of the five, because it has no test with a p value the software prints on its own, and the one that changes the interpretation most when it fails: a slope near zero can hide a perfect inverted U.
How you check it. With the plot of residuals against fitted values and a LOESS smoother on top, which is where curvature becomes visible; in logistic regression, with the Box-Tidwell test, which examines linearity in the logit rather than in the variable itself; and with a quick, very honest diagnostic that consists of splitting the predictor into quartiles, plotting the outcome means and looking at whether they rise in a straight line. That plot does not go in the paper, it goes into your decision.
What the alternatives are. Four, ordered from smallest to largest conceptual change. Transform the predictor (log, square root) when the effect is proportional rather than additive; this changes what the coefficient means and has to be spelled out. Add a quadratic term for the centred predictor: centring before squaring avoids the artificial collinearity between the linear and quadratic terms. Splines or a generalised additive model, flexible and harder to narrate: you report the effective degrees of freedom and always add a figure, because the coefficient means nothing on its own. And change the model family: a count of relapses is not a normal outcome, and a Poisson or negative binomial regression usually makes the “non-linearity” disappear alongside the heteroscedasticity, because both were symptoms of fitting the wrong model.
Model paragraph (non-linearity). Residual plots for the model of psychological distress on daily screen use showed systematic curvature, with a LOESS smoother flattening above roughly seven hours. A quadratic term for the centred predictor was therefore added. It improved fit, ΔR2 = .04, F(1, 197) = 9.12, p = .003, and was itself reliable, B = −0.045, SE = 0.015, t(197) = −3.02, p = .003, 95% CI [−0.074, −0.015], indicating that the positive association between use and distress flattened at around 7.8 hours per day. Because only 4% of the sample reported more than seven hours of daily use, the turning point is estimated from a sparsely populated region of the predictor and is described as a feature of this sample rather than as a threshold.
That last sentence is what separates a finding from a statistical anecdote. A quadratic term bends wherever you let it, and where you let it there are usually fifteen cases. Writing the limitation yourself, with the percentage in front of it, costs one line and defuses the objection before it arrives.
Independence: the one assumption you cannot patch afterwards
What the reviewer writes. “Patients were treated by 58 therapists and the analysis treats the 480 observations as independent.” Or, in educational research, the same sentence with classrooms; or, in a longitudinal design, with the measurements inside each person. It is the most serious violation of the five and the only one with no correction you can bolt on at the end: here you change the model or you do not fix it.
What sits underneath. A small piece of arithmetic worth knowing by heart. The intraclass correlation measures how much of the total variance sits between clusters; the design effect is 1 + (average cluster size − 1) × ICC, and the effective sample size is your N divided by that design effect. With ICC = .14 and an average of 8.3 patients per therapist, the design effect is 2.02: your 480 cases are worth 238. The standard error comes out too small by a factor of about 1.42, and it errs in exactly the direction that suits you, which is why almost nobody catches it on their own.
What the alternatives are. The mixed model with a random intercept for the clustering is the standard answer: you report the number of clusters, the ICC from the empty model, the random structure and the degrees of freedom method. Cluster-robust standard errors are the alternative when the clustering is a nuisance rather than an object of study, with the caveat that below roughly 40 clusters you need the small-sample CR2 correction with Satterthwaite degrees of freedom. In time series the dependency is called autocorrelation, is detected by a Durbin-Watson statistic far from 2 (a value of 1.12 indicates positive autocorrelation), and is handled with an AR(1) error structure. And if you genuinely cannot model it — six sites is too few clusters for almost anything — the honest way out is to report the ICC anyway and state which way it biases the result.
Model paragraph (non-independence). The 480 patients were treated by 58 therapists (median caseload = 8, range 3-19), so observations were not independent. The unconditional model gave an intraclass correlation of ICC = .14 for therapist which, with an average cluster size of 8.3, implies a design effect of 2.02 and an effective sample size of approximately 238 rather than 480. All models therefore include a random intercept for therapist, with Satterthwaite degrees of freedom. Working alliance predicted outcome, B = 0.27, SE = 0.09, t(418.6) = 3.00, p = .003, 95% CI [0.09, 0.45]. The same coefficient estimated by ordinary least squares carried a standard error of 0.06 and a p value below .001, which illustrates the direction of the bias introduced by ignoring the clustering.
Look at that last sentence, because it does most of the work. It does not apologise: it shows the number the naive analysis would have produced and makes clear that the correction cost you significance rather than handing it to you. A reviewer who sees that stops wondering whether you knew what you were doing. And if the clustering only became apparent after data collection, that same sentence, with the ICC in front of it, is also the limitation that goes in the Discussion.
Frequently asked questions
Can I still use a t test if my data are not normally distributed?
Usually yes, and the reason is that the t test does not assume your scores are normal, it assumes the sampling distribution of the mean difference is approximately normal; with groups of about thirty cases or more it usually is, even when the raw scores are skewed. What you cannot do is pass over it in silence. Report the skewness and the Shapiro-Wilk result, state that the comparison was retained by prior decision, put a bootstrap interval around the effect, and add a rank test as a sensitivity analysis. The problem is never the t test, it is the missing sentence explaining why it is still there.
Levene's test is significant. What do I do?
Use the Welch version of the test you were going to run — Welch's t for two groups, Welch's F for more — and say so. It is the same procedure recommended as a default even when Levene is not significant (Delacre, Lakens and Leys, 2017), so your decision does not look opportunistic. The visible signature is the degrees of freedom, which stop being whole numbers: t(89.14) instead of t(116). For follow-up comparisons, swap Tukey for Games-Howell. In regression, the equivalent move is HC3 standard errors.
Do assumption checks go in the Method or in the Results?
The rule goes in the Method and the numbers go in the Results. The Method states what you were going to do: “all group comparisons used Welch's t test; regression models used HC3 standard errors; the Greenhouse-Geisser correction was applied whenever epsilon fell below .75”. The Results state what happened, next to the test it qualifies and not in a block of checks at the top of the section. Your reader should never have to hold six assumption results in their head for two pages before finding out which test they belong to.
How do I write up that I used a non-parametric test?
Name the test, say what it compares, give the statistic, the exact p, an effect size and the descriptive statistics that match it: medians and interquartile ranges, not means, if the test works on ranks. Say whether it was planned or is a sensitivity analysis, and say whether it agreed with the parametric result. The sentence that raises suspicion is the one that reports a rank test and then discusses means. If formatting is what slows you down, the APA 7 results formatter hands back the string with the italics and leading zeros already right.
Is it wrong to change the test after seeing the assumption results?
It is not fraud, but it carries a cost that is rarely acknowledged: the pair “preliminary test, then choice of test” has a higher false positive rate than either test on its own, because the choice now depends on the data (Zimmerman, 2004). The way to keep the flexibility without the cost is to pre-specify the robust option and run it regardless. And if you did switch after looking, say so in those words and report both analyses: a declared switch with matching conclusions costs you nothing, and a hidden one can cost you the paper.
What if the violated assumption cannot be fixed?
Then it becomes a limitation with a direction, not an apology. Say what the violation is, say which way it biases the estimate, and quantify it if you can: clustering you could not model inflates precision, so your intervals are too narrow; a floor effect compresses variance, so your correlations are attenuated. “Results should be interpreted with caution” is not that sentence: it tells your reader to worry without telling them what about.
Which of your five assumptions is a sentence with no number in it?
That is the one the reviewer will circle. The reviewer here does it first, on your own Results section, while you can still change it.
Run your Results through the reviewer →With the three moves in place — what you checked, which rule you applied, which analysis you ran — a violated assumption stops being an exposed flank and starts reading as evidence that you knew what you had got into. If the one in front of you is among those with no clean patch, with nested data you cannot model or a curvature you cannot name, we can look at it together in my statistical consulting.