5 Common Statistical Errors in Psychology Theses

After having provided statistical consulting to dozens of doctoral students and having reviewed more theses than I can count, I can affirm that statistical errors in psychology theses follow a predictable pattern. They are not exotic or sophisticated errors. They are, for the most part, basic errors that repeat themselves over and over again, and that could be avoided with somewhat more solid methodological training or with a timely statistical consultation. What is frustrating is that many of these errors are detected during the defense or review phase, when it is already too late to go back through the data and the time invested cannot be recovered.

The most common error, by far, is failing to plan the statistical analysis before collecting the data. Many doctoral students design their study, collect the sample, and only then ask themselves what analyses they should perform. This leads to absurd situations: samples too small for the intended analyses, variables measured in a format that does not allow the desired analysis, or designs that do not include the control conditions necessary to answer the stated hypotheses. An a priori power analysis, which is a calculation that can be done in five minutes with G*Power or our effect size calculator, tells you how many participants you need. Not doing it is like building a house without blueprints and hoping it turns out well.

Errors in the selection and application of analyses

A frequent error is applying parametric tests without verifying their assumptions. If you are unsure which test fits your design, our statistical test selector walks you through it step by step. The t-test and ANOVA assume normality of residuals and homogeneity of variances, and although they are reasonably robust to moderate violations of these assumptions (especially with large samples), there are situations in which the violations are severe enough to invalidate the results. Paradoxically, many doctoral students apply normality tests (Shapiro-Wilk, Kolmogorov-Smirnov) but then do not know what to do with the results. If the normality test is significant, they automatically switch to a nonparametric test, without considering that normality tests have high power with large samples and can reject normality due to trivial deviations that do not affect the validity of the ANOVA.

Another classic error is the indiscriminate use of Pearson correlations without visually exploring the relationship between variables. A Pearson correlation measures the linear relationship between two variables, and it can be close to zero even when a strong but nonlinear relationship exists. It is also very sensitive to outliers: a single extreme point can dramatically inflate or deflate the correlation coefficient. The solution is simple: before calculating any correlation, create a scatterplot. If the relationship does not appear linear or if there are obvious outliers, the Pearson correlation is not the appropriate tool.

The confusion between correlation and causation is perhaps the most serious and most difficult conceptual error to eradicate. Many doctoral students write causal conclusions based on correlational data, with phrases such as "self-esteem produces better academic outcomes" when what they have actually shown is an association between the two variables. The direction of causality could be the reverse (good outcomes improve self-esteem), there could be a third variable that explains both (socioeconomic status, for example), or the relationship could be bidirectional. Only experimental designs with manipulation and randomization allow legitimate causal inferences.

Problems with multiple comparisons

When a doctoral student has a questionnaire with 20 scales and compares two groups on each of them, they are performing 20 simultaneous statistical tests. With a significance level of 0.05, the probability of obtaining at least one false positive is no longer 5% but approximately 64%. This is what is known as the multiple comparisons problem, and the number of theses that completely ignore it is alarming. The Bonferroni correction (dividing alpha by the number of comparisons) is the best known but also the most conservative, and it can cause you to miss real effects. Alternatives such as the Holm correction or false discovery rate (FDR) control offer a better balance between false positive control and statistical power.

A related problem is what we might call exploration disguised as confirmation. The doctoral student has a vague hypothesis, explores the data in multiple directions, finds a significant result, and writes the thesis as if that had been their hypothesis from the beginning. This practice, known as HARKing (Hypothesizing After Results are Known), is not necessarily dishonest in its intent, but it does produce misleading science. The solution is simple in principle: clearly distinguish in the thesis between confirmatory analyses (planned a priori) and exploratory analyses (conducted post hoc), and be transparent about the process.

How to avoid these errors

The best prevention against statistical errors in a thesis is to consult with a specialist statistician for doctoral theses before collecting the data, not after. A one-hour consultation during the design phase can save you months of lost work. If your university does not offer this service, seek external statistical consulting. It is an investment that pays for itself in time and frustrations avoided.

It is also important that your thesis supervisor has sufficient statistical competence to oversee your analyses, or that they openly acknowledge their limitations and refer you to someone who can help. Unfortunately, not all supervisors have this honesty, and some doctoral students inherit methodological errors from their supervisors without knowing it. If something does not convince you about the analyses you have been advised to use, seek a second opinion. In the end, the thesis bears your name, and the responsibility for the correctness of the analyses is yours. Which is exactly why it helps to know what you will be asked about them on defense day: simulate the committee with your own thesis and check whether any of these errors surface in the questions.

Keep reading

All blog articles