Calculating the sample size is one of the first things a thesis committee or a reviewer asks for, and one of the worst justified. It is not "the more, the better": an underpowered study fails to detect the effect it is looking for (and wastes the whole effort), and an oversized one burns resources and exposes participants for nothing. This guide walks through how to do it properly, before collecting data, with G*Power. If you just need a quick number for a standard test, use the sample size calculator; if you need to understand and defend the calculation, read on.
A priori, never post hoc
Sample size is calculated before collecting data (a priori). Computing power after you have the results, using the observed effect (post hoc), adds nothing and reviewers reject it, because it is just a restatement of the p-value you already have. If you are asked to justify the sample of a study already done, the right move is to reason about the minimum detectable effect, not to compute observed power. The detail of why post hoc power analysis does not work is in its own article.
The four ingredients
Any power calculation combines four quantities, and fixing three gives you the fourth. For sample size you fix the first three and solve for N:
- Expected effect size: how large the effect you expect is (a Cohen's d, an r, an f). It is the hardest ingredient and the one that moves N the most.
- Significance level (alpha): the false-positive risk, by convention .05.
- Power (1 minus beta): the probability of detecting the effect if it exists, by convention .80, and increasingly .90 in serious studies.
- The test and design: a t-test does not need the same N as a repeated-measures ANOVA or a correlation.
The hard step: where the effect size comes from
This is where 90% of the errors happen. The expected effect size should come, in order of preference, from prior studies on your exact question, a meta-analysis of the field, or your own pilot study. Only as a last resort do you use Cohen's conventions (small, medium, large), and using them blindly is dangerous: assuming a "medium" effect when your phenomenon produces small ones leaves you with a sample that falls short. The serious recommendation is to justify the effect you consider minimally important for your question, not the one that gives you the most comfortable N.
G*Power step by step
G*Power is free and the standard reviewers expect. The flow is always the same:
- Choose the test family (t, F, chi-square, z, exact) and the specific statistical test (for example, "Means: difference between two independent means" for a t-test).
- Set the type of analysis to "A priori: compute required sample size".
- Enter the expected effect size, alpha (.05), power (.80 or .90) and, where relevant, the allocation ratio between groups.
- Press Calculate and G*Power returns the total N and the N per group.
As a mental benchmark, to detect a medium difference between two groups (d = 0.5) with alpha .05 and power .80 you need about 64 per group (128 total); for a small effect (d = 0.2) the figure jumps above 390 per group. That sensitivity to the effect size is exactly why you cannot improvise it.
Special cases G*Power does not cover well
- SEM and factor analysis: not done in G*Power. Use rules of thumb (a common minimum of 200 cases, or 10 to 20 per estimated parameter) or, better, Monte Carlo simulation for the power of a specific model.
- Logistic regression: the practical rule is 10 to 20 events per predictor variable (EPV), not total N.
- Multilevel models: what matters is the number of clusters more than the number of individuals; estimated by simulation.
- Questionnaire validation: COSMIN guidelines ask for a minimum of 100 and 7 per item for the factor structure.
- Clinical trials with non-inferiority, equivalence or cluster designs: the formula changes from the standard superiority test. The clinical trial sample size calculator covers non-inferiority, equivalence (TOST), cluster designs and diagnostic accuracy, with dropout adjustment.
In all of these the number does not come out of a formula but out of the model you intend to fit, so it is worth showing both at once. Paste the project into the Methodologist and it tells you whether the sample size you are working with supports that particular model, along with the design problems worth fixing before you start recruiting.
Not only power: precision too
There is an increasingly recommended alternative: size the sample for a target precision, that is, a confidence interval narrow enough around the effect, rather than only to reach significance. It is especially useful in estimation studies (prevalence, validation, clinical studies). The logic is in the confidence intervals guide.
Common errors
Using Cohen's conventions without justifying them; computing post hoc power; forgetting the attrition rate (if you expect to lose 20%, inflate N so you end up with what you need); using the wrong test in the calculation (N depends on the final test); and not documenting the assumptions of the calculation, which is exactly what the committee wants to see. The sample size belongs inside the analysis plan, written and justified before you start, and if you are unsure which test you will run, the guide on how to choose the statistical test helps. And since "why this sample size?" is one of the safest bets among committee questions, if yours is an undergraduate or master's dissertation you can rehearse the answer in the defense simulator before they ask it for real.
Need to justify your sample size to a committee?
I am a PhD in Psychology and I calculate and justify your study's sample size (G*Power, simulation for SEM or multilevel, precision for clinical studies) with the text ready for the committee or the reviewer. In psychology or any health-sciences field. Free initial diagnosis.
See the statistical consulting →