The reviewer wants a larger sample. What are your options?

Of all reviewer comments, “the sample is too small; the authors should collect more data” is the one that feels least answerable. The funding period closed. The clinic stopped referring. The cohort graduated. The ethics approval expired. You have a manuscript and no realistic route to more participants, and the reviewer has written a sentence that appears to require one.

It is answerable more often than it looks, but only if you first work out what the comment is really about. “Too small” can mean too small to detect the effect, too narrow to support the generalization in the title, too thin to sustain the number of parameters in your model, or too small in one particular subgroup that the paper leans on. Those are four different problems and only one of them has more participants as the only solution.

This page covers how to decode the demand, the four legitimate responses and when each applies, how to argue credibly from population scarcity when your participants genuinely do not exist in large numbers, and how to write a polite refusal that editors accept, including when to take the disagreement to the editor rather than to the reviewer.

First, decode what is actually being demanded

Reviewers rarely explain the reasoning behind a sample size complaint, so you have to reconstruct it. Four distinct concerns produce the same sentence, and the surrounding comments in the review usually tell you which one you have.

The question to ask yourself before writing anything

Which specific claim in this manuscript does the sample fail to support? Write the answer down. If you can name it, you know what to do: either support it differently or stop making it. If you cannot name any, then the comment may be a reflex applied to a number that looked small in isolation, and the correct response is a reasoned defense rather than a concession, which is discussed at the end of this page.

Option one: collect more, when it is genuinely feasible

Sometimes the answer is yes, and authors dismiss it too quickly. If your recruitment channel still exists, if the instrument battery is short, and if the additional N required is a few dozen rather than a few hundred, extending data collection during the revision window is possible more often than it appears. Online panel studies in particular can add participants in days.

Before committing, check four things. The revision deadline, and whether an extension is available, which editors grant routinely if you ask before the deadline rather than after. Whether your ethics approval covers the extension or needs an amendment, and how long that takes at your institution, which is usually the binding constraint. Whether the new participants are drawn under identical conditions, because a second cohort recruited a year later through a different channel is a different sample and creates a batch problem you will have to model and report. And whether the additional data would actually change the conclusion, which a quick sensitivity power calculation will tell you.

If you do extend, report it transparently: state that data collection was extended during revision, give the dates and numbers for each phase, and show that the results are similar in the original and combined samples. Reviewers are fine with this when it is declared and deeply unimpressed when they detect it from a changed N with no explanation.

Option two: shrink the claim to fit the data

This is the most underused response and often the best. A sample too small for the paper you wrote may be entirely adequate for a slightly different paper, and rewriting the claim costs you nothing but the ambition you had for the abstract.

The moves available, in rough order of how much they cost you: remove the moderation or subgroup analysis from the abstract and present it as exploratory in the results; narrow the population in the title from a general category to the specific one you sampled; reframe the study as preliminary, feasibility or proof of concept, which many journals have a formal article type for; reduce a confirmatory framing to a descriptive or estimation framing, reporting intervals and refraining from hypothesis tests you are not powered for; and, at the limit, change the target journal to one whose scope accommodates smaller-scale work.

The reframing that works best is the one you propose yourself rather than the one you are forced into. A response letter that says “we agree, and we have accordingly retitled the paper, moved the moderation analysis to the supplement and removed the corresponding claim from the abstract” reads as competence. The same change extracted over two rounds reads as resistance.

One caution: a pilot or feasibility framing has to be genuine. Feasibility studies have their own reporting conventions, including explicit feasibility outcomes such as recruitment rate, retention and acceptability, and a stated intention about the definitive study. Relabeling an underpowered efficacy study as a pilot without adopting any of that is a move reviewers recognize.

Option three: show the design was adequate for the primary question

If the complaint is about power and you believe the design was adequate for what you claimed, you can make that case, but it requires numbers rather than assertion.

The materials are a sensitivity power analysis for the primary test, stating the smallest effect the design could detect at 80% power; confidence intervals on the primary estimates so the reader can judge precision directly; and, where the finding is null, an equivalence test or a Bayes factor that quantifies evidence for absence rather than leaving it as a failure to reject. Together these convert “the sample seems small” into a specific statement about what your study could and could not see.

You can also argue from design efficiency where it applies, and this is legitimate and frequently forgotten. Within-subject designs need substantially fewer participants than between-subject designs for the same power. Repeated measurements per person, as in diary or experience sampling designs, buy precision at level 1 that a cross-sectional design of the same N does not have. Well-matched or stratified designs reduce error variance. If your 40 participants each provided 60 observations, the reviewer comparing you to a 300-person survey is comparing the wrong quantity, and saying so with the effective sample size for your key parameter is a strong reply.

Option four: fit a model the sample can support

When the concern is complexity rather than N, the productive answer is to simplify, and the simplification is usually available without abandoning the research question.

Common substitutions: replace a full latent variable model with a path analysis on composite scores, or with a model using item parcels, both of which drastically reduce the number of estimated parameters at a known cost in measurement precision that you should state. Replace random slopes with random intercepts when the number of clusters is too small to estimate slope variance, and report that you did so and why. Replace a three-way interaction with a planned contrast that tests the specific comparison your theory predicts, which is both more powerful and more interpretable. Replace a model with twelve covariates with one carrying the three that a directed acyclic graph identifies as necessary, which is better practice regardless of sample size. Use regularized or Bayesian estimation with weakly informative priors where a maximum likelihood solution is unstable, and report the priors.

The framing in the response letter matters. Presenting the simpler model as a concession to sample size is weaker than presenting it as the appropriate model given the information available, with the more complex version retained in the supplement for readers who want it. The second version is honest, it preserves your work, and it gives the reviewer somewhere to look.

When the participants genuinely do not exist: defending a small N

Some samples are small because the population is small. Rare neurological conditions, inpatient units for a specific diagnosis, elite athletes in one discipline, professionals with a particular specialization, survivors of a specific event. Reviewers apply general sample size heuristics to these studies routinely, and a well-made scarcity argument is usually accepted, but it has to be made with evidence rather than asserted.

What a credible scarcity argument contains: an estimate of the size of the eligible population with a source, whether that is a national registry, a prevalence figure or the caseload of the participating services; the recruitment numbers showing you approached a substantial proportion of it; the time period over which recruitment ran; and a comparison to the sample sizes in the comparable published literature, which in rare conditions is often smaller than yours. A sentence such as “the 27 participants represent 61% of the individuals meeting criteria who attended the three national referral centres during the 22-month recruitment window” ends the conversation in a way that no amount of methodological argument can.

Pair it with an analysis strategy appropriate to the N rather than one borrowed from large-sample research. Options that reviewers accept in this space include single-case experimental designs with proper visual and statistical analysis, Bayesian estimation with informative priors derived from the existing literature, within-subject designs maximizing observations per participant, and case series reported against the established reporting guidelines for that design. What does not work is running a conventional null hypothesis significance test on 14 people and discussing the p value as if the design were standard.

Qualitative and mixed methods studies

A sample size objection to a qualitative study is usually a category error, and it is worth answering as such rather than defensively. Qualitative sampling logic aims at analytic rather than statistical generalization, and the number of participants is judged against the richness of the data, the specificity of the sample to the research question, the quality of the dialogue and the analytic strategy.

Do not lean on data saturation as an automatic defense. It has become contested: it derives from a specific methodology and sits awkwardly with reflexive thematic analysis, whose developers have argued explicitly against using it as a generic justification, and reviewers increasingly notice when it is invoked as a formula. The stronger contemporary framing is information power, which reasons about how narrow the study aim is, how specific the sample is, how much established theory the analysis is supported by, how strong the dialogue was, and whether the analysis is a case analysis or a cross-case one. Argue on those grounds and name them.

Writing a refusal the editor will accept

You are allowed to decline to collect more data. Authors overestimate how badly this goes. What determines the outcome is not whether you comply but whether you engage: editors reject responses that dismiss the comment, and they routinely accept responses that explain the constraint, offer everything else available, and adjust the claims.

The structure: acknowledge the concern as legitimate; state the specific and verifiable reason more data cannot be collected; list what you have done instead; state the claims you have withdrawn or qualified as a result; and, if relevant, note what the study still contributes.

A version that works: “We appreciate the concern and agree that a larger sample would allow stronger inference. Additional recruitment is not feasible: the study enrolled consecutive patients admitted to the two specialist units in the region over 26 months, the funded recruitment period closed in November, and the eligible incident population is approximately 30 per year nationally, so an appreciably larger sample would require a multi-year multi-centre study. In place of additional data we have (a) added a sensitivity power analysis showing the design could detect within-subject effects of d = 0.48 or larger, (b) replaced the latent model with a path analysis on composites, retaining the original model in Supplement B, (c) removed the subgroup comparison in former Table 4 and the corresponding claim from the Abstract, and (d) added recruitment figures to the Methods showing the sample represents 58% of eligible admissions in the window. We now frame the study as providing the first estimates in this population rather than as a confirmatory test.”

If the reviewer is immovable and you believe they are wrong, write to the editor. Editors arbitrate; that is their function, and a courteous note explaining that one reviewer requires data that cannot exist and asking the editor to weigh the request against the study’s contribution is a normal part of the process. Do not argue with the reviewer through three rounds hoping they will change their mind. They are anonymous, unpaid, and under no obligation to.

Frequently asked questions

Can I refuse to collect more data and still get published?

Yes, frequently. Editors accept well-argued refusals when the constraint is real and stated specifically, when you offer the alternatives that are available, and when you adjust the claims the sample cannot support. What fails is a refusal without alternatives, or one that treats the comment as unreasonable. The concession that matters is usually about what you claim, not about how many participants you have.

How do I know whether the reviewer means power or generalizability?

Read the rest of the review. Power concerns travel with comments about null results, missing power analyses and wide intervals. Generalizability concerns travel with comments about single sites, single countries, one clinic or one profession. The distinction is decisive, because more participants from the same source addresses the first and does nothing at all for the second.

Is it acceptable to collect additional data during revision?

Yes, provided you declare it. State that data collection was extended during revision, give dates and numbers for each phase, confirm that the procedure and eligibility were identical, and show that the results are consistent in the original and combined samples. Check first whether your ethics approval covers the extension, since an amendment is often the longest step. Undeclared additions detected by a reviewer are a serious problem.

How do I defend a small sample in a rare condition?

With numbers about the population, not with methodological argument. Give an estimate of the eligible population and its source, the proportion of it you recruited, the recruitment window, and a comparison with sample sizes in the published literature on the same condition. Then use an analysis strategy suited to the design, such as within-subject measurement, single-case methods or Bayesian estimation with informative priors, rather than a conventional test that the N cannot sustain.

Should I relabel my study as a pilot?

Only if you are willing to write it as one. Pilot and feasibility studies have their own conventions: explicit feasibility outcomes such as recruitment and retention rates, acceptability data, progression criteria and a stated plan for the definitive study. Applying the label to an underpowered efficacy study without adopting any of that is transparent to reviewers and usually makes the situation worse rather than better.

A reviewer asked for a larger sample in a qualitative study. What do I say?

Explain that qualitative sampling follows a different logic, aimed at analytic rather than statistical generalization, and then justify your N on the appropriate grounds: how narrowly the aim is defined, how specific the sample is to that aim, how much established theory supports the analysis, the quality of the interview dialogue, and whether the analysis is case-based or cross-case. Avoid leaning on saturation as a formula, since it is contested and increasingly recognized as invoked rather than demonstrated.

When should I go to the editor instead of arguing with the reviewer?

When the request is impossible rather than difficult, when two reviewers contradict each other, or when you have already answered the point once and it has returned unchanged. Write a short, courteous note to the handling editor setting out the constraint and what you have done instead, and ask them to weigh it. Editors arbitrate as a matter of routine, and using that route is far more productive than a third round of the same exchange.

How does it compare to the other tools?

You may also find this useful