Almost every psychology and health paper you have ever read used a convenience sample. Undergraduates from a subject pool, patients who happened to attend one clinic during one calendar year, nurses from three wards in one hospital, 400 respondents from Prolific. Probability sampling from a defined population is rare outside epidemiology and national surveys, and reviewers know that. So when a reviewer writes “the study used a convenience sample, which limits generalizability”, they are almost never objecting to the fact itself.
They are objecting to something more specific, and if you answer the wrong objection you lose the exchange. Sometimes the complaint is that your claim is bigger than your sample can carry. Sometimes it is that the selection mechanism plausibly correlates with your outcome, which is a bias problem, not a generalizability problem. Sometimes it is that you wrote one sentence about it in the limitations and nothing else, which reads as a box being ticked.
This page separates those objections, shows what a convenience sample actually forbids and what it still permits, and gives you the material to write a limitation paragraph that a reviewer accepts: the recruitment funnel, the comparison against a benchmark, and a stated direction of likely bias instead of the ritual sentence about caution in interpretation.
A convenience sample is any sample selected because its members were reachable, rather than by a mechanism that gives every member of a defined population a known probability of inclusion. That definition covers the subject pool, the clinic caseload, the online panel, the snowball referral chain and the conference attendee list. It does not, by itself, invalidate anything.
The reason it draws fire is that it breaks a specific chain of reasoning. Statistical inference asks what a population value is, given a sample drawn from that population by a known process. Remove the known process and the sampling distribution you are using to compute your confidence interval no longer describes the population you want to talk about. The p value still computes; what it refers to becomes ambiguous. Reviewers rarely put it in those words, but that is the machinery underneath the complaint.
In practice the objection arrives in one of three forms, and they need different answers. The first is a coverage complaint: your sample cannot represent the population named in your title, so the title has to change. The second is a selection bias complaint: the reason people ended up in your sample is itself related to your outcome, so the association is distorted, not merely narrow. The third is a transparency complaint: you did not report enough about recruitment for anyone to judge either of the first two. Read the reviewer’s sentence again and decide which one you got.
Coverage is about who is missing. If you sampled psychology undergraduates and want to talk about adults, you are missing almost everyone, and the fix is to narrow the claim. This is uncomfortable but not fatal, and it is the objection most authors already know how to answer.
Selection is about why people are present, and it is the more serious one. If you recruited for a study advertised as being about social anxiety, people high in social anxiety who are also comfortable enough to volunteer are overrepresented, and the association between anxiety and your outcome can be biased in a direction you can reason about. If you recruited from a clinic waiting list, everyone in your sample crossed a help-seeking threshold, which conditions on a variable that sits between your predictor and your outcome. Reviewers who understand this are not asking you to caveat; they are asking you to think about the direction of the distortion and say so.
The single most useful correction to make in your own head is that random assignment and random sampling are different things and buy different things. Random assignment buys internal validity: it licenses a causal claim about the effect of your manipulation within the people you studied. Random sampling buys external validity: it licenses generalizing a number to a population. A randomized experiment on a convenience sample supports a causal claim and does not support a population estimate. A national probability survey supports a population estimate and does not support a causal claim. Reviewers conflate these constantly, and so do authors.
That distinction tells you which claims survive in your paper and which do not.
Reviewers treat these very differently, and a limitation paragraph written for one of them sounds hollow attached to another.
The classic objection is age, education and cultural narrowness, and the classic defense is that the process under study is basic enough to be invariant. That defense is only credible when you can name the process and point to evidence it replicates outside the pool. Where it fails badly is anything involving life experience, occupational context, clinical severity or accumulated exposure: work stress, parenting, chronic illness adaptation, retirement, financial strain.
The version of the defense that actually works is comparative. Report the pool’s composition, note where it plausibly differs from the target population on the constructs in your model, and reason about what that does to your estimate. Restricted range on the predictor attenuates the correlation, which means your estimate is a conservative floor rather than an inflated one. That is a real argument. “Future research should replicate in other populations” is not.
These are more demographically diverse than a subject pool and less naive than either. Reviewers in 2026 expect you to address data quality directly, not to treat panel data as a solved problem. The specific things they look for: attention and instructional manipulation checks stated in advance with a preregistered exclusion rule, screening for automated responses and server farm traffic, completion time floors, duplicate IP or worker ID handling, and how many people you excluded and why, reported as a funnel rather than as a final number.
Non-naivety is the underdiscussed one. Panel workers have often seen the Cognitive Reflection Test, the trolley problem, the ultimatum game and several standard deception paradigms many times. If your design relies on surprise or on an unfamiliar cover story, a reviewer is entitled to ask what proportion of your sample had seen something like it before. Ask, and report it.
The payment and consent conditions also matter to health journals and ethics reviewers in a way they do not to psychology journals. If your sample includes people reporting clinical symptoms, describe the risk protocol: what a participant who endorses suicidal ideation on your screener sees on the next page.
One hospital, three schools, a single company, the patients treated in a service between two dates. This is the strongest kind of convenience sample and the one authors defend worst, because the defense here is nearly always available and nearly always omitted.
If you sampled a whole service across a defined window, you did not sample conveniently in the loose sense. You took a consecutive series, which is a defined and describable process, and you should say so in exactly those words: consecutive patients attending X between date and date, with these inclusion criteria, of whom this many were approached, this many consented and this many completed. That is a sampling frame. It has a coverage limitation, which is that it describes one service, but it has no selection mystery, and it lets you compare your sample to the service’s full caseload on the variables the records contain.
A limitation paragraph is convincing when it contains information the reviewer did not have. Four pieces of information do most of the work, and most manuscripts contain none of them.
Post-stratification weighting to census margins is available and is sometimes requested, particularly in health services research. It fixes coverage on the variables you weight on and nothing else, so it helps when the imbalance is demographic and the outcome relates to demographics. It does nothing about selection on unmeasured variables, it inflates your standard errors, and it can look like a technical answer to a conceptual problem. If you do it, report both weighted and unweighted estimates and treat the comparison as a sensitivity analysis rather than a correction.
Here is the version that appears in most submitted manuscripts: “A limitation of this study is the use of a convenience sample, which limits the generalizability of the findings. Future research should replicate these results in more representative samples.” It contains no information. Every reviewer has read it several hundred times, and its main effect is to signal that the authors have not thought about their sampling at all.
Here is the same limitation, rewritten with the four elements above, for a hypothetical study of burnout and sleep quality in hospital nurses: “Participants were nurses working in three medical-surgical units of a single tertiary hospital, invited during shift handover between March and June. Of 214 eligible staff, 168 were reached, 141 consented and 129 provided complete data (60.3% of those eligible). Compared with the hospital’s staff register, our sample was similar in age and years of experience but underrepresented night-shift-only staff (11.6% versus 19.4%), who are both harder to reach at handover and, in prior work, report worse sleep. This is likely to have attenuated the burnout-sleep association we report, so the coefficient should be read as a conservative estimate. Because the sample comes from one hospital with a single rostering system, the prevalence figures in Table 2 describe this workforce and should not be read as national estimates.”
The second version does not have better data. It has the same data, described well enough that a reviewer can evaluate it, and it makes a directional claim the authors are prepared to defend. That is the whole difference, and it is usually the difference between a comment that reappears in round two and one that does not.
When the comment arrives after review rather than before, you have three legitimate moves and one illegitimate one. The illegitimate one is adding a sentence to the limitations and reporting that you have addressed the comment, which is the most common response and the one that produces a second round.
The first legitimate move is to narrow the claim. Change the title, the abstract’s final sentence and the first paragraph of the discussion so they refer to the population you actually sampled. This costs you nothing scientifically and is often exactly what the reviewer wanted; many convenience sample comments are really title comments in disguise.
The second is to add the evidence you left out: the funnel, the benchmark comparison, the sensitivity analysis with and without weighting. If you have administrative data on non-participants, even on two variables, a non-response comparison is disproportionately persuasive.
The third is to disagree, which is legitimate when the reviewer has conflated design with sampling. If your causal claim rests on random assignment, say so plainly: the convenience sample constrains how far the effect size generalizes but does not threaten the internal validity of the randomized comparison, and you have restricted your generalization claims accordingly. Cite the relevant point, acknowledge the constraint you are accepting, and show the reviewer where in the revised manuscript you accepted it. A disagreement that comes with a concession attached is usually accepted.
On its own, almost never in psychology and health sciences, because most published work in these fields uses one. It becomes a rejection reason when it is paired with a claim the sample cannot support, such as a prevalence estimate or a statement about a national population, or when recruitment is described so thinly that the reviewer cannot rule out serious selection bias. The sample is rarely the problem; the mismatch between the sample and the claim is.
Better on demographic breadth, worse on naivety, and roughly equivalent on the underlying inferential problem, since neither is a probability sample of anything. Online panels bring their own set of reviewer expectations around data quality: preregistered attention checks, bot screening, duplicate detection and a transparent exclusion funnel. A panel sample described without those details often draws more criticism than a subject pool sample described carefully.
Do not justify it there, describe it there. State the recruitment channel, the eligibility criteria, the dates, the numbers at each stage of the funnel and any incentive. The justification belongs in the discussion, where you connect the sample to the target population and reason about direction of bias. Methods sections that argue instead of describing make reviewers suspicious.
Yes, and everyone does. The honest framing is that your inference is about a hypothetical population that your sample can be treated as representing, or about the effect of a manipulation within the sample you have. What you cannot do is present the resulting numbers as estimates of parameters in a named real-world population without an argument connecting the two. Confidence intervals remain informative about sampling variability; they do not fix selection.
It solves the part of the problem that runs through the variables you weight on, and only that part. If your sample is 70% female and your population is 51%, weighting fixes the sex imbalance and does nothing about the fact that people who volunteer for a study on wellbeing differ from those who do not on wellbeing itself. Report weighted and unweighted results side by side and treat the difference as information about how sensitive your conclusion is to composition.
Where the participants came from, when, what made someone eligible, how many were approached, how many consented, how many completed and how many were excluded with reasons. Six numbers and two sentences. Their absence is the single most common cause of an escalating convenience-sample exchange with a reviewer, because without them the reviewer has to assume the worst.
Say so directly, explain the constraint, and offer what you can do instead: narrow the claim to the population you sampled, add a benchmark comparison, run the analysis with and without weights, and move any population-level statement to a clearly labelled speculation. Editors accept a well-argued refusal far more often than authors expect. What they do not accept is silence, or a change so small it is obviously cosmetic.