“The authors should conduct a sensitivity analysis to demonstrate the robustness of their findings.” One sentence, no further detail, and six genuinely different analyses it could be asking for. Run the wrong one and you will have done real work that does not answer the comment, which is worse than doing nothing, because the reviewer now has to repeat themselves.
The term is used differently across fields. In epidemiology it usually means quantifying how much unmeasured confounding would be needed to overturn your result. In clinical trials it means checking whether the conclusion survives different handling of missing data and protocol deviations. In psychology it often means rerunning the model with and without outliers or covariates. In meta-analysis it means leave-one-out and publication bias diagnostics. In power analysis, confusingly, it means something else entirely: solving for the smallest detectable effect.
This page separates them, gives you the diagnostic questions that identify which one you were asked for, sets out how each is actually run, and covers the reporting convention that matters most: a sensitivity analysis is interpreted by the spread of the estimates, not by whether the p value stayed under .05 in every specification.
Work through these and identify yours before you open the data file. In most cases the surrounding comments in the same review make it obvious, because reviewers rarely raise a robustness concern out of nowhere.
The most common meaning in psychology. The reviewer wants to know whether your conclusion depends on decisions you made along the way that could reasonably have gone otherwise: which covariates entered the model, whether you log-transformed the skewed outcome, whether you treated the Likert scale as continuous or ordinal, where you set the cut-off on the screening instrument, whether you used listwise deletion.
The tell is that the reviewer has questioned a specific choice elsewhere in the review. If they asked why you adjusted for baseline severity, the sensitivity analysis they want is your model with and without that adjustment. Rerun the primary analysis under the alternative and report both estimates with intervals.
Standard in trials, longitudinal designs and any study with meaningful attrition. Your main analysis probably assumed data were missing at random, whether or not you said so, since that is the assumption underlying multiple imputation and maximum likelihood. The reviewer wants to know what happens if that assumption fails.
The analyses that answer it: complete case analysis alongside the imputed analysis, so the reader can see whether they agree; a pattern-mixture or delta-adjusted model in which imputed values for dropouts are shifted by a specified amount in the unfavorable direction, with the shift varied across a plausible range; and, at the limit, extreme case scenarios where all dropouts in one arm are assigned the worst outcome. Report the point at which the conclusion changes, because that is the informative quantity: “the treatment effect remained statistically significant until imputed post-treatment scores for dropouts in the intervention arm were shifted by more than 4 points, which exceeds the minimal clinically important difference”.
The reviewer suspects the finding rests on a handful of observations, often after seeing a scatterplot or a small N. Standard diagnostics answer this: Cook’s distance, DFBETAS, leverage, and the simple and persuasive approach of refitting the model with each observation removed in turn and plotting the resulting coefficient.
The reporting rule here is important. Do not delete influential cases and report only the cleaned analysis. Report both, state how many cases were influential, and describe what they were: a data entry error is a different matter from a genuine participant with an extreme but real score, and the second is not an outlier to be removed but a person your model does not fit well.
The epidemiological meaning, and increasingly common in health psychology and health services research. You have reported an adjusted association and the reviewer wants to know how strong an unmeasured confounder would have to be to explain it away.
The E-value is the accessible modern tool: it reports the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need with both the exposure and the outcome, above and beyond the measured covariates, to fully account for the observed effect. It is computed from your point estimate and interval, it needs no additional data, and it converts a vague worry into a number that can be compared against the strength of known confounders in your literature. If the E-value is 1.4 and the measured covariates in your field routinely show associations that strong, that is worth saying honestly.
Common in multilevel, longitudinal and structural equation modeling. Does the conclusion survive a random intercept model versus random slopes, a different covariance structure for repeated measures, a different link function, robust standard errors, a different estimator such as MLR versus WLSMV, or in Bayesian work a different prior?
Prior sensitivity deserves a specific mention because reviewers of Bayesian analyses ask for it nearly every time. Report the posterior under your preferred prior and under at least two alternatives spanning a defensible range, including a wide, weakly informative one. Presenting a Bayes factor without a prior sensitivity check is now an incomplete report in most journals that publish Bayesian work.
If your paper is a systematic review, the reviewer means the standard suite: leave-one-out analysis, subgroup analysis by study quality or risk of bias, fixed versus random effects comparison, restricting to studies at low risk of bias, and publication bias assessment through funnel plot asymmetry, Egger’s test, trim-and-fill or selection models. Reviewers in this area have a checklist in their heads, so omitting any of them generates a comment.
Three questions usually resolve it. First, what did the same reviewer complain about elsewhere in the review? A sensitivity request almost always follows a specific doubt raised two comments earlier, and reading the review as a whole rather than comment by comment is the single most reliable diagnostic.
Second, what discipline is the reviewer from? A comment that also mentions confounding, adjustment or directed acyclic graphs comes from an epidemiological training and means the fourth meaning. A comment that mentions researcher degrees of freedom, preregistration or forking paths means the first. A comment appearing next to a query about dropout means the second.
Third, what is the weakest defensible decision in your own analysis? You usually know. If you excluded 14 participants on a rule you invented after seeing the data, if you chose a cut-off because it gave cleaner groups, if you dropped a covariate because it made the effect disappear, the sensitivity analysis the reviewer wants is almost certainly about that, whether or not they identified it.
And if none of this resolves it, you are allowed to ask. A short note to the editor requesting clarification of a comment is normal and does not count against you. It is considerably better than spending three weeks on the wrong analysis.
When many defensible analytic decisions exist and no single one is clearly correct, the systematic option is to run all reasonable combinations and report the distribution of results. In the multiverse framing the emphasis is on data processing choices, in the specification curve framing on model specification, but the logic is the same: enumerate the defensible choices, run the full grid, and display the resulting estimates ordered by size with the choices marked underneath.
It is the right answer when the decisions are numerous and genuinely arbitrary, when the field disagrees about them, or when your effect is modest and a reviewer suspects it depends on the path taken. It is overkill when your design has two or three decision points, when one specification is clearly preregistered and the others are alternatives, or when the paper is short. Running a 480-specification multiverse for a comment about whether you should have adjusted for age is not a proportionate response, and reviewers occasionally read it as an attempt to bury the question.
Two cautions if you do run one. Every specification in the grid must be defensible in advance; padding with implausible ones dilutes the result and is detectable. And the interpretation must engage with the whole distribution, including the specifications where the effect vanishes, rather than reporting the median and declaring robustness.
The reporting conventions here are specific enough that getting them wrong produces a second round even when the analysis was correct.
This happens, and how you handle it determines whether the paper survives. The temptation is to bury the discrepancy or to argue that the primary specification is the correct one and the others are irrelevant. Neither works, because you now hold information about the fragility of your finding and concealing it is the beginning of a much worse problem.
The publishable response is to report the discrepancy prominently, explain which decision drives it, and downgrade the claim to match. A finding that holds under six specifications and disappears under three, all defensible, is a finding whose abstract should say it is sensitive to how the outcome was operationalized. Journals do publish that. What they do not publish, and what retracts later, is the version where the fragility was known and unreported.
Two things belong in the answer: what you ran, and what it changed. The second is what reviewers actually read for.
A model response: “We have added a set of sensitivity analyses reported in Supplementary Table S4 and summarized on p. 16. The primary model is unchanged. We refitted it (a) without the 11 participants excluded for incomplete diary data, using multiple imputation instead; (b) treating the outcome as ordinal with a cumulative link model rather than as continuous; (c) with and without adjustment for baseline symptom severity, which the reviewer questioned; and (d) with cluster-robust standard errors at the clinic level. The adjusted coefficient ranged from 0.18 to 0.26 across the four, with confidence intervals excluding zero in three of the four; the exception was the unadjusted model, where the interval was [-0.02, 0.31]. We now state in the Results that the association is attenuated and no longer statistically distinguishable from zero without adjustment for baseline severity, and we have qualified the corresponding sentence in the Abstract. We have also added an E-value of 1.71 for the adjusted estimate.”
That answer volunteers the specification where the result weakened. Doing so is counterintuitive and it is the reason the response works: a reviewer who sees the authors reporting against their own interest stops looking for what else might be hidden, which is the actual mechanism by which robustness comments escalate.
In practice the terms are used interchangeably, though sensitivity analysis has a stricter meaning in epidemiology and trials, where it refers to varying a specific assumption such as the missing data mechanism or the presence of unmeasured confounding. Robustness check is the looser term for rerunning the analysis under alternative reasonable specifications. If a reviewer uses either, describe precisely what you varied rather than relying on the label.
As many as there are decisions a competent colleague could reasonably have made differently, and no more. For most papers that is between three and six. The number matters less than the selection: analyses that test the decisions the reviewer actually doubts, and the ones you privately consider weakest, are worth more than a long list of variations on choices nobody disputes.
Report it and change the conclusion. A finding sensitive to a defensible analytic choice is still publishable if you say so, identify which choice drives it, and qualify the abstract accordingly. Concealing the discrepancy is the response that risks the paper, and increasingly it is discoverable, because more journals require data and analysis code to be shared.
A summary of how strong an unmeasured confounder would have to be, in association with both the exposure and the outcome and beyond the covariates already adjusted for, to fully explain away an observed effect. It is computed directly from your estimate and interval with no extra data, and it turns an unfalsifiable objection about confounding into a number readers can compare against the strength of known confounders in the field. Report it for the point estimate and for the interval limit closest to the null.
No. It is the right tool when many defensible choices exist and none is clearly privileged, and it is disproportionate when the analysis has two or three decision points or one preregistered specification. A large grid can also obscure the specific question the reviewer asked. Targeted checks that address the doubts raised in the review are usually more persuasive than a large automated grid that addresses none of them directly.
The detailed table belongs in the supplement, and a two- or three-sentence summary belongs in the main text next to the primary result, including the range of estimates and any specification where the conclusion changed. Reviewers object when the only trace in the main text is a pointer to an appendix, because the robustness of a headline claim is part of the claim.
Yes. A short note to the handling editor asking for clarification of a comment is a normal part of the process and carries no penalty. It is far better than delivering a well-executed analysis of the wrong question, which costs you a review cycle and leaves the original doubt unanswered.