Everything in your study came from one questionnaire, filled in by one person, in one sitting. A reviewer writes that common method variance may account for the observed relationships, and asks what you did about it. If you have read a few papers in your field you know the standard reply: run Harman’s single factor test, report that the first factor explains less than 50% of the variance, conclude that common method bias is not a concern.
That reply has not worked for some years, and increasingly it makes things worse. Harman’s test has been criticized in the methodological literature for two decades and is now explicitly rejected by editors at several organizational, health and psychology journals, because it lacks the sensitivity to detect method bias at the levels that actually distort estimates. Reporting it signals that you reached for the ritual rather than the argument.
This page covers what common method variance is and is not, why the popular statistical remedies do less than their reputation suggests, what the procedural design features are that genuinely reduce it, the counter-arguments you are entitled to make when the criticism is applied reflexively, and how to write a response when the data are already collected and no new design is possible.
Common method variance is variance attributable to the measurement method rather than to the constructs the measures represent. When two variables are measured by the same method, anything that systematically influences responses to both, independent of the constructs, becomes shared variance that the correlation cannot distinguish from a substantive relationship.
The sources are well catalogued. Some sit in the rater: consistency motifs, where respondents try to appear coherent across their answers; implicit theories about how the constructs relate; social desirability; transient mood at the moment of completion; acquiescence, where some people agree with everything. Some sit in the items: ambiguous wording that invites the respondent to fall back on a general attitude, common scale anchors, item priming from proximity, complexity that pushes people toward heuristic answering. Some sit in the context: all measures on one page, all in the same order for everyone, all in a single sitting with no break.
The important qualification, which authors under criticism should know and reviewers sometimes forget, is that method variance does not act uniformly. It inflates some correlations and attenuates others, depending on how the method factor relates to each measure. It is most damaging to bivariate associations between constructs measured with similar-looking items, and it behaves differently for interaction terms, which is the basis of one of the strongest defensive arguments available and is covered further down.
The test as usually run loads all items from all measures into an unrotated exploratory factor analysis and checks whether a single factor accounts for the majority of the variance. If it does not, authors declare the problem absent.
The logic does not hold. The criterion is arbitrary: nothing establishes that method bias severe enough to distort a correlation would produce a first factor above 50%. The test is insensitive in practice, and simulation work has shown it failing to detect method variance at levels that meaningfully bias parameter estimates. It also has a structural absurdity, in that a study with more distinct constructs will almost automatically pass, since more genuine factors mechanically reduce the share taken by the first, meaning the test rewards exactly the designs where more method contamination is possible.
The confirmatory version, fitting a one-factor model and showing it fits badly, has the same problem in different clothes. Bad fit for a single factor demonstrates that your constructs are distinguishable, which is a discriminant validity result, not evidence about method bias. Reviewers who make this distinction, and there are more of them every year, will say so.
If you have already reported Harman’s test in the submitted manuscript, the best move in revision is usually to remove it rather than defend it, and to replace it with something that engages the actual question.
Three post hoc approaches are commonly requested. All of them are better than Harman’s test, and none of them is a solution.
You add a latent factor on which every indicator in the model loads, alongside its substantive factor, and compare parameter estimates with and without it. The appeal is that it needs no extra measurement, so it can be done to data already collected.
The difficulties are technical and serious. The models are frequently underidentified or converge only with constraints that carry their own assumptions, such as equal method loadings across all items. The method factor absorbs any misspecification in the model, not only method variance, so a change in the substantive estimates is ambiguous evidence. And the approach cannot separate method variance from a genuine higher-order construct that your measures share. Reviewers with structural equation modeling experience know all of this. If you fit one, report the identification constraints explicitly, report fit for both models, and describe the result as a robustness check rather than as a correction.
You include a variable that is theoretically unrelated to everything in your model, measured with the same method. Any correlation it shows with your substantive variables is attributed to shared method, and that correlation is used to adjust the others, in the simplest version by partialling the smallest observed correlation out of the matrix.
This is a stronger idea, and the confirmatory variants that model the marker as a latent factor within the measurement model are stronger still, because they allow the method effect to differ across constructs instead of assuming it is constant. The catch is decisive for authors under review: the marker has to have been measured, and chosen for theoretical irrelevance before you saw the data. You cannot select it afterwards by looking for the variable with the lowest correlations, which is what post hoc marker analyses effectively do and what reviewers ask about first. If you have a defensible marker in your battery, use it. If you do not, say you do not rather than improvising one.
Frequently the honest position is that no post hoc statistical remedy is available for the design you ran, and that stating this plainly is stronger than performing a test that does not work. What you can still do is bound the problem: report the correlation matrix in full so readers can see the pattern, report the discriminant validity evidence you have, such as heterotrait-monotrait ratios or the comparison of average variance extracted against squared inter-construct correlations, and reason explicitly about which specific associations in your model are most at risk and which are structurally protected.
The methodological consensus is that procedural remedies do the work and statistical remedies mop up. If you are still designing, these are the features that change the size of the problem, and if you have already collected data, some of them may already be present in your study and simply undescribed. Read the list as an audit of what you can claim.
Many manuscripts that receive this comment already contain two or three of these features and never described them, because they seemed like housekeeping. Anonymity guaranteed, order randomized, one outcome taken from records, a two-week gap between sections of the protocol. Going back through the procedure and writing a short paragraph titled procedural steps to reduce method effects is frequently the single most effective revision available, and it costs no new data.
Common method variance is sometimes raised as an automatic objection to any self-report study, and there is a serious methodological literature arguing that its effects have been overstated and that the criticism is often applied without evidence that bias is present in the particular case. You may use these arguments, provided you use them precisely and do not overreach.
The strongest one concerns interactions. Method variance shared between predictor and moderator tends to attenuate rather than inflate the estimate of their product term, which means a significant interaction is difficult to explain as an artifact of shared method and is more plausibly an underestimate of the true effect. If your key finding is a moderation, say this explicitly and cite the work establishing it. It is the cleanest defense available in this area.
A second is differentiation. If method bias were driving your results, it should inflate associations across the board, so a pattern in which some theoretically expected correlations are strong and others, measured identically, are near zero is not what an undifferentiated method factor would produce. Point to the specific null associations in your own matrix; they are evidence, and they are already in your paper.
A third is construct type. The bias is most credible for evaluatively similar attitudinal constructs measured with similar item wording. It is far less credible for factual or behavioral reports with a defined reference period, for measures with distinct response formats, or where one construct is affectively neutral. Argue this for your specific measures rather than in general.
What none of these justifies is the sentence “common method bias was not a concern in this study”. The defensible sentence names the mechanism, explains why it is unlikely to account for the particular association at issue, and concedes the associations where it remains a live alternative explanation.
The structure that works has four moves: acknowledge without ritual, describe the procedural features already present, report whatever legitimate statistical check you can run while naming its limits, and then qualify the specific claims that remain vulnerable.
A worked version: “We agree that all constructs were self-reported in a single session and that common method variance is a plausible alternative explanation for part of the observed covariation. We have removed the Harman single factor test from the revised manuscript, since it lacks the sensitivity to support the conclusion we drew from it. In its place we have added a paragraph to the Procedure (p. 8) describing the steps taken at design stage: anonymity was explicitly guaranteed, the presentation order of the three instrument blocks was randomized across participants, response formats differed across blocks, and the sickness absence outcome was obtained from personnel records rather than self-report. We also note in the Results that the central finding is a cross-level interaction, and that shared method variance attenuates product terms rather than inflating them, so method bias is an unlikely account of that effect. For the two bivariate associations that do rest entirely on same-source data, we have qualified the claims in the Discussion (p. 21) and identified them as requiring multi-source replication.”
That response concedes what cannot be defended, defends what can be defended with a reason rather than an assertion, and leaves the reviewer with something to verify. It also removes a test rather than adding one, which is unusual enough that reviewers tend to read the rest of the letter more sympathetically.
It is still widely published and widely criticized, and an increasing number of editors and reviewers reject it outright. The test lacks the sensitivity to detect method variance at levels that bias estimates, its 50% threshold has no principled basis, and studies with more constructs pass it more easily regardless of contamination. If you have already reported it, the strongest revision is usually to remove it and replace it with a description of procedural safeguards.
A marker is a variable measured with the same method as your constructs but theoretically unrelated to them, so its correlations with them estimate shared method variance. It has to be planned: choosing one after the fact by looking for the lowest correlations in your matrix converts the technique into a data-driven exercise, and that is the first thing a reviewer will check. If a defensible marker exists in your battery, use it and explain why it is theoretically unrelated. If not, say so.
No, and this matters for your defense. Method variance inflates some associations and attenuates others depending on how the method factor relates to each measure. It systematically attenuates estimates of interaction terms, which is why a significant moderation effect is hard to explain as a method artifact. Reviewers who state that method bias necessarily inflates everything are overstating the case, and you can say so with citations.
Not necessarily, though temporal separation is one of the most effective remedies. What the comment requires is either a design feature that breaks the shared source or moment, an argument for why the specific association at issue is unlikely to be an artifact, or an honest qualification of the claim. Adding a second wave is the strongest answer and often impossible; the other two are available immediately.
It is disproportionately valuable. A single outcome taken from records, from a different rater or from an objective source removes the shared-rater mechanism for the associations involving it, and lets you contrast those associations with the fully self-reported ones inside your own data. That contrast is evidence a reviewer can inspect, which is worth far more than any post hoc statistical test you could run on same-source data.
Social desirability is one source of common method variance, not a synonym for it. Others include acquiescence, consistency motifs, implicit theories about how the constructs relate, transient mood, item priming and shared scale anchors. Controlling statistically for a social desirability scale addresses only one channel and is often taken by authors as addressing the whole problem, which reviewers will point out.
Only with a stated reason attached. A bare assertion that method bias was not a concern is the version reviewers reject. A defensible version identifies the mechanism, explains why it is implausible for your specific constructs and design, points to features of your own results that an undifferentiated method factor would not produce, and concedes the specific associations where it remains a live alternative.