Causal claims in a cross-sectional study, and how to fix them

One reviewer comment appears more often than any other in cross-sectional psychology and health research: the design does not support the causal language used in the abstract and the discussion. It is easy to dismiss as pedantry, because you know your design, you wrote the limitation, and you never used the word “cause” anywhere in the paper.

That is exactly why the comment keeps arriving. Causal language almost never enters a manuscript through the word “cause”. It enters through “predicts”, “leads to”, “impacts”, “drives”, “reduces”, “the effect of”, through a mediation model that assumes a temporal ordering the data cannot see, and through a practice implication that only makes sense if the association is causal. Authors write the limitation on page 24 and then contradict it in the title, the last line of the abstract and the first line of the discussion.

This page maps where the causal language actually hides, explains why cross-sectional mediation is the single most criticized case and why the bias there is not small, sets out what you are entitled to say, and gives you a rewriting table you can run over your own manuscript before a reviewer does it for you.

Where causal language actually enters the manuscript

Run this as a search-and-inspect pass rather than a reading pass. The offending words are few and they cluster in five predictable places.

The particular problem with “predictor”

The regression vocabulary works against you here. Statistically, “predictor” is a neutral term for a variable on the right-hand side of the equation, and there is nothing wrong with it in a methods section or a table header. In a discussion, however, most readers hear it temporally: predicting implies coming first. Reviewers vary in how much they care, and some will accept it while others will insist on “correlate” or “associated variable”. The safe practice is to keep “predictor” where the model is being described and to avoid it where the finding is being interpreted.

The same applies to “determinant”, which sounds technical but is causal, to “risk factor”, which epidemiologists reserve for prospectively established associations, and to “explains variance in”, which is arithmetically accurate and rhetorically causal. If your sentence would read strangely with the two variables swapped, it is asserting direction.

Why cross-sectional mediation draws the sharpest criticism

Of all the causal structures fitted to cross-sectional data, mediation is the one that gets papers rejected, and the reason is not stylistic. A mediation model is a claim about a process unfolding in time: X changes, which changes M, which changes Y. Every parameter in it is defined with reference to that sequence. When all three variables are measured in the same fifteen-minute survey, the model estimates a quantity that has no straightforward relationship to the longitudinal process it is meant to represent.

The methodological work on this is unambiguous and has been for close to two decades. Cross-sectional estimates of indirect effects are biased relative to the longitudinal parameters they are taken to stand for, the bias can be large, and its direction is not predictable from the cross-sectional data. Estimates can be substantially inflated, substantially attenuated, or the wrong sign, depending on the stability of the variables and the lag over which the true process operates. Crucially, the fit of the cross-sectional model tells you nothing about this: a mediation model with excellent fit indices can be badly wrong about the process.

A second problem compounds it. In cross-sectional data, the statistical model X to M to Y typically fits about as well as M to X to Y and as Y to M to X, because they can be equivalent models with identical covariance implications. Choosing among them is done by theory, not by the data, and reviewers know that. When authors present a fitted mediation model as if the data had selected it, they invite the reviewer to fit the reversed model and ask why theirs is worse.

None of this makes mediation analysis forbidden. It makes it a modeling exercise whose causal interpretation depends entirely on assumptions you have to state: temporal ordering assumed from theory, no unmeasured confounding of any of the three paths, no reverse causation, and correct specification of the lag. State them, and reviewers will usually let the analysis stand as exploratory. Present the indirect effect as a demonstrated mechanism and they will not.

What to do if the mediation analysis is the paper

If the indirect effect is your headline finding and the data are cross-sectional, you have three realistic options. First, keep the analysis and reframe it explicitly as a test of the pattern of associations implied by a theoretical model, reporting that the data are consistent with that pattern and equally consistent with alternatives you should name. Second, fit the plausible alternative orderings and report them side by side, which is more persuasive than it sounds because it shows you know the models are statistically indistinguishable and are not hiding it. Third, and best when possible, find any source of temporal separation in the design, even a partial one: outcomes measured a week later, register data on a subsequent event, a retrospective anchor with a defined reference period.

What does not work is the sentence “although the design is cross-sectional, the model was based on theory, and therefore the direction of the paths is justified”. Theory justifies specifying the model. It does not turn the estimate into an unbiased estimate of the causal quantity, and that is what the reviewer is objecting to.

Why adding covariates does not buy you causality

The most common repair authors attempt is to add control variables and then use firmer language on the grounds that confounders have been handled. It rarely satisfies a reviewer with an epidemiology or causal inference background, for three reasons worth understanding.

The first is straightforward: adjustment removes confounding only from the variables you measured, measured well, and specified with the right functional form. Measurement error in a covariate leaves residual confounding behind, which is why adjusting for “socioeconomic status” using a single self-reported income band controls for much less than the label suggests.

The second is that some adjustments actively create bias. Conditioning on a variable that lies on the causal path between exposure and outcome removes part of the effect you are trying to estimate. Conditioning on a common consequence of exposure and outcome, a collider, induces an association where none existed. The habit of putting every available variable into the model as a control is a reliable way to do both, and “we controlled for all available demographic and clinical variables” now reads to informed reviewers as a warning rather than a reassurance.

The third is temporal. In cross-sectional data you often cannot tell whether a covariate is a confounder, a mediator or a consequence, because everything was measured at once. Adjusting for current medication use when studying symptom severity is a clear case: medication may confound the association or may be a downstream consequence of severity, and the data cannot distinguish them. Say which you assume it is and why.

The honest framing is that adjustment reduces confounding under stated assumptions and does not license causal language. If you want to go further, sensitivity analysis for unmeasured confounding is the accepted tool: quantifying how strong an unmeasured confounder would have to be to explain away your association gives the reader something concrete and is increasingly expected in health journals.

The rewriting pass, sentence by sentence

Most of the fix is mechanical. These are the substitutions that carry the load, and running them across a manuscript typically takes an hour.

What you are entitled to claim, and it is more than nothing

Cross-sectional designs support several claims outright, and papers weaken themselves by hedging these too. You can report the magnitude and precision of an association, and if the association is large and precisely estimated, that is a finding. You can report the pattern of associations across several variables and evaluate whether it matches or contradicts what a theory predicts, which makes your study a test of the theory even without temporal data. You can rule things out: a theory predicting a strong association is embarrassed by a tight interval around zero, and disconfirmation is logically stronger than confirmation here.

You can also make legitimate causal claims from a cross-sectional survey when the design contains a randomized element. A vignette or survey experiment embedded in a single wave randomizes the manipulation, and the causal claim about that manipulation is clean, though everything else in the survey remains correlational. Instrumental variable and natural experiment designs applied to single-wave data are another route, with their own assumptions to defend. If your design has one of these, say so early and prominently, because reviewers apply the cross-sectional reflex to the whole paper unless you stop them.

The limitation paragraph, and answering the reviewer

The standard limitation, “the cross-sectional design precludes causal inference”, is necessary and insufficient. It is a disclaimer, and reviewers have learned that a disclaimer sitting next to causal verbs is worth nothing. What upgrades it is specificity about which alternative explanations are live in your particular case.

A stronger version names them: reverse causation with an argument about which direction is more plausible and why, the specific unmeasured confounders that would be most damaging given your variables, and the timescale problem, meaning whether the process you are theorizing about would even be visible in a single measurement. A limitation that says “rumination and insomnia plausibly maintain each other, and our design cannot separate the two directions; prospective work using daily diaries would be needed because the hypothesized process operates over days rather than months” is doing analytic work rather than covering the authors.

For the response letter, the effective structure is to concede the general point immediately, then show the audit rather than argue about it. Something like: “We agree, and on rereading we found that causal phrasing had survived in the title, in the final sentence of the abstract and in five places in the Discussion. We have revised all of them, listed with line numbers below. The mediation model is now presented as an exploratory test of a theoretically specified pattern rather than as evidence of a mechanism, and we have added the reverse-ordered model to the supplement to make the equivalence explicit. The practice implication has been rewritten as conditional on the association being causal, and we now identify the prospective design that would be needed to establish it.”

That response works because it gives the reviewer something to check. A letter that says “we have softened the language throughout” without saying where is the single most common cause of the comment returning in round two, usually with a sharper tone.

Frequently asked questions

Can I use the word “predictor” in a cross-sectional paper?

In the methods and in table headers, yes, because there it is standard regression vocabulary for a right-hand-side variable. In the discussion it is riskier, since most readers hear a temporal claim in it. Reviewers differ on how strictly they enforce this. The low-cost approach is to keep it where the model is being described and to use “correlate” or “associated with” where the finding is being interpreted.

Is mediation analysis with cross-sectional data always wrong?

Not wrong to compute, but wrong to interpret as evidence of a mechanism. Cross-sectional indirect effects are biased estimates of the longitudinal quantities they are meant to represent, and the bias can be large in either direction. The workable position is to present the analysis as a test of whether the pattern of associations is consistent with a theoretically specified model, to state the assumptions that would have to hold, and to acknowledge that alternative orderings fit equally well.

Does controlling for confounders let me make causal claims?

No. Adjustment removes confounding only for variables you measured well and specified correctly, and it can introduce bias when the covariate is a mediator or a collider. In cross-sectional data you frequently cannot tell which of the three a variable is. Adjustment is worth doing and worth reporting, but the honest sentence is that the association persisted after adjustment for the measured covariates, not that the effect is causal.

Does adding a limitation paragraph solve the problem?

Only if the rest of the manuscript agrees with it. The comment usually arrives precisely because the paper contains both a disclaimer and a set of causal claims, and reviewers read the contradiction as the authors trying to have it both ways. Fix the title, the abstract and the discussion first; the limitation paragraph is the last step, not the fix.

Can a cross-sectional study ever support a causal claim?

Yes, when the design carries something extra. A randomized manipulation embedded in a single-wave survey, such as a vignette or framing experiment, supports a causal claim about that manipulation. Instrumental variable designs and natural experiments can support causal claims under strong and stateable assumptions. In these cases make the design element visible early, because reviewers otherwise apply the correlational reflex to the whole paper.

What can I legitimately conclude from cross-sectional data?

The magnitude and precision of associations, whether the pattern of associations across variables matches or contradicts a theoretical prediction, and whether a proposed association can be ruled out within a stated interval. Disconfirmation is particularly strong: a theory that requires a substantial association is genuinely threatened by a tight interval around zero, and that is a real contribution rather than a consolation prize.

The reviewer wants longitudinal data I do not have. What do I say?

Say that you agree the design cannot answer the directional question, name the prospective design that could and the timescale it would need, and then show what you have done instead: reframed the claims, presented alternative orderings, and, if available, added a sensitivity analysis for unmeasured confounding. Editors accept a study that knows its own limits. What they reject is a study that states a limit and then ignores it two paragraphs later.

How does it compare to the other tools?

You may also find this useful