The reviewer wants measurement invariance. What do you actually run?

You compared men and women on a wellbeing scale, or three countries, or the same participants before and after treatment, and a reviewer has asked whether measurement invariance was established. If the answer is no, the objection is more serious than it sounds: without invariance, a difference in scale scores between your groups could reflect a difference in the construct, a difference in how the groups use the items, or any mixture of the two, and nothing in your comparison distinguishes them.

The concrete worry is easy to state. Suppose an item asking about crying is a strong indicator of depression in one group and a weaker one in another because of norms about emotional expression. The two groups can then differ on the total score while being identical on depression, or be identical on the total score while differing on depression. Every mean comparison, every correlation with an external variable and every regression coefficient you report is contaminated by this, and no amount of reliability evidence detects it, because coefficient alpha is computed within a group and says nothing about comparability across groups.

This page walks the nested sequence of models, states plainly what each level licenses you to claim, gives the fit criteria reviewers actually apply and where they break down, covers the ordered-categorical case that most Likert data really falls into, and sets out the options when invariance fails, which is a common and publishable outcome rather than a dead end.

The four nested models, and what each one buys you

Invariance testing fits an increasingly constrained series of confirmatory factor models across groups, and compares each to the one before. The constraints are imposed on different parts of the measurement model, and each level licenses a different kind of comparison. Knowing which level you need for your specific claim is what keeps this from being a ritual.

Configural: the same structure in both groups

The baseline model fits your factor structure simultaneously in all groups with no cross-group constraints: same number of factors, same items loading on the same factors, everything else free to differ. If this model fits acceptably, the construct is organized the same way in both groups, which is the minimum condition for going further.

A configural model that fits badly is a real finding and stops the sequence. It means the instrument measures something structurally different in your groups, and comparing scores is not meaningful in any form. This happens more often with translated instruments and with clinical versus community samples than authors expect.

Metric (weak): equal loadings

Loadings are constrained equal across groups, which tests whether a one-unit change in the latent variable produces the same change in each item in every group. Establishing metric invariance means the construct has the same meaning and the same metric in both groups.

What this licenses: comparing correlations and unstandardized regression coefficients involving the latent variable across groups. If your research question is whether the association between burnout and turnover intention differs between nurses and physicians, metric invariance is what you need, and you do not need to go further. Many papers test scalar invariance unnecessarily and then panic when it fails, for a claim that never required it.

Scalar (strong): equal intercepts or thresholds

Intercepts are additionally constrained equal, testing whether people at the same level of the latent trait give the same expected item score regardless of group. Failure here means one group systematically endorses an item more than the other at equal levels of the underlying construct, which is precisely the differential item functioning that biases mean comparisons.

What this licenses, and only this: comparing latent means, and by extension interpreting differences in observed scale scores or sum scores as differences in the construct. Every paper that reports a t test on a total scale score between two groups is implicitly assuming scalar invariance. This is why reviewers raise it, and why the objection applies to studies that never fitted a factor model at all.

Strict (residual): equal residual variances

Residual variances are additionally constrained, testing whether the items are measured with equal precision across groups. This level is required if you want to compare observed variances or reliabilities across groups, and it is rarely necessary for the questions psychology and health papers usually ask. Many methodologists consider it unnecessary for latent mean comparison. Report it if you have it and do not treat its failure as a problem unless your claim depends on it.

The criteria reviewers apply, and where they break down

The chi-square difference test between nested models is the formal test, and almost nobody relies on it alone, because it is sensitive to sample size in a way that makes it reject trivially small differences in large samples and miss substantial ones in small samples. Reviewers expect to see it reported and then supplemented.

The conventional supplementary criteria are a change in CFI no larger than .01 between nested models, accompanied by a change in RMSEA no larger than .015 or a change in SRMR no larger than .030 for loading invariance and .010 for intercept invariance. These cut-offs come from simulation work and have become the de facto standard in applied papers.

Two limitations are worth knowing, because a sophisticated reviewer will raise them. The cut-offs were derived under specific conditions, including reasonably large and balanced group sizes, continuous indicators and a particular range of model sizes, and their performance degrades outside those conditions, particularly with unequal group sizes and small samples. And they are rules of thumb, not tests, so a change of .011 is not qualitatively different from .009. Report the actual values rather than only a pass or fail verdict, and let the reader see how close the calls were.

One further practical point: report the fit of each model in the sequence, not just the differences. A table with model, chi-square, df, CFI, RMSEA with its interval, SRMR, and then the deltas against the previous model, is the expected format and it is what a reviewer will look for. Reporting only “measurement invariance was established” with no table is treated as an unsupported assertion.

Likert items are ordinal, and that changes the sequence

Most invariance testing in psychology is run with maximum likelihood on items with five or seven response options treated as continuous. With five or more categories and roughly symmetric distributions this is often tolerable, and robust maximum likelihood with a scaled chi-square handles the non-normality reasonably. With fewer categories, with strong floor or ceiling effects, or with clinical symptom items where most people select the lowest option, it is not.

The alternative is to treat items as ordered categorical and estimate with a weighted least squares mean and variance adjusted estimator, which models thresholds rather than intercepts. The invariance sequence then changes shape: because thresholds and loadings are not separately identified in the usual way, the recommended approach constrains thresholds and loadings together rather than in the two-step loadings-then-intercepts order familiar from the continuous case. Getting this wrong produces a sequence that looks conventional and tests the wrong constraints, and it is a specific thing psychometrically trained reviewers check.

Practically: in R, the semTools function that builds invariance syntax handles the ordinal parameterization correctly and is worth using rather than writing constraints by hand. In Mplus the categorical option with the appropriate model constraints does the same. jamovi and SPSS Amos are more limited here, and if your indicators are ordinal with few categories, this is a reason to move the analysis to lavaan or Mplus rather than to argue the point.

When invariance fails, which is often

Failure at the scalar step is the most common outcome in cross-cultural and cross-language comparisons, and treating it as a catastrophe leads authors to hide it. It is a result, it is informative, and there are four legitimate paths forward.

The one thing not to do

Do not run the sequence, find that scalar invariance fails, and then report the mean comparison anyway with a sentence in the limitations. This is the version reviewers reject hardest, because the analysis you ran established that the comparison is not interpretable and you made it regardless. If you cannot establish at least partial scalar invariance and your claim requires it, the claim has to change, not the caveat.

Longitudinal invariance: the version people forget

The same logic applies across time within the same people, and it is asked about less often than group invariance but matters just as much. If you claim that symptoms decreased from pre-treatment to follow-up, you are assuming the instrument measured the same construct in the same metric at both time points. Response shift, where participants recalibrate their internal standard after an intervention, is a real phenomenon in health outcomes research and is exactly what longitudinal invariance testing detects.

Two technical differences from the multi-group case. The model is fitted as a single-group model with the same items at each occasion, so constraints are imposed within one model rather than across groups. And the residuals of the same item at different occasions must be allowed to correlate, because item-specific variance persists in the same person over time; omitting these correlated uniquenesses biases the results and is one of the most common errors in longitudinal invariance work.

If your design has three or more waves, test the sequence across all of them jointly rather than pairwise, and report the same table format. For intensive longitudinal designs with many occasions, full invariance testing becomes impractical and the honest approach is to test it on a subset of occasions and say so.

Reporting it, and answering the reviewer

The expected report is compact: a paragraph in the analysis section naming the estimator, the software and version, the sequence you will fit and the criteria you will apply, stated before the results; a table with the fit of every model and the deltas; and a sentence in the results stating the level of invariance achieved and, if partial, exactly which parameters were freed.

Where authors lose marks is in stating a conclusion without the table, in reporting only chi-square difference tests, in claiming invariance from a single-group CFA, and in reporting the sequence for a scale while making comparisons on a different scale that was never tested. That last one is common in papers with several instruments: reviewers check whether the invariance evidence covers the measure the key claim rests on.

A response letter that closes the comment: “We agree that the group comparison in Table 3 assumes measurement invariance, which we had not tested. We have added a multi-group confirmatory factor analysis using WLSMV, since the items are five-point ordinal with pronounced floor effects. Configural fit was acceptable (CFI = .968, RMSEA = .044). Constraining loadings and thresholds simultaneously produced ΔCFI = .004 and ΔRMSEA = .003, supporting invariance for 14 of the 16 items. Items 7 and 12 required freeing; both refer to somatic complaints and the difference is consistent with the reporting differences described in [reference]. Latent mean comparisons were therefore conducted under partial invariance and are reported in the new Table 4; the direction and approximate magnitude of the group difference are unchanged, and we now state the partial invariance and its implications explicitly in the Results and Discussion.”

If you cannot run the analysis because the sample is too small for a multi-group model, say that instead of running it badly. Group sizes below roughly 100 to 150 per group make invariance testing unstable for models of typical size, and a reviewer will accept a stated constraint with a corresponding limitation more readily than a set of fit indices from an underidentified model.

Frequently asked questions

Do I need measurement invariance to compare group means?

Yes, at the scalar level, at least partially. Comparing observed scale scores between groups assumes that people at the same level of the underlying construct give the same expected item responses regardless of group, and that is exactly what scalar invariance tests. A t test on a total score is making this assumption whether or not a factor model was ever fitted, which is why the objection applies to papers with no structural equation modeling in them.

What level of invariance do I need to compare correlations across groups?

Metric invariance is sufficient. Equal loadings mean the latent variable has the same metric in both groups, which is what makes an unstandardized association comparable. Many papers test scalar invariance for a question that only requires metric, and then treat scalar failure as a barrier when it is not relevant to the claim being made. Decide what your claim needs before you run the sequence.

What are the accepted cut-offs for invariance testing?

The widely applied convention is a change in CFI no greater than .010 between nested models, together with a change in RMSEA no greater than .015 or a change in SRMR no greater than .030 for loadings and .010 for intercepts. Chi-square difference tests should be reported but not relied on alone, because they are highly sensitive to sample size. Report the actual values so readers can see how close the decision was rather than only a verdict.

What is partial invariance and how much is acceptable?

Partial invariance means some parameters are constrained equal across groups and the non-invariant ones are freed. Latent mean comparison remains defensible when at least two indicators per factor, including the reference indicator, are invariant, and when the freed items are a minority. Report which items were freed and offer a substantive explanation, because a pattern that makes sense (translation, cultural idiom, service availability) is far more convincing than a list of item numbers.

Is Cronbach’s alpha in both groups enough evidence?

No, and this is a common misunderstanding. Alpha is computed within a group and describes internal consistency there. Two groups can have identical alphas while the items function completely differently across them, and can have different alphas while being fully invariant. Reliability evidence and invariance evidence answer different questions, and reviewers who ask for the second will not accept the first.

My sample is too small for multi-group CFA. What are my options?

Say so explicitly and state the constraint rather than fitting an unstable model. Group sizes below roughly 100 to 150 make invariance testing unreliable for models of typical size. Reasonable alternatives include restricting the claim to within-group associations, using an item-response-theory-based differential item functioning approach designed for smaller samples, or citing existing invariance evidence for your instrument in comparable populations while acknowledging it was not established in yours.

Does invariance apply to comparisons over time as well as between groups?

Yes, and it is asked about less often than it should be. Claiming that scores changed after an intervention assumes the instrument measured the same construct in the same metric before and after, which response shift can violate. Longitudinal invariance is tested within a single model with constraints across occasions, and the residuals of the same item at different occasions must be allowed to correlate; omitting those correlated uniquenesses is the most frequent error in this analysis.

How does it compare to the other tools?

You may also find this useful