In health sciences, a questionnaire is rarely a formality: a quality-of-life scale, a PROM (patient-reported outcome measure) or an adherence instrument ends up informing clinical decisions and measuring whether a treatment works. That is why validating them well is not optional, and why journals in the field expect validation to follow a recognized standard. In psychology the reference is the Standards (AERA, APA, NCME); in health, the de facto standard is COSMIN.
This guide covers the measurement properties COSMIN requires and the order in which to evaluate them, with indicative sample sizes for each step. It applies to nursing, medicine, physiotherapy, dentistry or any health discipline that uses questionnaires. If your instrument is in pure psychology, the questionnaire validation guide for psychology fits better.
What COSMIN is and what properties it evaluates
COSMIN (COnsensus-based Standards for the selection of health Measurement INstruments) is a set of internationally consensus-built standards for evaluating and reporting the quality of health measurement instruments. It organizes properties into three domains:
- Reliability: internal consistency, test-retest reliability, inter-rater reliability, and measurement error.
- Validity: content validity, structural validity (factor structure), construct validity (hypothesis testing), cross-cultural validity (invariance) and criterion validity.
- Responsiveness: the ability of the instrument to detect real change over time.
The underlying logic is the same as in general psychometrics: validity is not a fixed property of the questionnaire, but of the interpretations you make of its scores in a specific population and use. A PROM validated for chronic pain is not, by default, validated for post-surgical pain.
Step 1: content validity (the most important and most neglected)
Before any numbers, a panel of experts and a patient sample must judge whether the items are relevant, comprehensible and sufficient for the construct. In COSMIN, content validity is considered the most important property and the worst reported. Quantify it with Aiken's V or Lawshe's CVR; the full procedure is in the questionnaire validation by experts guide, and the indices can be calculated with the Aiken's V and CVI calculator. If you are adapting an instrument from another language, add the cross-cultural adaptation process of the International Test Commission (translation, back-translation, expert committee and cognitive interviews with patients).
Step 2: structural validity (factor analysis)
Structural validity checks that the instrument's dimensional structure holds up with data. If the instrument is new or being adapted, the ideal is an exploratory factor analysis in one sample and a confirmatory factor analysis in an independent one. COSMIN considers a minimum of 7 subjects per item and at least 100 people adequate for this step; below that, the factor solution is unstable.
Step 3: reliability (not only Cronbach's alpha)
This is the most common error in health: reporting only Cronbach's alpha and leaving it at that. Internal consistency (alpha, and preferably McDonald's omega) is one piece, but COSMIN also distinguishes test-retest reliability, which for continuous scores is evaluated with the intraclass correlation coefficient (ICC) in a sample of stable patients measured twice (a typical interval is 1 to 2 weeks). Also report the measurement error (SEM and the smallest detectable change, SDC), because that is what allows you to interpret whether an individual change is real or noise.
Step 4: construct validity through hypothesis testing
Rather than claiming generically that the instrument "is valid", COSMIN asks you to formulate a priori hypotheses about how scores should behave and test them. For example: expecting high correlations with instruments of the same construct (convergent validity), low correlations with different constructs (discriminant validity) and differences between known clinical groups. Defining hypotheses before looking at the data is what distinguishes a serious validation from a correlation fishing expedition.
Step 5: cross-cultural validity and measurement invariance
If you are comparing countries, languages or clinical groups, you need evidence of measurement invariance (configural, metric and scalar). Without invariance you cannot interpret a mean difference between groups as a real difference in the construct: it could be an artefact of the instrument. This step is omitted with surprising frequency in cross-cultural health studies.
Step 6: responsiveness and interpretability
A PROM you intend to use for evaluating a treatment must detect changes when they occur. Responsiveness is also evaluated with a priori hypotheses (for example, that the clinically improving group shows greater change on the scale than the non-improving group). And for change to be interpretable you need the Minimal Important Change (MIC or MID): how many points on the scale correspond to a change that the patient notices. Without that, a "statistically significant" change may be clinically irrelevant.
Summary table: properties, how assessed, and indicative sample size
| Property | How assessed | Indicative sample |
|---|---|---|
| Content validity | Experts + patients (Aiken's V, CVR) | 5 to 10 experts |
| Structural validity | EFA / CFA | ≥ 100 and ≥ 7 per item |
| Internal consistency | Alpha, omega | ≥ 100 |
| Test-retest reliability | ICC | ≥ 50, two measurements |
| Construct validity | Hypotheses (convergent, known-groups) | ≥ 75 |
| Responsiveness | Hypotheses about change + MIC | ≥ 50 with change |
Common errors in health validation
Reporting only alpha and calling it "validation"; skipping content validity; not measuring test-retest reliability with ICC; comparing groups without checking invariance; and failing to calculate measurement error and the MIC, which is exactly what a clinician needs to interpret scores. Another classic is confusing internal consistency with reliability: a high alpha does not demonstrate temporal stability or unidimensionality.
How it all fits together
A complete COSMIN validation is, in practice, a publishable study in its own right. You do not need to evaluate every property in a single paper, but you should state which properties you cover and which remain pending. If the instrument measures a psychological construct, most of the analysis (structure, reliability, invariance) overlaps with validation in psychology; what changes in health is the emphasis on test-retest reliability, measurement error and responsiveness.
Validating a PROM or a clinical scale for your thesis or article?
I am a PhD in Psychology and I run COSMIN validations from start to finish (content validity, factor structure, reliability and ICC, invariance and responsiveness), with R code and the write-up ready for the journal. I work with nursing, medicine, physiotherapy and dentistry. Free initial diagnosis.
See the statistical consulting →