Questionnaire Validation by Experts: Aiken's V, CVR and Delphi (Step by Step)

Before administering a questionnaire to hundreds of participants, before computing an alpha, and before fitting a factor model, there is a step that determines much of the instrument's quality and is often dismissed in two lines: content validity, that is, the judgment of an expert panel on whether the items adequately represent what you say you are measuring. It is the first validity evidence a thesis committee or a health-sciences reviewer looks for, the cheapest to get right and the most expensive to fix late. If you only need to run the numbers, go straight to the Aiken's V and CVI calculator.

Content validity vs face validity

Content validity is the degree to which the items of an instrument are a representative and relevant sample of the construct being measured. Do not confuse it with face validity, which is merely the surface impression that the questionnaire "looks like" it measures the right thing. Content validity is a reasoned judgment by experts about relevance, clarity and sufficiency against an operational definition of the construct. It comes first because no subsequent analysis can rescue an item that should not be there.

How many experts and how to select them

Standard practice is 5 to 10 experts. Below 5, the indices are unstable; above 10, marginal returns fall while logistics grow. An important nuance: Lawshe's CVR has critical-value tables starting at 5 judges, and CVI thresholds assume 6 or more. If you can only recruit 3 or 4 experts, you can still report the indices, but the criterion becomes more stringent (close to perfect agreement).

The panel should combine two profiles: content experts (clinicians or researchers in the area) and measurement experts (psychometricians or methodologists). Choose them for real experience (publications, years of practice, prior work with similar instruments), not availability. They should be independent of each other, and any conflicts of interest should be declared. Provide them with the construct definition and the definition of each dimension: an expert cannot judge relevance without knowing what they are judging against.

What the experts rate

The cleanest approach is to ask each judge to rate each item on several criteria using an ordinal scale (typically 4 points, to avoid a comfortable midpoint):

  • Relevance: the item is pertinent to the construct or the dimension to which it is assigned.
  • Clarity: the item is unambiguous and well written.
  • Coherence: the item fits the theoretical dimension where you placed it.
  • Sufficiency: rated at the dimension level (not item level), asking whether the set of items is sufficient to cover that dimension.

Always accompany numeric ratings with a space for qualitative comments. A large part of the value of expert judgment lies in wording suggestions and missing items, not only in the scores.

Quantifying agreement 1: Aiken's V

Aiken's V summarizes expert agreement on an item as a number between 0 and 1. The formula is V = S / [n (c - 1)], where S is the sum of expert scores after subtracting the minimum possible scale value, n is the number of judges, and c is the number of response categories. A value of 1 means all judges gave the maximum score.

The most widely used threshold for retaining an item is V ≥ .70, with many clinical-instrument committees requiring V ≥ .80. The serious recommendation is not to stop at the point estimate: report the confidence interval of V (via the Penfield and Giacobbi method), because with few judges the lower bound is what really tells you whether agreement is solid. The Aiken's V calculator returns V with its confidence interval and also the CVI.

Quantifying agreement 2: the CVI (I-CVI and S-CVI)

The Content Validity Index starts from a 4-point relevance scale and dichotomizes: it counts how many experts rate the item as relevant (3 or 4). The I-CVI (item level) is that proportion. The S-CVI/Ave (scale level) is the mean of the I-CVIs across all items.

Reference thresholds (Polit and Beck): with 6 to 10 experts, I-CVI ≥ .78; with 3 to 5 experts, I-CVI = 1.00 is required; for the scale as a whole, S-CVI/Ave ≥ .90. A known limitation of the CVI: it does not correct agreement for chance, so it should be accompanied by a modified kappa when the number of judges is low.

Quantifying agreement 3: Lawshe's CVR

The Content Validity Ratio is used when you ask experts to classify each item as "essential", "useful but not essential" or "not necessary". The formula is CVR = (ne - N/2) / (N/2), where ne is the number of experts who mark "essential" and N is the total number of experts. It ranges from -1 to +1: if more than half considers it essential, CVR is positive.

The strength of CVR is that it has critical values by panel size. An item is retained if its CVR exceeds the minimum from the table:

Experts (N)Minimum CVR
5.99
8.75
10.62
15.49
20.42
40.29

The mean of the retained items' CVRs is the Lawshe CVI (not to be confused with the Content Validity Index above, which shares the abbreviation but has a different formula). Reporting the critical value table alongside your CVRs is one of those small touches that signal rigor to a reviewer.

The Delphi method: when you need consensus

The three tools above work with a single rating round. The Delphi method is for when you want to build consensus iteratively, typically because the construct is new or contested. It consists of successive anonymous rounds: in each round experts rate, receive a summary of the group's responses (without knowing who said what) and rate again, so positions converge without a dominant expert pulling others along.

Anonymity is the key (it prevents authority and bandwagon effects), and you must fix the stopping criterion in advance: an objective consensus level (for example, Kendall's W or an agreement percentage) or the stability of responses across rounds. Two or three rounds usually suffice.

How to decide which items stay, are revised, or are removed

The decision is never purely numerical. Combine the quantitative threshold with the qualitative comments:

  • Retained: passes the threshold (e.g. V ≥ .80 or CVR above the critical value) with no substantive comments.
  • Revised: relevance is high but clarity is low, or comments flag ambiguous wording. The item measures the right thing but is poorly written.
  • Removed: does not meet the relevance threshold, or several experts agree it is redundant or irrelevant to the construct.

Document every decision and its justification. That decision table is exactly what a committee or an editor wants to see.

Common errors

Too few or non-independent experts; stopping at face validity and calling it content validity; conflating content validity (expert judgment) with construct validity (assessed afterwards with data through factor analysis and correlations with other variables); setting retention criteria after seeing the results; ignoring qualitative comments; and reporting Aiken's V without its confidence interval.

What comes after expert judgment

Expert judgment closes the content phase, but validation does not end there. Next comes a pilot study with the target population and then the psychometric core: factor structure with exploratory and confirmatory factor analysis, reliability with alpha, omega and ICC, and measurement invariance if you compare groups. The complete route is in the questionnaire validation guide.

Validating a questionnaire for your thesis or article?

I am a PhD in Psychology and I guide questionnaire validation from start to finish (expert judgment, Aiken's V, factor structure, reliability and invariance), with the write-up ready for the committee or the reviewer. Psychology, nursing, medicine or physiotherapy. Free initial diagnosis.

See the statistical consulting →

Keep reading

All blog articles