Exploratory Factor Analysis (EFA): Complete Guide with Examples

Exploratory factor analysis (EFA) is one of the most widely used multivariate techniques in psychological research, and with good reason. When a researcher develops a new questionnaire or adapts an existing one for a different population, they need, after evaluating item content validity with tools such as the content validity calculatora tool that allows them to discover the underlying structure of the responses, that is, to identify how many dimensions or factors lie behind a set of items. EFA fulfills exactly that function: it analyzes the correlations among observed variables to identify groupings that suggest the existence of latent constructs not measured directly.

However, the apparent simplicity of running an EFA in any statistical software conceals a series of methodological decisions that can drastically influence the results. The choice of extraction method, the criterion for determining the number of factors, the type of rotation, and the interpretation of factor loadings are steps where inadequate decisions can lead to artificial or difficult-to-replicate factor solutions. Understanding the logic behind each of these decisions is what separates a rigorous EFA from a mechanical exercise of clicking through a software menu.

If what you need at the end of this process is the results paragraph and the table already drafted in APA 7 style, the EFA APA 7 report generator builds them from your loading, KMO, Bartlett and variance-explained values.

Pattern matrix · Oblique rotation (Promax) · n = 412 Shading proportional to |λ| · loadings < .30 attenuated F1 Anxiety F2 Depression i1 · I feel tense .78 .05 i2 · I worry about everything .82 −.02 i3 · Racing heartbeat .71 .12 i4 · Hard to concentrate .48 .42 cross-load i5 · I no longer enjoy things .07 .83 i6 · I feel sad .10 .79 i7 · Trouble sleeping −.04 .74 i8 · Lack of energy .06 .76 α = .85 · ω = .87 α = .82 · ω = .84 Factor correlation: r(F1, F2) = .42
Pattern matrix after oblique rotation. Simple structure on 7 of 8 items; i4 cross-loads and deserves a review (drop it, reword it, or accept it if theory supports it).

Before you begin: prerequisites

Before running an EFA, it is necessary to verify that the data meet certain conditions. The first is sample size. Although the classic rule of "at least 10 participants per item" is still frequently cited, recent methodological research has demonstrated that what is relevant is not the subject-to-item ratio but rather the communalities of the items and the clarity of the factor structure. With high communalities (above .60) and well-defined factors (at least 3-4 items with high loadings per factor), samples of 100-150 participants may be sufficient. With low communalities and a diffuse structure, not even 500 participants guarantee a stable solution. In practice, the recommendation of aiming for a minimum of 200-300 participants remains sensible as a starting point.

The adequacy of the correlation matrix is evaluated using two classic tests: Bartlett's test of sphericity and the Kaiser-Meyer-Olkin (KMO) measure of sampling adequacy. Bartlett's test tests the null hypothesis that the correlation matrix is an identity matrix, that is, that the variables are not correlated with each other. If this hypothesis is not rejected, the EFA is pointless because there are no correlations to explain. The KMO, for its part, evaluates the extent to which partial correlations between variables are small, which indicates that the observed correlations are due to common factors. KMO values above .80 are considered good, and values below .60 suggest that the EFA may not be appropriate for those data.

Another aspect worth evaluating is the nature of the variables. If the items have an ordinal response format with few categories (4- or 5-point Likert scales, for example), using the Pearson correlation matrix can underestimate the true associations between items. In these cases, it is preferable to work with the polychoric correlation matrix, which estimates the correlations as if the underlying variables were continuous. Programs such as R (with the psych package) or FACTOR allow this option, while SPSS only works with Pearson correlations natively.

Key decisions: extraction and number of factors

The extraction method determines how the factors are estimated from the correlation matrix. The two methods most recommended in the current literature are principal axis factoring and ordinary or weighted least squares. Principal component analysis (PCA), although frequently confused with EFA, is not technically a factor analysis method because it does not distinguish between common variance and specific variance. PCA seeks components that maximize total explained variance, while EFA seeks factors that explain the shared variance among the variables. This difference, apparently technical, has practical consequences: PCA tends to overestimate factor loadings, especially when communalities are low.

Determining how many factors to retain is probably the most critical decision in the entire process, and unfortunately it is where the most errors are made. Kaiser's criterion (retaining factors with eigenvalues greater than 1) remains the most commonly used default in most statistical software, but methodological research has repeatedly demonstrated that it tends to overestimate the number of factors, especially with many variables. The scree plot is a visual alternative that looks for the "elbow" where the eigenvalue curve levels off, but its interpretation is subjective and can vary between analysts.

The currently most recommended methods are Horn's parallel analysis and Velicer's MAP (Minimum Average Partial) method. Parallel analysis compares the eigenvalues obtained from the real data with the eigenvalues expected from random data matrices of the same size. Only factors whose real eigenvalues exceed the random ones are retained. This method has consistently demonstrated greater accuracy than Kaiser's criterion in simulation studies. The MAP method, for its part, seeks the number of components that minimizes the average partial correlation between variables after extracting the factors. The ideal approach is to triangulate several methods and look for convergence, rather than blindly relying on a single one.

Rotation and interpretation

Once the factors are extracted, rotation facilitates their interpretation by redistributing the explained variance among them. The main decision here is whether to allow the factors to be correlated with each other (oblique rotation) or to force their independence (orthogonal rotation). In psychology, where the measured constructs are usually related to each other (anxiety and depression, for example, or the different dimensions of personality), oblique rotation is generally more appropriate. The most widely used oblique rotations are Promax and Direct Oblimin, while Varimax is the most common orthogonal rotation.

A pragmatic strategy that many methodologists recommend is to always start with an oblique rotation. If the correlations between factors turn out to be negligible (below .15 or .20 in absolute value), the oblique solution will converge with the orthogonal one and the latter can be reported for simplicity. But if there are substantial correlations between factors, forcing an orthogonal rotation would distort the solution and conceal real relationships between the measured dimensions.

The interpretation of the rotated solution is based on factor loadings, which indicate the strength of the relationship between each item and each factor. Conventionally, loadings above .30 or .40 are considered significant, although the exact threshold depends on the sample size and the context. Items with high loadings on a single factor and low loadings on the rest are the easiest to interpret. Items with cross-loadings (substantial loadings on more than one factor) may indicate that the item is ambiguous or that it measures aspects of several constructs simultaneously. The decision to eliminate these items or keep them depends on theoretical and practical considerations that go beyond statistics.

Reporting the results correctly

A complete EFA report should include several elements. First, the justification of the decisions made: why that extraction method was chosen, what criteria were used to determine the number of factors (and the results of each criterion), and what type of rotation was applied and why. Second, the measures of sampling adequacy (KMO and Bartlett). Third, the total variance explained by the factor solution. Fourth, the table of rotated factor loadings, clearly indicating which items load on each factor. And fifth, the communalities of each item, which report what proportion of its variance is explained by the extracted factors.

An aspect that is often neglected is the need for replication. An EFA performed on a single sample can produce a solution specific to those data that does not hold up in independent samples. The recommended practice is to split the sample in half (if the size permits), perform the EFA on one half, and verify the structure through a confirmatory factor analysis on the other. When this is not possible due to sample limitations, at least it is important to be transparent about the exploratory nature of the results and the need for subsequent confirmation. The EFA is a tool for generating hypotheses about the structure of an instrument, not for confirming them. This distinction, fundamental but frequently ignored, marks the difference between an appropriate use and a misuse of the technique.

Reporting EFA results in APA 7 style

Reviewers in psychology and the health sciences expect the results of an exploratory factor analysis to follow APA 7th edition conventions, and a reporting section that omits the expected elements is one of the most common reasons for a revise-and-resubmit. A complete APA-style EFA report states, in this order: the extraction method and rotation used; the Kaiser–Meyer–Olkin (KMO) measure and Bartlett's test of sphericity; the criteria used to decide the number of factors and what each one indicated; the total variance explained; and a table of factor loadings. Statistical symbols such as χ², p, r and are italicized; KMO and the factor labels are not.

A results paragraph following these conventions reads, for the two-factor solution shown above: "The eight items were submitted to an exploratory factor analysis using principal axis factoring with Promax (oblique) rotation. The Kaiser–Meyer–Olkin measure indicated adequate sampling, KMO = .89, and Bartlett's test of sphericity, χ²(28) = 1,742.6, p < .001, confirmed that the correlations were large enough for factor analysis. Parallel analysis and Velicer's MAP test converged on two factors, which together explained 58.4% of the variance. Item i4 cross-loaded on both factors (.48 and .42) and was flagged for revision. The factors, labelled Anxiety and Depression, were moderately correlated, r = .42."

The loading matrix itself is presented as a table: the table number (e.g., Table 1) appears above an italicized, title-case caption such as Pattern Matrix of the Eight Items After Promax Rotation; items occupy the left-hand stub column and the factors form the remaining columns. A general note below the table reports the extraction and rotation methods and the threshold below which loadings were suppressed (commonly .30). Boldfacing the salient loading for each item and adding a final column with the communalities () makes the table far easier to read. After reporting the structure, it is good practice to include each factor's reliability; our Cronbach's alpha calculator returns the coefficient with its confidence interval. For a copy-paste template built specifically for this analysis, see how to report an exploratory factor analysis in APA 7, with a ready-made results paragraph and loading table; our general APA 7 reporting guide walks through the other tests, and the questionnaire validation service covers the full EFA-to-CFA pipeline.

Keep reading

All blog articles