Traditional meta-analysis, also called pairwise meta-analysis, compares two interventions at a time: a treatment against a control, one drug against another, or a therapy against a waitlist. This structure works well when the clinical question is binary, but in real practice clinicians and health policy decision-makers face more complex questions. A clinical psychologist treating anxiety disorders does not simply need to know whether cognitive behavioral therapy (CBT) is better than placebo; they need to know how CBT compares with selective serotonin reuptake inhibitors (SSRIs), with mindfulness, with combined therapy, and with every other available alternative. Network meta-analysis (NMA), also known as mixed-treatment comparison meta-analysis, allows us to answer this type of question.
The concept of indirect comparison
The central idea of NMA rests on an apparently simple logical principle: if study A compares CBT versus placebo and study B compares SSRIs versus placebo, we can indirectly compare CBT with SSRIs through their common comparator, placebo. This idea was formalized mathematically by Bucher et al. (1997) and subsequently developed in more sophisticated statistical frameworks by Lu and Ades (2004, 2006) and Salanti et al. (2008, 2011).
For an indirect comparison to be valid, the transitivity assumption must hold: the studies comparing A versus C and B versus C must be sufficiently similar in all relevant characteristics (population, duration, dosage, outcome measures) so that the difference between A and B can be estimated in an unbiased way through C. If the studies comparing CBT versus placebo were conducted with mild patients and those comparing SSRIs versus placebo were conducted with severe patients, the indirect comparison will be confounded by patient severity.
When both direct comparisons (studies that directly compare A versus B) and indirect ones (estimates of A versus B through C) exist, NMA combines both sources of evidence into a single estimate. This increases the precision of the estimate and, in principle, allows the detection of inconsistencies: if the direct and indirect evidence do not agree, something is wrong, whether it is the transitivity assumption, the presence of effect modifiers, or quality problems in the studies.
The evidence network
The most characteristic visual component of an NMA is the network graph. In this graph, each node represents an intervention and each edge (line connecting two nodes) represents the existence of at least one study that directly compares those two interventions. Node size is typically proportional to the total number of participants assigned to that intervention, and edge thickness is proportional to the number of studies or participants in that direct comparison. This graph allows one to assess at a glance the structure of the evidence: which comparisons have abundant direct evidence, which rely on a single study, and which can only be estimated indirectly.
A well-connected network is essential for the validity of the NMA. Each intervention must be connected to the rest of the network, directly or indirectly, through at least one common comparator. Star-shaped networks (where all comparisons pass through a central node, such as placebo) depend entirely on indirect evidence for comparisons between active treatments. Denser networks, with multiple loops, allow the assessment of consistency between direct and indirect evidence.
The consistency assumption
The consistency assumption states that the direct estimate of A versus B and the indirect estimate (through C) should agree. When they do not, we speak of inconsistency, which is the statistical manifestation of a violation of the transitivity assumption. Inconsistency can be assessed globally (Higgins design-based test) or locally, comparison by comparison (node-splitting or back-calculation methods). Detecting inconsistency is crucial because it invalidates the NMA estimates for the affected comparisons.
Interpreting results: league tables and rankings
NMA results are typically presented in three ways. The first is the forest plot, which shows all comparisons against a reference treatment (for example, all interventions versus placebo). The second is the league table, a triangular matrix where each cell contains the effect estimate and its confidence or credibility interval for the comparison between two treatments. This table allows one to evaluate any pairwise comparison at a single glance.
The third is the treatment ranking, which orders interventions from best to worst. SUCRA (Surface Under the Cumulative Ranking curve) is the most widely used ranking index: a SUCRA of 100% indicates that a treatment is always the best across all simulations, and a SUCRA of 0% indicates that it is always the worst. P-scores (frequentist) are analogous to SUCRA in the Bayesian framework.
However, rankings must be interpreted with extreme caution. A treatment may have the highest SUCRA and yet not be statistically different from several other treatments. Rankings exaggerate differences that may not be clinically significant and are especially unstable when confidence intervals are wide. Cialdini et al. (2022) and Salanti et al. have repeatedly warned that SUCRA should not be used as the sole criterion for recommending a treatment.
When is an NMA appropriate
An NMA is appropriate when several conditions are met. First, there must be a connected network of interventions: each treatment must be connected to the rest, directly or indirectly, through at least one common comparator. Second, the included studies must be sufficiently similar in their clinical and methodological characteristics for the transitivity assumption to be reasonable. Third, there must be enough studies to estimate effects with acceptable precision; with one or two studies per comparison, estimates will be very imprecise.
Situations where an NMA is not appropriate include disconnected networks (where no path connects all interventions), insufficient evidence (few direct comparisons), extreme clinical heterogeneity between studies from different comparisons, or when interventions are not genuinely comparable (for example, comparing surgery with psychotherapy for chronic pain without clinical justification).
Software and reporting guidelines
NMAs can be conducted within a frequentist or Bayesian framework. In R, the netmeta package implements the frequentist approach (graph-theoretical model of Rucker and Schwarzer), while gemtc implements the Bayesian approach via JAGS. In Stata, the network command supports both approaches. Bayesian models can also be fitted directly in WinBUGS or OpenBUGS using code provided by the NICE Decision Support Unit.
For reporting, the PRISMA-NMA extension (Hutton et al., 2015) provides a specific checklist that includes elements such as the presentation of the network graph, assessment of transitivity, consistency tests, and presentation of rankings with their credibility intervals. Following these guidelines not only improves reporting transparency but also facilitates quality assessment by reviewers and readers, which is essential when the goal is to publish in Q1 journals.
In conclusion, NMA is a powerful extension of traditional meta-analysis that allows the simultaneous comparison of multiple interventions and generation of treatment hierarchies. Its value lies in the ability to integrate direct and indirect evidence to inform complex clinical decisions. However, its assumptions are more demanding than those of pairwise meta-analysis, and the interpretation of rankings requires prudence. Like any advanced statistical tool, it is only as reliable as the quality of the data and the reasonableness of the assumptions on which it is built.