Network Meta-Analysis: Comparing Treatments That Were Never Directly Compared
Network meta-analysis can rank every treatment in a therapeutic area from a single analysis - even when most of them have never been tested head-to-head. The method is powerful, and the assumptions it rests on are frequently overlooked.
In most therapeutic areas, the number of approved treatments is large and the number of direct head-to-head randomised trials is small. Cardiologists prescribing anticoagulants, oncologists choosing between PD-1 inhibitors, and rheumatologists selecting between biologics all face the same problem: the evidence they need - a direct randomised comparison of the treatments they are actually choosing between - usually does not exist. Network meta-analysis (NMA) was developed to address this gap. It has become central to health technology assessment, and its use carries assumptions that are often stated briefly and assessed inadequately.
The Logic of Indirect Comparison
The basic intuition behind indirect comparison is straightforward. If treatment A has been compared to placebo in randomised trials, and treatment B has also been compared to placebo in randomised trials, then the relative effects of A versus B can be estimated by combining the two sets of comparisons through the common comparator (placebo). This is an indirect comparison. Its validity rests on one critical assumption: that the patients in the A-vs-placebo trials and the patients in the B-vs-placebo trials are sufficiently similar that the placebo arm can serve as a valid link between them. This assumption is called transitivity.
Transitivity is not a statistical property that can be tested directly - it is a clinical judgement about whether the studies connected through the network are comparable enough in terms of population, disease severity, concomitant treatments, and outcome definitions to support the mathematical borrowing that NMA performs.
How NMA Works Technically
In a network meta-analysis, all available pairwise comparisons are synthesised simultaneously in a coherent statistical model. Each treatment is assigned a relative effect versus a reference treatment, and these effects are estimated jointly, incorporating both direct evidence (where it exists) and indirect evidence (inferred through the network). The model produces a full set of pairwise comparisons and, from these, treatment rankings - often presented as "surface under the cumulative ranking curve" (SUCRA) or P-score values that summarise each treatment's probability of being best, second-best, and so on.
The statistical framework for NMA was formalised in a series of methodological papers in the early 2000s, and it is now implemented in standard statistical packages. Both frequentist and Bayesian approaches are used in practice, with Bayesian NMA offering natural incorporation of prior information and probabilistic ranking outputs.
When NMA Conclusions Should Be Trusted
NMA conclusions are most reliable when the transitivity assumption is well-supported - when the trials in the network are genuinely comparable in terms of the populations and interventions they study, when the network is well-connected (most treatment pairs have at least some indirect evidence connecting them), and when there is consistency between direct and indirect evidence wherever both exist. Inconsistency - a statistical test comparing direct and indirect estimates for the same treatment pair - is a red flag that either the transitivity assumption is violated or there are systematic biases that differ across the studies in different parts of the network.
NMA conclusions deserve more scepticism when the network is sparse - a small number of trials connected through a single common comparator - when there are large differences in patient populations, disease definitions, or outcome measures across the trials in the network, when heterogeneity within individual comparisons is high, or when the ranking is being used to claim that one treatment is "best" based on a small numerical difference that is within the uncertainty bounds of the estimates.
Its Role in HTA Submissions
Health technology assessment bodies - particularly NICE in the UK and IQWIG in Germany - routinely require NMA as part of the comparative effectiveness evidence for new drug submissions. The NICE reference case explicitly requires an indirect treatment comparison when no direct randomised evidence is available against the relevant comparator. NMA has become a standard submission component, which means that sponsors who do not plan for it from the beginning of clinical development may find themselves submitting NMAs built on networks with structural weaknesses that drive the HTA outcome.
The PRISMA-NMA reporting guideline provides a 32-item extension to the standard PRISMA checklist covering the specific reporting requirements for network meta-analyses, including the description of the network geometry, the transitivity evaluation, the assessment of consistency, and the presentation of results.
NMA and Evidence Synthesis Quality
The quality of an NMA is bounded by the quality of the primary studies that enter it. A technically sophisticated NMA built on trials with high risk of bias, heterogeneous populations, and inconsistent outcome definitions will produce precise-looking estimates that are not reliable. The primary evidence synthesis - the systematic review of the included trials, the risk of bias assessment, the characterisation of clinical heterogeneity - is the foundation on which the NMA sits, and its quality is at least as important as the statistical model chosen.
At NousLab, we work with teams building NMAs for HTA submissions where the review methodology and the analytical model are developed in parallel, not sequentially. The decisions made in the systematic review directly affect the structure of the network and the validity of the transitivity assessment. See NousLab's approach to HTA evidence packages.