30 Jun 2026 · 5 min read

Meta-analysis: When Pooling Studies Makes Things Worse

Meta-analysis is often treated as the gold standard of evidence - but pooling the wrong studies can produce a precise answer to the wrong question. Understanding when not to pool is as important as knowing how.

Meta-analysis: When Pooling Studies Makes Things Worse

In 1904, Karl Pearson published what is often cited as the first formal attempt to combine results across studies - an analysis of typhoid vaccination data from five different populations. He was trying to solve a real problem: individual studies were too small to be conclusive. The instinct was right. The method, however, carries risks that are still underappreciated more than a century later.

The Appeal of Pooling

Meta-analysis offers something seductive: statistical power that no single study can achieve on its own. When a treatment effect is modest and trials are small, combining them into a pooled estimate feels like common sense. And often it is. The logic is sound when studies are asking the same question in the same way in comparable populations with comparable outcomes measured the same way. The trouble is that this condition is almost never perfectly met, and the field has developed an unhealthy tolerance for pooling studies that do not belong together.

What the I² Statistic Actually Tells You

The I² statistic, introduced by Higgins and colleagues in 2003, estimates the proportion of variability in effect estimates that is due to heterogeneity rather than sampling error. An I² of 0% suggests that variability is consistent with chance alone. An I² of 75% or above is conventionally labelled "considerable" heterogeneity. But I² is a relative measure, not an absolute one - it tells you nothing about the magnitude of between-study variance, only its share of total variance. A large trial with tight confidence intervals can have a high I² even when the absolute between-study differences are small. A set of small trials can have a low I² even when the actual treatment effect varies wildly across contexts.

The Cochrane Handbook for Systematic Reviews of Interventions is explicit on this point: I² should be interpreted alongside the estimate of tau² (the between-study variance on the natural scale) and the prediction interval, which tells you the range of effects you would expect in a new study. A prediction interval that crosses the line of no effect is a signal that the pooled estimate, however statistically significant, may not apply to any specific clinical context.

Clinical vs Statistical Heterogeneity

Statistical heterogeneity - the kind I² measures - is downstream of clinical and methodological heterogeneity, and the latter is harder to quantify but more important to understand. Clinical heterogeneity refers to meaningful differences in populations, interventions, comparators, and outcomes across studies. A meta-analysis of antihypertensive drugs that pools trials conducted in patients with mild hypertension alongside trials in patients with severe hypertension, using different outcome definitions and different follow-up durations, may produce a statistically homogeneous result - low I² - simply because the noise from different directions cancels out. That pooled estimate is not more reliable. It is quietly misleading.

The most dangerous meta-analyses are those with low I² and high clinical heterogeneity - they look clean on the surface while concealing fundamental incompatibilities beneath.

When Pooling Actively Harms Conclusions

The history of evidence-based medicine contains well-documented cases where pooling obscured rather than clarified. The early meta-analyses of antiarrhythmic drugs after myocardial infarction showed trends toward benefit that were later reversed by the CAST trial, which found that the drugs increased mortality. The pooled signal had been confounded by short follow-up durations and surrogate endpoints. Similarly, early pooled analyses of hormone replacement therapy and cardiovascular disease drew heavily on observational data with a healthy user bias that randomised trials later contradicted directly.

These are not failures of the meta-analytic technique per se - they are failures of the decision to pool without adequately interrogating the clinical and methodological differences between contributing studies.

The Decision to Pool: A Framework

Before pooling, a review team should ask a set of questions that go beyond statistical tests. Are the populations clinically comparable? Are the interventions the same in terms of dose, duration, and delivery? Are the comparators equivalent? Are the outcome definitions identical, or have they been harmonised in a way that introduces assumptions? Is the follow-up duration consistent, and does it matter for the outcome in question? If the answers reveal meaningful variation in any of these dimensions, a narrative synthesis or a subgroup analysis is often more honest than a pooled estimate.

Random-effects meta-analysis is not a solution to clinical heterogeneity - it is a statistical acknowledgement of it. A random-effects model assumes that the true effects are distributed around a mean, but it says nothing about whether that mean is clinically meaningful or whether the distribution makes sense given what we know about the biology and context of the studies included.

What This Means in Practice

At NousLab, we work with research teams that are building systematic reviews and evidence syntheses to support regulatory submissions, HTA dossiers, and clinical guideline development. One of the most consistent patterns we observe is the pressure to produce a pooled estimate - because funders, regulators, and guideline developers expect one. Resisting that pressure when the evidence does not support pooling is one of the hardest and most important judgements in evidence synthesis.

The responsible use of meta-analysis requires being willing to say, clearly and explicitly, that the studies in front of you should not be combined - or that if they are combined, the result should be interpreted with named caveats that are not buried in supplementary material. The I² statistic is a starting point, not an answer. The real work is clinical and methodological judgement, applied before the analysis begins.

If you are designing a systematic review or evidence synthesis and want to think through the pooling decision carefully, get in touch with the NousLab team.

Jesus Arias
Jesus Arias
Founder & CEO at NousLab
Connect on LinkedIn

Join the Research Revolution

See how NousLab can accelerate your organization's research

AI