The Grey Literature Problem in Systematic Reviews
The studies that never reach PubMed - unpublished trials, internal reports, conference abstracts - can be decisive for the conclusions of a systematic review. Ignoring them is not a neutral choice.
In 2008, Erick Turner and colleagues published a landmark analysis in the New England Journal of Medicine examining antidepressant trials submitted to the FDA. Of 74 registered studies, 38 had positive results - and 37 of those were published. Of the 36 studies with negative or questionable results, 22 were not published at all, and 11 were published in a way that conveyed a positive outcome. The effect sizes in the published literature were, on average, 32% higher than those in the FDA database. This was not fraud. It was the ordinary operation of scientific publishing, selecting for results that fit a particular direction.
What Grey Literature Is
Grey literature is a broad category covering any evidence-based material that has not been published through conventional commercial or academic publishing channels. This includes unpublished clinical trial reports, conference abstracts and posters, dissertations and theses, government and regulatory agency reports, internal industry documents released through litigation or freedom of information requests, and technical reports from research institutions. For a systematic review to be truly comprehensive, it must actively search for and incorporate grey literature - treating PubMed as the only source is a methodological flaw, not a shortcut.
The Cochrane Handbook dedicates substantial guidance to grey literature searching, recommending searches of trial registries (ClinicalTrials.gov, WHO ICTRP), regulatory databases (FDA, EMA), and conference proceedings as standard components of a complete search strategy.
Publication Bias and Its Consequences
Publication bias - the tendency for positive studies to be published and negative studies not to be - systematically inflates effect estimates in the published literature. The problem is compounded by outcome reporting bias, where studies are published but only the significant outcomes are reported fully, and time-lag bias, where positive trials are published faster than negative ones. All three biases operate in the same direction: they make interventions look more effective than they are.
The consequence for systematic reviews is significant. A review drawing only on published literature will inherit these biases. Its pooled estimate will be inflated. Its conclusions will be too optimistic. And if that review is being used to support a clinical guideline, a reimbursement decision, or a regulatory submission, those inflated conclusions have real-world effects.
The Funnel Plot: Useful, Limited, Misunderstood
The funnel plot is the standard visual tool for assessing publication bias in meta-analysis. It plots effect estimates on the horizontal axis against a measure of study precision (often standard error or sample size) on the vertical axis. In the absence of bias, the plot should be approximately symmetrical - small studies scattered widely around a central estimate, large studies clustered closely. Asymmetry, particularly the absence of small negative studies in the lower-left quadrant, suggests publication bias.
Funnel plot asymmetry is a useful signal, but it has important limitations: it requires at least ten studies to be interpretable, it cannot distinguish publication bias from true small-study effects, and it is not a test for grey literature completeness. A symmetrical funnel plot does not mean that the grey literature has been adequately searched - it means that, within the included studies, there is no obvious asymmetry. These are different things.
Searching for What Has Not Been Published
Finding grey literature requires a different set of search strategies than database searching. Trial registries reveal studies that were registered but never published, and comparing registered primary outcomes with published outcomes can identify outcome reporting bias directly. Regulatory agencies - the FDA in the United States and the EMA in Europe - publish clinical study reports for approved products that contain far more granular data than the corresponding journal publications. Researchers who have obtained these reports through freedom of information requests have repeatedly found that the published literature understates harms and overstates efficacy.
Conference abstract databases - including those from major specialty societies - capture work that may never reach full publication but was considered significant enough to present publicly. Contacting study authors and relevant pharmaceutical companies directly is also a standard Cochrane recommendation, though the yield is variable.
The Practical Challenge
Grey literature searching is time-consuming, inconsistently indexed, and frequently incomplete even when done carefully. This creates a tension in systematic review production: the more rigorous the grey literature search, the more resource-intensive the review, and the harder it is to replicate. Review teams face pressure to deliver conclusions quickly, and grey literature searching is the component most likely to be abbreviated when time and budget are constrained.
At NousLab, we see this tension regularly. The answer is not to abandon grey literature searching - it is to be explicit about what was searched, what was not, and what the implications of that limitation are for the conclusions. A systematic review that acknowledges the boundaries of its search is more trustworthy than one that implies completeness it cannot demonstrate.
If your review is being used to support a regulatory or HTA decision, the grey literature question is not optional. The agencies will ask about it. Our platform is built to support comprehensive evidence searches, including structured grey literature workflows that are documented and reproducible.