04 Aug 2026 · 5 min read

The Reproducibility Crisis in Preclinical Research

When Bayer scientists tried to replicate preclinical oncology findings, they succeeded in fewer than a quarter of attempts. This is not an anomaly - it is a structural feature of how preclinical research is conducted and reported.

The Reproducibility Crisis in Preclinical Research

In 2012, Glenn Begley and Lee Ellis published a paper in Nature reporting that Amgen's scientists had attempted to replicate 53 landmark studies in haematology and oncology. They were able to reproduce the findings of six. That is an 89% failure rate in studies that had, in many cases, been the basis for clinical trial programmes and substantial R&D investment. Around the same time, Bayer HealthCare reported that only 20–25% of published preclinical studies could be confirmed in their internal replication attempts. These numbers are striking, and they have not substantially improved in the decade since.

Why Preclinical Research Fails to Replicate

The causes of irreproducibility in preclinical research are multiple and interact with each other in ways that make them hard to disentangle. The first is statistical: preclinical studies routinely use sample sizes that are underpowered to detect true effects reliably, and the same pressure toward positive results that operates in clinical research operates with even less constraint in animal studies. With n=6 per group in a rodent experiment, a positive result is compatible with both a true effect and chance variation.

The second is biological: animal models are imperfect proxies for human disease, and the conditions under which animals are housed and handled vary significantly across laboratories. A finding in a specific mouse strain, under specific housing conditions, with a specific pathogen burden and specific microbiome composition, may not replicate even in the same species in a different facility. These are not exotic problems - they are routine features of in vivo research that are rarely reported with enough detail to allow replication.

Cell Line Contamination and Reagent Variability

In vitro research has a contamination problem that has been known since the 1960s but continues to undermine published findings. The misidentification and cross-contamination of cell lines has affected thousands of published studies. HeLa cells - derived from Henrietta Lacks in 1951 - have contaminated dozens of other cell lines that researchers believed were different. Studies published on "cell line X" were actually studying HeLa. International databases of authenticated cell lines exist, and authentication by short tandem repeat profiling is increasingly required by journals, but compliance is incomplete.

Reagent variability - differences in antibody specificity, compound purity, viral titre, and assay kit performance across batches and suppliers - is a further source of non-reproducibility that is almost never reported in publications. A positive result obtained with one antibody lot may not replicate with a different lot of nominally the same antibody.

HARKing in Animal Research

HARKing - Hypothesising After Results are Known - is at least as prevalent in preclinical research as in clinical research, and possibly more so, because the pre-registration culture that is now well-established in clinical trials has barely reached animal research. An experimenter who runs multiple outcome measures and reports only the significant ones is not being dishonest in any simple sense - they are following the informal norms of their field. But the result is an inflated false positive rate and a literature that systematically overstates effect sizes.

The absence of pre-registration in animal research means that the distinction between confirmatory and exploratory findings is almost never clearly marked, and most published animal studies present exploratory findings as though they were confirmatory.

The NC3Rs and ARRIVE Guidelines

The National Centre for the Replacement, Refinement and Reduction of Animals in Research (NC3Rs) developed the ARRIVE guidelines (Animal Research: Reporting of In Vivo Experiments) to address the reporting deficiencies in animal studies. ARRIVE 2.0, published in 2020, provides a 21-item checklist covering study design, sample size calculation, randomisation, blinding, outcome measures, and statistical analysis. Many journals now require or recommend ARRIVE compliance, and compliance has demonstrably improved the reporting quality of animal studies - though it has not yet resolved the underlying practices that generate irreproducible results.

Alongside ARRIVE, the Experimental Design Assistant (EDA) developed by the NC3Rs helps researchers plan animal studies with appropriate sample sizes, randomisation, and blinding before the experiment begins. These tools are useful, but they are adopted voluntarily, and their adoption is uneven.

What This Means for Translation

The failure rate of preclinical-to-clinical translation in drug development is not explained entirely by the inherent complexity of human biology. A meaningful portion of it reflects the fact that clinical programmes are sometimes built on preclinical findings that could not be reproduced even before they crossed species boundaries. Evidence synthesis in the preclinical space - systematic reviews of animal studies using tools like the SYRCLE risk of bias tool - is still a minority practice, but it is one that consistently reveals a more uncertain evidence base than individual studies suggest.

At NousLab, we have worked with teams preparing evidence packages for first-in-human decisions where a systematic review of the preclinical literature revealed that the animal data supporting the programme was less consistent than the summary papers had implied. That kind of honest appraisal before a clinical investment is made is exactly the kind of work that saves time and resources in the long run. See how NousLab supports preclinical evidence synthesis.

Jesus Arias
Jesus Arias
Founder & CEO at NousLab
Connect on LinkedIn

Join the Research Revolution

See how NousLab can accelerate your organization's research

AI