How AI Is Changing Systematic Reviews in Medical Research
AI is not automating systematic reviews. It is handling the parts that slow researchers down and the difference matters.
If you have ever been part of a systematic review team, you know the specific kind of exhaustion it produces. Not the good kind that comes from solving a hard problem. The other kind, the one that comes from screening your four hundredth abstract on a Tuesday afternoon, trying to remember whether you already excluded a paper with a near-identical title three weeks ago.
The average systematic review takes between 12 and 18 months to complete. Some take longer. That number has barely moved in decades, even as the volume of published biomedical literature has grown exponentially. There are now more than 1.5 million new biomedical papers indexed on PubMed every year. The tools most research teams use to manage that volume have not kept up.
AI is starting to change that. Not in the way it is often hyped, not as a magic system that produces a finished review at the click of a button, but in a quieter, more useful way. It is getting better at handling the parts of the process that are repetitive, error-prone, and genuinely do not require a trained researcher's judgment.
The part nobody talks about: screening
Title and abstract screening is where most of the time goes in the early stages of a review. A typical search might return 3,000 to 8,000 records. Each one needs to be read, assessed against inclusion criteria, and either kept or excluded with a reason. Most protocols require two independent reviewers. Disagreements need to be resolved. The whole thing is necessary, methodologically important, and mind-numbing.
AI-assisted screening tools can now handle a significant portion of this workload. Systems trained on prior inclusion and exclusion decisions learn to rank records by relevance, pushing the most promising ones to the top and flagging likely exclusions for rapid confirmation. A 2023 study published in Systematic Reviews found that semi-automated screening tools reduced reviewer workload by an average of 57% while maintaining recall rates above 95% for included studies. BioMed Central has published several evaluations of these tools if you want to dig into the methodology.
That does not mean you hand the screening over to the machine and walk away. The researcher still makes the final call. What changes is the volume of records requiring careful human attention. Instead of reading 5,000 abstracts with equal effort, a reviewer focuses their attention on the 400 the system flagged as genuinely ambiguous. That is a different job, and a better use of time.
Data extraction is the next bottleneck
Once studies are selected, the work shifts to extracting structured information: study design, population characteristics, interventions, outcomes, risk of bias assessments. This is where errors accumulate quietly. A reviewer misreads a confidence interval. An outcome measure gets categorised inconsistently across studies. A table footnote that changes the interpretation of a result goes unnoticed.
Large language models are surprisingly capable at structured data extraction from scientific text, not perfectly, but well enough to produce a first pass that a reviewer can check and correct rather than build from scratch. Correcting is faster than creating. Teams that have integrated AI-assisted extraction into their workflows report meaningful reductions in the time spent on this phase, though published benchmarks vary depending on the complexity of the review question and the quality of the source papers.
The honest caveat here is that AI extraction still struggles with ambiguity in methods reporting. When a paper describes its randomisation procedure in vague terms, or when outcome definitions shift between the methods and results sections, a model will often produce a confident-sounding extraction that is subtly wrong. Human review of these cases remains essential, which is why the better implementations flag low-confidence extractions for closer attention rather than burying them.
Synthesis and the limits of automation
This is where I would urge caution about overclaiming. Narrative synthesis, making sense of a heterogeneous body of evidence, identifying patterns, explaining inconsistencies, drawing conclusions that are proportionate to the quality of the data, is still a deeply human task. It requires domain knowledge, methodological judgment, and a willingness to say "the evidence does not support a firm conclusion" even when funders and editors would prefer otherwise.
AI can assist with structuring a synthesis, identifying inconsistencies across extracted data, and drafting sections of the write-up. It cannot replace the intellectual work of interpretation. Teams that have tried to use language models to generate synthesis sections directly have generally found the output superficially plausible but methodologically shallow. The citations look right. The conclusions are often not.
The appropriate role for AI in this phase is as a drafting assistant and consistency checker, not as an autonomous analyst. That distinction matters more in medical research than in most fields, because the downstream consequences of a poorly conducted review can affect clinical guidelines and patient care.
What this means in practice
The research teams getting real value from AI tools in systematic reviews are not the ones trying to automate the whole process. They are the ones who have identified the specific bottlenecks in their workflow and applied AI assistance precisely there.
A team doing a rapid review for a health technology assessment body might use AI screening to compress the title and abstract phase from three weeks to four days, then apply the same rigour as always to the rest of the process. A university group running a full Cochrane-style review might use AI extraction to produce first drafts of evidence tables, freeing senior researchers to spend more time on risk of bias assessment and synthesis. The gains are real, but they are specific and contextual.
One thing worth noting: methodological transparency has not kept up with adoption. Many published reviews that used AI-assisted screening or extraction do not report this clearly in their methods sections. As these tools become more common, reporting standards will need to evolve. The EQUATOR Network, which maintains reporting guidelines for health research, is one body working on this, though guidance specific to AI-assisted reviews is still developing.
The 12-month timeline is not inevitable
Systematic reviews will probably always take longer than anyone wants. The process is rigorous by design, and rigor takes time. But the idea that a well-conducted review must consume the better part of a year is partly a product of tooling that has not changed much since the 1990s.
The teams doing this work deserve better tools. Not tools that cut corners or that obscure what was done and why, but tools that handle the mechanical load so that researchers can spend their time on the work that actually requires their expertise.
That shift is already happening in some research environments. It will take longer to reach others. But the direction is clear, and the evidence that it works, when implemented carefully, is growing.