26 May 2026 · 6 min read

Clinical Trial Protocol Design: The Decisions That Determine Everything

Most clinical trial failures are not failures of execution. They are failures of design that were locked in before the first patient was enrolled.

Clinical Trial Protocol Design: The Decisions That Determine Everything

When a clinical trial produces an uninterpretable result, there is a natural tendency to look for what went wrong during execution. Enrolment was slower than expected. Protocol deviations accumulated. The statistical analysis did not perform as planned. These are real problems, and they deserve attention.

But in most cases, the trials we would call failures, ones that could not answer their primary question, trials that were stopped early for futility, trials that answered a question nobody needed answered, failed because of decisions made in the protocol before a single patient was enrolled. The data could not produce a useful answer because the design could not support one.

This is not a new observation. The ICH E9(R1) addendum on estimands, the FDA's guidance on adaptive designs, and a generation of clinical trial methodology literature all point to the same problem: protocol design is where the science actually happens, and it requires as much care as any other phase of a trial.

The estimand problem: asking the question you mean to ask

The ICH E9(R1) addendum, finalised in 2019 and now shaping regulatory submissions globally across the FDA, EMA, and PMDA, introduced the concept of the estimand as a formal framework for aligning the research question with the statistical analysis. The idea is straightforward: before choosing an analysis method, specify precisely what treatment effect you are trying to estimate, in what population, under what conditions, accounting for what intercurrent events.

Intercurrent events are the things that happen during a trial that complicate the simple comparison between treatment and control: patients discontinue treatment, switch to rescue medications, die of causes unrelated to the treatment being studied. How you handle these events in your analysis determines what question your trial is actually answering, which may or may not be the question you intended.

The most common form of estimand misalignment is designing a trial around an intention-to-treat analysis and then discovering that the ITT analysis does not answer the clinical question because intercurrent events are too frequent and too heterogeneous to be treated as noise. A trial of a treatment for a chronic condition where 40 percent of patients discontinue treatment within the follow-up period is not a trial of the treatment's efficacy. It is a trial of something more complex that the ITT analysis may not capture in a clinically useful way.

Primary endpoint selection: the decision that cannot be undone

The primary endpoint is the single most consequential design decision in a clinical trial. It determines sample size, drives power calculations, defines what counts as success, and shapes how regulatory agencies will assess the trial's results. Getting it wrong at the design stage is a problem that no amount of analytical creativity can fix after the fact.

The most common endpoint selection errors are choosing an endpoint because it is measurable rather than because it is meaningful, and choosing an endpoint that is meaningful but not sensitive enough to detect the expected treatment effect within the planned follow-up period.

The first error produces trials that generate statistically significant results around endpoints that clinicians and patients do not care about. Biomarker endpoints are the canonical case: a treatment that improves a surrogate endpoint without affecting clinical outcomes is a treatment whose benefit is uncertain. Regulatory guidance has tightened around surrogate endpoints precisely because of the gap between statistical success and clinical relevance that this error produces.

The second error produces trials that are underpowered not because the sample size calculation was wrong, but because the assumed effect size was wrong, and the assumed effect size was wrong because the evidence base for the assumption was incomplete. A literature-based effect size estimate that does not account for publication bias is likely to be inflated. Building a power calculation on an inflated estimate produces a trial that is underpowered before recruitment begins.

Sample size: the assumptions are the problem

Sample size calculations are taught as a mathematical exercise, and the mathematics is straightforward. The difficulty lies entirely in the assumptions: the expected effect size, the expected variability, the anticipated dropout rate, the significance threshold, the desired power. Each of these is an estimate, and the estimate for the most influential input, the effect size, is frequently drawn from the same biased literature described above.

A 2014 analysis in BMJ found that effect sizes in trials frequently fall below the estimates used for sample size calculations, leading to systematic underpowering. The discrepancy was largest in areas with few prior trials, where effect size estimates relied on early-stage studies with small samples and inflated results. The methodologically correct response is to use more conservative effect size estimates and to pre-specify interim analyses that allow for sample size re-estimation if the observed effect size during the trial differs from the assumed one.

Adaptive designs, which allow pre-specified modifications to sample size, allocation ratios, or endpoint selection based on interim data, are one approach to managing this uncertainty. They are increasingly accepted by regulatory agencies, including through FDA guidance on adaptive designs, though they require substantially more methodological rigour and pre-specification than conventional designs.

The evidence review that precedes protocol writing

Every design decision described above depends on a thorough understanding of the existing evidence. The right primary endpoint cannot be chosen without knowing what endpoints have been used in prior trials, what their sensitivity characteristics are, and how they have performed as surrogate markers for clinical outcomes. The right effect size cannot be assumed without a critical appraisal of the prior literature that accounts for publication bias. The right patient population cannot be defined without understanding the heterogeneity of treatment effects across subgroups in previous studies.

This evidence work is not incidental to protocol design. It is the foundation of it. And it is frequently done too quickly, too narrowly, or too late in the process. A protocol that is written before the evidence review is complete will reflect the assumptions held at the start of the process rather than the picture that emerges from a complete assessment of the literature.

We have seen this pattern directly with research teams that come to NousLab at the protocol development stage. The most common finding is not that the proposed design is wrong, it is that it is based on a less complete picture of the prior evidence than a comprehensive search would support. Adjusting the effect size assumption based on a more complete literature review, refining the patient population based on subgroup data from trials the team had not found, identifying a validated endpoint that had been used successfully in an adjacent indication, these are the kinds of findings that change protocols before they are locked rather than after the trial fails.

If you are at the protocol development stage for a clinical trial and want to stress-test your evidence base before committing to a design, our clinical research use cases describe how we approach this kind of work.

Jesus Arias
Jesus Arias
Founder & CEO at NousLab
Connect on LinkedIn

Join the Research Revolution

See how NousLab can accelerate your organization's research

AI