HARKing: The Research Practice Nobody Talks About Enough
HARKing - presenting post-hoc hypotheses as if they were a priori - is widespread, rarely acknowledged, and one of the most direct contributors to the replication crisis. It is also not discussed honestly enough in methods training.
In 1998, Norbert Kerr published a paper in Personality and Social Psychology Review coining the term HARKing - Hypothesising After Results are Known. He defined it as presenting post-hoc hypotheses as if they were a priori in scientific communications. The paper was careful and measured. It estimated that the practice was common. It identified the structural incentives that produce it. And it was largely ignored for the better part of a decade, until the replication crisis made the consequences undeniable.
What HARKing Is and Is Not
HARKing is not fraud in the straightforward sense of fabricating data. It is a more subtle distortion: a researcher runs an analysis, finds a pattern that was not hypothesised in advance, and writes the paper as though that pattern was the predicted outcome. The data is real. The analysis is real. The distortion is in the narrative framing - the paper reads as a confirmatory test of a hypothesis that was actually generated by the data it purports to test.
This matters because of the logic of hypothesis testing. A p-value of 0.05 means that, if the null hypothesis is true, there is a 5% chance of observing a result this extreme or more extreme by chance. But this probability is only valid if the hypothesis was specified before looking at the data. If the hypothesis was selected from among many potential hypotheses after looking at the data - choosing the one that happened to be significant - the false positive rate is no longer 5%. It may be 30%, 50%, or higher, depending on how many implicit hypotheses were tested.
How Widespread It Is
Direct evidence on the prevalence of HARKing is hard to obtain precisely because it requires researchers to report honestly about practices they may not even recognise as problematic. Survey data from various fields suggests that a substantial minority of researchers acknowledge doing it, and that the actual prevalence is almost certainly higher than self-report surveys capture. Analysis of registered trial reports versus published trial reports - comparing pre-specified outcomes with reported outcomes - routinely finds that outcome switching is common: primary outcomes are downgraded, secondary outcomes that happened to be significant are elevated, and the paper's narrative is built around the positive findings rather than the pre-specified analysis plan.
A 2014 analysis by Dwan and colleagues in PLOS ONE examined discrepancies between registered and published outcomes in RCTs and found that outcome reporting bias was present in the majority of trials examined. This is HARKing at the level of the primary outcome - not a subtle analytic choice but a fundamental misrepresentation of what the study was designed to test.
The Relationship to the Replication Crisis
The replication crisis - the systematic failure of published findings to hold up when independently tested - is not a single problem with a single cause. But HARKing is one of the most direct contributors. When a finding is reported as confirmatory but was actually exploratory, the expected replication rate is the prior probability of the hypothesis being true times the power of the replication study - a very different calculation from the one implied by the original paper's p-value. Fields where HARKing is most prevalent are the same fields where replication rates are lowest. This is not a coincidence.
The solution is not to stop doing exploratory research - exploratory research is how science generates new hypotheses. The solution is to clearly label exploratory findings as exploratory, and to require independent confirmatory tests before treating them as established.
Pre-Registration as a Structural Solution
Pre-registration - the practice of publicly recording the hypothesis, design, and analysis plan of a study before data collection - is the primary structural response to HARKing. Major pre-registration platforms include ClinicalTrials.gov (required for clinical trials in many jurisdictions), the WHO ICTRP, and the Open Science Framework (OSF), which is used widely in psychology, social science, and increasingly in biomedical research. The AsPredicted platform offers a simple, brief pre-registration format that reduces the barrier to entry.
Pre-registration does not prevent exploratory analyses - it simply makes the distinction between confirmatory and exploratory analyses transparent. A pre-registered study can still report unexpected findings; it must just be clear that they are unexpected. This is not a high bar scientifically. It is a cultural and incentive problem: journals still preferentially publish surprising, clean, positive results, and the incentive to present post-hoc hypotheses as a priori will persist as long as that preference does.
Writing Honestly About Exploratory Findings
Writing honestly about exploratory findings requires a specific kind of epistemic courage in the current publishing environment. Framing a finding as hypothesis-generating rather than hypothesis-confirming signals to reviewers that the result requires replication - which can make it harder to publish in high-impact journals that prefer clean narratives. But the alternative - a literature full of overconfident findings that do not replicate - is worse for science and, ultimately, for the researchers who build on those findings.
Practical guidance for authors includes using language that calibrates to evidence strength ("these findings suggest," "this observation warrants replication" rather than "we demonstrated," "our results confirm"), reporting all outcomes that were measured and not just those that were significant, and using the registered analysis plan as the primary analysis while clearly labelling secondary and exploratory analyses. At NousLab, evidence synthesis that is honest about what is known and what is speculative is central to how we think about building a useful knowledge base. See how NousLab structures evidence quality assessments.