Misreading the Margins: How Clinicians Systematically Misinterpret Statistical Significance and What It Costs Patients
The Number That Launched a Thousand Prescriptions
There is perhaps no figure in clinical research more frequently cited—and more consistently misunderstood—than the p-value. Its threshold companion, 0.05, has achieved a kind of mythological status in biomedical publishing: a numerical gate through which findings must pass before they are granted the credibility of significance. Yet decades of methodological scholarship, and a growing chorus of statisticians, have made clear that this threshold communicates something far more narrow than the clinical community has traditionally assumed.
The problem is not merely academic. When clinicians misinterpret statistical significance as a proxy for clinical meaningfulness, the downstream effects manifest in treatment decisions, formulary inclusions, and practice guideline endorsements that may not reflect the actual magnitude of benefit a therapy provides. Understanding how this misreading occurs—and why it persists—requires a closer examination of what these statistics were ever designed to convey.
What a P-Value Actually Says
At its most precise definition, a p-value represents the probability of observing a result at least as extreme as the one obtained, assuming the null hypothesis is true. It does not quantify the probability that the null hypothesis is true. It does not measure the size of an effect. It does not indicate that a finding is clinically important. And it emphatically does not confirm that a result will replicate in practice.
These distinctions matter enormously. A trial enrolling tens of thousands of participants can produce a statistically significant result for a treatment effect so small that it carries no meaningful benefit for individual patients. Conversely, a smaller study may fail to reach significance not because the treatment is ineffective, but because it was underpowered to detect a genuine effect. In both scenarios, a clinician relying on the p-value as a summary verdict is operating on incomplete—and potentially misleading—information.
Surveys of practicing physicians and even biomedical researchers have repeatedly demonstrated that these misinterpretations are not fringe occurrences. A widely cited study published in PLOS ONE found that the majority of surveyed researchers endorsed at least one incorrect interpretation of a significant p-value. Among clinicians without formal statistical training, the rate of misinterpretation is presumed to be higher still.
The Confidence Interval: A Richer Tool, Poorly Used
Confidence intervals were designed, in part, to address the limitations of binary significance testing. Rather than reducing a result to a pass/fail judgment, a 95% confidence interval provides a range of plausible values for the true effect size—offering information about both the direction and the precision of an estimate. In principle, this richer representation should support more nuanced clinical reasoning.
In practice, confidence intervals are frequently misread in ways that replicate the same errors associated with p-values. The most common misinterpretation holds that a 95% confidence interval means there is a 95% probability that the true parameter falls within the stated range. This is incorrect. The interval is a property of the estimation procedure, not a probabilistic statement about any single interval from a single study. Across many replications of a study, 95% of such intervals would contain the true value—but for any given published interval, the true value either is or is not within the bounds.
Perhaps more consequential for clinical practice is the tendency to focus exclusively on whether a confidence interval excludes the null value, rather than examining the clinical relevance of the full range of plausible effects. An interval that spans from a trivially small benefit to a moderately meaningful one may be statistically significant—its lower bound excludes zero—while simultaneously signaling substantial uncertainty about whether the treatment produces effects large enough to justify its use, its cost, or its risk profile.
How Guidelines Amplify the Distortion
The misinterpretation problem does not stop at the level of individual clinicians. It is institutionalized in the processes by which clinical practice guidelines are developed and communicated. When guideline panels assign recommendation grades based partly on whether supporting trials achieved statistical significance, they risk laundering a methodological shortcut into authoritative clinical guidance.
Consider the common scenario in which a guideline endorses a pharmacological intervention based on a trial that achieved p < 0.05 for its primary endpoint. If that endpoint was a surrogate marker—a laboratory value or imaging finding rather than a patient-centered outcome—and if the confidence interval for the absolute risk reduction spans from negligible to modest, the recommendation may carry far more authority than the underlying evidence warrants. Clinicians reading the guideline encounter a letter grade or a strong recommendation, not a nuanced probabilistic statement about uncertain effects.
This dynamic is particularly visible in cardiovascular and metabolic disease management, where surrogate endpoints are common and absolute risk reductions for primary outcomes are frequently small. Statin prescribing patterns, for example, have been shaped substantially by trials reporting relative risk reductions that appear impressive in isolation but translate to absolute benefits that, for lower-risk patients, may be clinically marginal.
The Absolute Versus Relative Risk Divide
One of the most persistent contributors to the misinterpretation problem is the preferential reporting of relative rather than absolute risk reductions in published trials. A therapy that reduces the rate of an adverse event from 4% to 2% has produced a 50% relative risk reduction—a figure that sounds transformative. The absolute risk reduction is 2 percentage points, and the number needed to treat is 50. These figures tell a substantially different story about the practical value of the intervention.
Research consistently shows that clinicians presented with relative risk data are more likely to recommend a treatment than those presented with equivalent absolute risk data. This is not a failure of individual reasoning; it reflects how information is framed and presented in journals, press releases, and continuing medical education materials. When journals permit—or even encourage—the leading presentation of relative metrics without mandatory disclosure of absolute equivalents, they participate in a system that systematically inflates perceived treatment value.
Toward More Statistically Literate Clinical Practice
Addressing this problem requires intervention at multiple levels of the clinical research enterprise. Journals can mandate the reporting of absolute effect sizes, confidence intervals, and numbers needed to treat alongside any presentation of relative risk data. Guideline panels can adopt more rigorous standards for distinguishing statistical significance from clinical significance, explicitly addressing effect size magnitude in their deliberations.
Medical education represents perhaps the most durable lever for change. Graduate and postgraduate training programs in the United States have historically underemphasized biostatistics, and what training exists often focuses on hypothesis testing mechanics rather than on the interpretive skills clinicians need at the point of care. Integrating applied statistical literacy into residency curricula, board examination content, and continuing medical education requirements would equip practitioners to engage more critically with the evidence they encounter.
For the broader research community, the growing movement toward effect size reporting, pre-registration of analysis plans, and Bayesian interpretive frameworks offers a set of methodological tools that may, over time, reduce the field's dependence on a single threshold as the arbiter of credibility.
The Stakes of Getting This Right
Statistical misinterpretation is not a harmless intellectual error. It shapes which treatments patients receive, which drugs health systems purchase, and which interventions receive continued research investment. In a healthcare environment defined by resource constraints and competing therapeutic options, the ability to distinguish a genuinely meaningful effect from a statistically decorated modest one is a clinical competency, not a statistical luxury.
The p-value will likely remain a feature of biomedical research for the foreseeable future. The question is whether the clinical community can develop the interpretive discipline to read it accurately—and to demand the additional information necessary to make it useful.