The Unmined Record: Why Frontline Clinical Observations Fail to Reach the Evidence Base
Photo: Wellcome Institute for the History of Medicine; Arnold, Ken, 1960-; Porter, Roy, 1946-2002; Wilkinson, Lise, Lady, No restrictions, via Wikimedia Commons
Consider the hospitalist who notices, across dozens of patients over two years, that a particular antibiotic regimen produces unexpected glycemic instability in diabetic patients not flagged in the current prescribing literature. Or the community oncologist who observes a pattern of treatment response in a demographic subgroup underrepresented in the pivotal trial that established the standard of care. These observations exist—documented in clinical notes, reflected in lab values, embedded in the data architecture of electronic health record systems. What they almost never become is evidence.
The gap between clinical observation and published knowledge is not new, but the scale at which it now operates is. The widespread adoption of electronic health record (EHR) systems across US healthcare institutions has created an unprecedented repository of longitudinal patient data. The Office of the National Coordinator for Health Information Technology reported that by 2021, more than 96% of non-federal acute care hospitals had adopted certified EHR technology. In theory, this digitization should have accelerated the translation of frontline clinical insight into generalizable evidence. In practice, the vast majority of that data remains institutionally siloed, clinically underutilized, and scientifically invisible.
The Publication Bottleneck and Who It Excludes
Academic medical centers, which house the majority of funded clinical researchers, operate within a publication incentive structure that systematically devalues incremental, observational, or single-institution findings. Peer-reviewed journals—particularly high-impact publications—favor large randomized trials, multi-site studies, and findings with broad generalizability. A case series of fifteen patients observed at a community hospital in rural Tennessee may contain clinically significant information, but it faces nearly insurmountable obstacles to peer-reviewed publication: limited statistical power, questions about selection bias, and the absence of the institutional infrastructure—biostatistics support, research coordinators, grant funding—that facilitates manuscript preparation.
The clinicians most likely to accumulate meaningful observational data are frequently those least positioned to publish it. Frontline practitioners at non-academic institutions carry patient volumes that leave minimal protected time for research activities. A 2020 survey conducted by the American College of Physicians found that a majority of practicing internists reported spending fewer than two hours per week on any form of scholarly activity. For those without academic appointments, the path to publication involves navigating journal submission systems, responding to peer review, and managing revision cycles—a process that can extend over years and offers no direct professional reward for clinicians outside of academic promotion tracks.
What Gets Lost, and Where It Goes Instead
The knowledge that does not enter the peer-reviewed literature does not simply disappear. It circulates through informal channels—conference hallway conversations, departmental grand rounds, clinical listservs, and the accumulated tacit expertise of experienced practitioners. This informal transmission is not without value; clinical mentorship and collegial knowledge-sharing represent meaningful vectors for practice improvement. However, informal knowledge transfer is geographically constrained, systematically unverifiable, and inaccessible to clinicians who are not embedded in the right professional networks.
The consequences for clinical decision-making are material. Treatment protocols and diagnostic guidelines are built on the published evidence base, which reflects a particular kind of clinical knowledge—prospectively designed, rigorously controlled, and conducted largely within academic research environments. The patient populations in those environments are not always representative of the patients seen in community health centers, rural hospitals, or safety-net institutions. When the evidence base is constructed from a narrow slice of clinical experience, protocols derived from it carry embedded assumptions that may not hold universally.
Diagnostic error is one domain where this gap has particularly well-documented consequences. The National Academy of Medicine's 2015 report on diagnostic error in healthcare identified knowledge gaps—specifically, the failure to incorporate emerging clinical insights into diagnostic frameworks—as a contributing factor in a significant proportion of diagnostic failures. Many of those knowledge gaps exist not because the relevant observations have never been made, but because the mechanisms for converting observation into actionable evidence are inadequate.
The EHR as Latent Research Infrastructure
The irony of the current situation is that the infrastructure for capturing clinical observations at scale already exists. EHR systems contain structured and unstructured data on patient presentations, treatment decisions, and outcomes that, in aggregate, could support observational research of considerable methodological rigor. The emergence of large-scale clinical data networks—including PCORnet, the National Patient-Centered Clinical Research Network, and various health system–specific data collaboratives—demonstrates that EHR data can be harmonized and analyzed across institutions to generate meaningful evidence.
Yet these networks, while valuable, are not designed to capture the kind of granular, clinician-generated observational insight that constitutes the ghost literature problem. They operate on structured data fields—diagnosis codes, medication records, laboratory values—rather than the interpretive observations documented in clinical notes. Natural language processing technologies capable of extracting meaning from unstructured clinical text have advanced considerably, but their deployment for systematic knowledge extraction remains limited to well-resourced academic institutions with dedicated informatics capacity.
Furthermore, even where the technical infrastructure exists, the governance frameworks for using patient data in research contexts impose compliance burdens that can deter clinician-researchers from pursuing observational studies. Institutional Review Board processes, while essential for patient protection, are not always calibrated to the risk profile of retrospective observational analyses using de-identified EHR data. The administrative overhead of obtaining approval for a low-risk chart review can be substantial enough to discourage the inquiry entirely.
Proposed Pathways Toward Recovery
Several frameworks have been proposed to address the systemic underutilization of frontline clinical knowledge. Rapid publication formats—brief communications, structured case series, and registered reports for preliminary findings—have gained traction at some journals as mechanisms for lowering the barrier to entry for clinician-generated observations. Journals including BMJ Open and JAMA Network Open have explicitly expanded their scope to accommodate real-world observational data from non-academic settings, though uptake among community practitioners remains limited.
Some health systems have experimented with embedded research programs that provide practicing clinicians with protected time, biostatistical support, and IRB facilitation for small-scale observational studies. These programs, where adequately resourced, have demonstrated that clinician-researchers outside of academic medicine can produce publishable findings when structural barriers are systematically reduced. The challenge is scaling such programs in an environment where health system margins are under sustained pressure.
At the policy level, there is a compelling argument for treating clinician-generated observational data as a public health resource requiring active stewardship. The National Institutes of Health and the Agency for Healthcare Research and Quality have both articulated interest in expanding the evidence base through real-world data, but funding mechanisms specifically designed to support community clinician research remain limited relative to the opportunity.
Toward a More Complete Evidence Base
The medical knowledge infrastructure in the United States is, in important respects, built on an incomplete foundation. The published evidence base represents the clinical experiences of a particular subset of patients, observed by a particular subset of clinicians, in institutional contexts that do not reflect the full diversity of American healthcare delivery. The observations that fall outside that subset are not inherently less valid—they are simply less visible.
Addressing this gap requires changes at multiple levels simultaneously: journal practices that accommodate incremental evidence, health system structures that enable clinician scholarship, informatics capabilities that extract knowledge from existing data, and funding priorities that recognize the value of community-based observational research. None of these changes is straightforward. All of them are necessary if the evidence base that guides clinical decision-making is to reflect the full complexity of the patients it is meant to serve.