Clinical data abstraction remains one of the biggest bottlenecks in biomedical research, making it a prime candidate for AI. However, accuracy, privacy and security have remained barriers thus far.
A recent pilot by healthcare technology leader Verily Health put that challenge to the test. In collaboration with UCHealth, the University of Colorado Anschutz and RefinedScience, the project evaluated whether AI could transform complex, unstructured clinical data into structured, research-ready variables.
Why clinical data abstraction is a bottleneck to progress
Healthcare organizations are rich in clinical data, but much of it is fragmented and unstructured, consisting of physician notes, pathology and genomic reports, and other clinical documents. This data is often required to understand the full patient journey, but isn’t easy to extract and make available for advanced analytic modeling (e.g., from scanned PDFs).
The responsibility of data preparation often falls to the experts, including researchers. These top innovators regularly find themselves acting as “data wranglers,” spending valuable time dealing with messy data instead of generating new scientific insights.
Beyond wasted time and effort, the bottleneck is also an obstacle to potentially valuable discoveries. As Kathryn Twyman, director of AI and Data Science at Verily Health, explained:
“Clinical data resolution is the essential first step to unlocking the latent value within clinical datasets, from identifying new biomarkers to uncovering therapeutic opportunities that might otherwise remain hidden. When abstracting this data is a manual process, it slows the path from data to insight. If we can automate that process in a trusted, expert-led way, we can speed up that time to value while also freeing up our innovators to do what they do best.”
Why traditional AI has underdelivered
The promise of AI to automate clinical data abstraction has remained limited by the challenge of clinical interpretation. AI must understand clinical context, terminology and relationships across multiple records. This is where the technology by itself hits a ceiling.
“Standard AI models treat clinical notes as flat, unstructured text, lacking the grounding to distinguish between a clinician’s hypothesis and a biological fact. Without a structural ‘map’ of medical relationships, standard AI is prone to reasoning drift, where it conflates anecdotal patient history with universal clinical truth,” said Twyman.
Traditional “zero-shot” AI models, where AI is asked to perform a task without being specifically trained or adapted for it, struggle to make these important distinctions.
“In this test case, for example, we initially saw accuracy of only around 72% because the models did not understand that bone marrow aspirate holds precedence over a peripheral blood smear when blast percentages differ,” noted Twyman. “While traditional AI approaches can certainly save time, it takes a healthcare-native approach to transform unstructured clinical information into trusted abstractions.”
Testing a healthcare-native approach against one of the toughest medical challenges
In the case of the pilot, acute myeloid leukemia (AML) provided an ideal test case because of its heterogeneous patient populations, complex genomic information, multiple report types and nuanced clinical interpretation.
As Steve Hess, CIO of UCHealth, described, “We gave you the ‘hardest of the hard’ problem for this pilot, and it was on purpose to see how your tech would perform and how our teams would work together. The pilot was a success.”
At the center of the project was a clinical knowledge-augmented approach that helped the AI distinguish biological facts based on clinically guided reasoning. Twyman explained:
“Our healthcare-native architecture serves as a ‘clinical foundation layer’ that forces the AI to ground its extractions in established medical ontologies. By using a clinical knowledge-augmented approach, the AI can independently verify if a relationship is medically plausible before it is confirmed by a clinician. This is the structural mechanism that bridges the accuracy gap. It stops the AI from ‘guessing’ and forces it to ‘reason’ based on a clinical protocol backed biomedical framework. For example, when extracting critical clinical variables like myeloblast percentage — a key AML metric — a simple ‘zero-shot’ AI approach achieved 72% accuracy. Once clinical knowledge was incorporated into the extraction pipeline, accuracy rose to more than 95%.”
Beyond improving accuracy, the pilot projected that an estimated 1,200 hours of manual clinical data abstraction could be reduced to just 40 hours. By making high-quality clinical abstractions faster and more scalable, approaches like this have the potential to accelerate scientific discovery.
“We’ve already seen what’s possible when abstracting concepts from unstructured clinical data. For example, using a manually curated dataset, RefinedScience identified a complex precision biomarker for a specific subgroup of AML patients who may respond to cusatuzumab, an anti-CD70 antibody drug previously deprioritized after a Phase II study,” said Twyman. “As AI makes those abstractions faster and more scalable, researchers can spend less time preparing data and more time uncovering these kinds of insights across much larger datasets.”
What this pilot demonstrates for precision medicine
By solving the abstraction bottleneck, this approach does more than accelerate timelines. It redefines the role of the researcher, allowing them to shift from managing fragmented data to unlocking the high-value insights that will help advance precision medicine.
For a deeper dive into the technology, methodology and results behind the pilot, download the full case study from Verily Health.





