Can we trust AI-generated RWD? A systematic validation approach in prostate cancer

Published on

September 30, 2026

By

Eunice Hankinson, MSN, FNP-C, Clinical Director

The growing use of large language models (LLMs) to extract clinical detail from electronic health records (EHRs) has created new opportunities to generate real-world data (RWD) at scale. For oncology researchers and life sciences decision makers, this unlocks valuable information contained within unstructured data sources, including clinical notes, pathology reports, and other sources that are often difficult to access through traditional abstraction methods. However, as AI-enabled methods are increasingly integrated into evidence generation, questions arise around quality, reliability, and fitness for purpose.

At ISPOR 2026, the Flatiron team presented “Assessing Quality of a Large Language Model (LLM)-Derived Prostate Cancer (PC) Real-World Dataset: An Application of the Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework.” The analysis uses the VALID framework to assess the quality of Flatiron’s prostate Panoramic dataset across multiple key dimensions. In this interview, Eunice Hankinson discusses the rationale behind the development of the VALID framework, the study's key findings, and the future of LLM-extracted data in real-world evidence (RWE) generation.

validposter-image

The use of LLMs and other approaches to generate structured RWD is expanding rapidly. What challenges in evaluating AI-extracted data led to the development of the VALID Framework, and what gap was it designed to address?

The honest answer is that for a long time, there wasn't a transparent, systematic way to answer the question researchers kept asking us: How do I know I can trust this dataset? And that's a fair question. When you use LLMs to extract clinical information from EHRs at scale, the potential is enormous – but so is the responsibility to demonstrate that what comes out the other end is actually accurate and reliable.

What we noticed was that existing approaches to data quality tended to be piecemeal. You might check individual variables against a reference set, or you might make sure there were no obvious logical errors in the data – but there wasn't a comprehensive, repeatable framework that put all of those dimensions together in a structured way. The VALID Framework, published in JCO Clinical Cancer Informatics earlier this year, was designed to fill that gap.

We needed something transparent, reproducible, and rigorous enough to stand up to scrutiny for the kinds of decisions our data informs – regulatory submissions, comparative effectiveness analyses, external control arms, and research that ultimately informs clinical decision-making. Those are high stakes applications, and the rigor of the framework needed to reflect that.

For organizations considering the use of LLM-extracted datasets, what are the core principles of the VALID Framework, and how does it extend beyond traditional data quality assessment approaches?

The VALID Framework has three pillars that are designed to work together cohesively to assess data quality.

  1. Variable-level performance metrics – how accurately does the LLM extract specific clinical variables compared to expert human abstractors? That gives you a granular, measurable benchmark.
  2. Verification checks – are there internal inconsistencies or implausible patterns in the data itself? These are the kind of things that might slip past a variable-level accuracy check but still indicate something is off.
  3. Replication and benchmark analyses – does the LLM-extracted data actually reproduce established clinical findings? That last piece is, in some ways, the most meaningful, because it asks whether the data behaves the way we'd expect it to based on what we already know from expert human-abstracted data or clinical literature.

What distinguishes the VALID framework from more traditional approaches is that it's not only examining whether we got a number right – it's looking holistically at whether the data is fit for the scientific and regulatory decisions it will be used to support. Those are different questions, and both matter.

In this study, you applied the framework to an LLM-extracted prostate cancer dataset. Could you outline the validation approach and explain why prostate cancer provided a useful test case for assessing LLM-extracted data?

Prostate cancer is one of the most clinically complex disease areas to work in from a RWD perspective, which is part of what made it such a meaningful test case. The treatment landscape has evolved rapidly, with approvals of androgen receptor pathway inhibitors in the early disease setting, PARP inhibitors, and more recently AKT inhibitors for androgen pathway modulation-naive/sensitive (APMN/S; formerly referred to as HSPC). Radioligand therapies are also moving earlier in the disease trajectory. As a result, the clinical distinctions that matter most for research, such as whether a patient is APMN/S or androgen pathway modulation-resistant (APMR; formerly referred to as CRPC), and how lines of therapy are sequenced, require careful, nuanced extraction from the patient record.

We applied the VALID framework to our prostate Panoramic dataset of nearly 400,000 patients to assess whether the data was accurately capturing exactly those kinds of clinical distinctions that are both critical to research and difficult to extract reliably from unstructured records. For variable-level metrics, we used doubly-abstracted test sets of 349–500 patients and compared LLM output against expert human abstraction on variables like initial and metastatic diagnosis and date, and castration-resistant versus hormone-sensitive status. For verification checks, we examined the proportion of patients receiving more than one line of systemic therapy in the metastatic (m) APMN/S setting. And for replication and benchmark analyses, we looked at real-world overall survival (rwOS) in treatment-selected cohorts – specifically patients receiving ARPIs in first-line mAPMR, and PARPi in the second-line setting.

What were the key findings regarding the accuracy of the extracted variables, and were there particular clinical concepts that proved more or less challenging for the model to identify reliably?

The results were really encouraging. Across all three key clinical variables, the LLM performed within a very narrow margin of expert human abstractors. For initial diagnosis and date, the F1 score (a standard metric that balances how often the model is correct with how often it captures everything it should) was 2.10 percentage points lower than human abstraction; for metastatic diagnosis and date, 2.11 percentage points lower; and for APMR versus APMN/S, just 0.52 percentage points lower.

The verification check showed the LLM-extracted dataset had a 3.3% higher proportion of patients receiving more than one line of therapy in the mAPMN/S setting compared to the human-abstracted comparator – a small but meaningful difference that the framework was specifically designed to surface. The VALID Framework did exactly what it was designed to do: flag variations that warrant closer examination.

The replication results were arguably the most compelling. In treatment-selected cohorts, the LLM-extracted and human-abstracted datasets produced virtually identical rwOS estimates – median rwOS of 25.3 versus 24.4 months for first-line ARPI in mAPMR, and 15.8 versus 15.9 months for second-line PARPi. That level of concordance gave us confidence that the dataset can support the kinds of downstream research questions it was built to answer.

The VALID framework emphasizes evaluating performance across different patient subgroups and data contexts. Did you observe any meaningful variation in accuracy across patient populations, disease characteristics, or clinical settings?

This poster was focused on demonstrating the overall framework application and establishing baseline validation metrics for the dataset, so we didn't report granular subgroup-level breakdowns in this particular study. With that said, understanding how performance varies across patient populations – including by race, community versus academic setting, or treatment groups – is something we continuously investigate. Internally, we use VALID as the structure for how we ask and investigate these questions systematically and repeatedly over time, not just once at dataset launch. We know that if a dataset performs well on average but has meaningful gaps for specific subpopulations, that matters enormously for the equity of the research it enables. That's an active area of focus for us.

How should researchers interpret these findings when determining whether an AI-extracted dataset is fit-for-purpose for a specific research question?

Fitness-for-purpose and data quality are two related but distinct questions. Data quality asks if the information in the dataset is accurate, internally consistent, and clinically plausible. VALID is designed to answer these questions. Fitness-for-purpose is a separate judgment that researchers have to make for their specific study: is a dataset the right tool for a specific research question – does it have the right variables, the right population, enough follow-up time to capture the outcomes you care about? You need both, but they're not the same thing, and it's worth being clear about which one you're evaluating.

Looking ahead, how do you see validation frameworks such as VALID evolving as LLMs become more sophisticated, and what additional evidence will be needed to support broader adoption of AI-generated RWD?

My view is that the work of validation never really ends – and that's not a limitation; that's actually the point.

The value of a framework like VALID isn't just what it tells you today about a specific dataset. It's that it creates a repeatable, transparent process for continuously reassessing data quality as models evolve, as new variables are added to a dataset, and as the research questions we're trying to answer become more complex.

But what I hope for most is that the field continues to move toward greater transparency – not just Flatiron publishing its validation approach, but a broader norm where anyone generating or using LLM-extracted RWD can point to rigorous, publicly available evidence for how the data was built and evaluated. That's what it will take to build the durable trust this kind of evidence needs.