Five published studies each built their own definition of peripheral artery disease as a standalone patient cohort. Replicated against the same national claims database, those definitions identified anywhere from 5.15 to 6.91 million patients.
The reason isn't code list length. The shortest definition ran to a few dozen codes and the longest to several hundred, yet the long lists barely found more patients. What separated them was a single code that some lists left out, and it accounted for about a third of the largest cohort.
That is the kind of variance most evidence generation teams never see. It's also the subject of Precision or Illusion? Solving the definition-chaos crisis in real-world evidence, a four-part webinar series from PurpleLab® and Navidence. The series is part of PurpleLab CARES, our newly launched Center for Advancing Real-World Evidence Studies.
Clinical definitions for disease conditions and patient cohorts are inconsistent and unstandardized, and that imprecision carries through to the data, the AI models built on it, and the conclusions drawn from both. Two teams can study the same condition, follow accepted practice, and build cohorts that share only a fraction of the same patients.
"When two teams define the same condition differently, they can arrive at completely different answers, and most never realize it," said Scott DuVall, PhD, SVP, Real-World Evidence at PurpleLab. "This series makes this problem visible, and then shows what precision actually looks like when validated definitions are deployed against real-world data at scale."
The consequences land squarely on the teams doing the work. A market access dossier that another analyst can't replicate. A launch forecast built on the wrong denominator. An HEOR finding that shifts when a reviewer swaps in a different code list. The definition is upstream of all of it, and it's the piece least likely to get scrutinized.
This is shifting from best practice to expectation. FDA's real-world data guidance asks for the full operational definition in the protocol and study report, including the coding system, the rationale, and the limitations, and is explicit that how well a definition performs depends on the data source, the population, and the time frame it's applied to.
Each episode is built around an ISPOR 2026 poster and what it means in practice for HEOR, AI, commercial analytics, and medical affairs teams.
Navidence brings the validated phenotype definitions. PurpleLab brings the claims data infrastructure to show what they mean at scale.
"A definition that lives in a methods section can't be checked, reused, or trusted," said Aaron Kamauu, MD, MS, MPH, CEO and co-founder of Navidence. "Founding it on pragmatic validation, making it computable, and ensuring transparency is what lets another team run the same definition in their study and know they're looking at the same patients or outcome."
The definitions are computable and operationalized, with context on how each has been validated and used. The data shows how many real patients fall under each one, how much individual codes weigh, and how cohort demographics shift depending on the definition you choose. Patient-level findings in the series were generated in PurpleLab CLEAR Claims, a US national claims database.
Scott DuVall, PhD, SVP, Real-World Evidence at PurpleLab, leads the PurpleLab CARES initiative, championing the generation of reliable real-world evidence. He actively participates in the Observational Health Data Sciences and Informatics (OHDSI) community where he helps standardize and organize international network studies. He gained his love of acronyms (and the realization that data could be used to provide answers for real patients) as he formerly directed the Department of Veterans Affairs (VA) Informatics and Computing Infrastructure (VINCI) infrastructure and was a professor at the University of Utah School of Medicine.
Aaron Kamauu, MD, MS, MPH, is CEO and co-founder of Navidence, where he leads development of Computable Operational Definitions (CODefs) — expert-curated phenotype and outcome definitions another team can pick up, inspect, and run. He first met this problem as a medical intern rotating through radiology, emergency medicine, the ICU, and cardiology, where the same patient could read as four different patients depending on which service was writing the note. He left clinical practice for informatics in 2006, after graduate training in biomedical informatics and public health alongside medical school. He co-hosts Real World Wednesdays™ (RWW) podcast, a weekly reminder of how much of this work turns on definitions.
Episodes release weekly on Tuesdays through September.
Episode 1: Definitions Matter - September 8
You say ASCVD. But what does your data actually reflect? Why the same term can mean different things across datasets.
Episode 2: Precision Definitions - September 15
"You do not make cheesecake with cheddar cheese. The ingredient list matters." Five published ASCVD definitions, and cohorts that share only about a third of the same patients.
Episode 3: Balance at Scale - September 22
Garbage in, garbage out. Even one code can move the count by millions. Peripheral artery disease, five published definitions, and why code list size tells you almost nothing about what a cohort contains.
Episode 4: The Road Ahead - September 29
What does good actually look like? What the research means for pharma teams, and why definitions matter more as companies lean harder on AI.
The studies land on the same recommendation: publish complete definitions and code lists, which is what regulators are increasingly asking for anyway.
Need real-world data to back your definitions? Reach out to us here.
Want support with precision data definitions? Reach out to Navidence here.
The series draws on companion posters presented at ISPOR 2026 in Philadelphia by researchers at Navidence and PurpleLab. The full posters can be downloaded alongside each new episode.
PurpleLab® is a health-tech company driven by one clear philosophy: outcomes matter most. As your trusted partner for real-world data, we help organizations drive decisive action based on precise insights — with the ultimate goal of giving everyone a fighting chance at the best possible health outcome. purplelab.com
Navidence is a technology company focused on supporting researchers in defining health data through the deployment of Computable Operational Definitions. They help healthcare and life sciences organizations explore, access and select codes lists, data definitions and deploy them to design and assess the use of real-world data across all phases of clinical research through its platform. navidence.com