![]()

Key Takeaways
- Claims data offers unmatched scale and longitudinal reach across payer networks, but lacks the clinical depth needed for nuanced outcomes research.
- EHR data provides rich clinical context – vitals, lab results, physician notes – yet is often fragmented across health systems and harder to standardize at scale.
- The strongest real-world evidence (RWE) strategies typically integrate both datasets, combining cost visibility with clinical granularity to produce insights neither source can deliver alone.
- Choosing the wrong data source for a given research question doesn’t just slow programs down – it can distort findings that inform payer negotiations and regulatory submissions.
- Understanding where each dataset fits in the drug development and market access lifecycle is the foundation of any credible RWE program – and where expert advisory support pays dividends.
Two Datasets, One Critical Decision
Every real-world evidence program in life sciences ultimately comes down to a foundational question: which data tells the story that needs to be told? Claims data and Electronic Health Records (EHR) data are both powerful, but they are built for different purposes – and using the wrong one for the wrong question carries real consequences.
The stakes have risen sharply as regulators, payers, and Health Technology Assessment (HTA) bodies now expect robust real-world evidence alongside traditional clinical trial data. The FDA has formally recognized RWE as scientifically valid for certain regulatory decision-making, including its use in post-market surveillance and label expansion decisions. That shift means dataset selection is no longer a back-office data science decision – it’s a strategic one.
What follows is a practical breakdown of how each dataset works, where each one excels, where each falls short, and how integrating both unlocks the strongest evidence base available.
Claims Data: Scale and Longitudinal Reach
What Claims Data Actually Captures
Claims data is generated every time a healthcare provider submits a billing record to an insurer for reimbursement. It’s transactional by design, and that design produces one of its greatest strengths: a consistent, time-stamped record of every reimbursable encounter a patient has across different providers, settings, and care episodes.
A standard claims record captures:
- Patient demographics – age, gender, geographic location, insurance type
- Diagnoses – ICD-10 codes assigned at each encounter
- Procedures – CPT and HCPCS codes reflecting services rendered
- Medications – NDC codes, dose, prescription fill dates and refill patterns
- Costs – total billed amounts, payer reimbursement, and patient out-of-pocket responsibility
Large aggregators compile this data across millions of patients. IBM MarketScan databases, for instance, contain claims information on over 240 million US patients drawn from employers, health plans, and government programs. That kind of scale makes claims data the default choice for population-level analyses, treatment pattern studies, and health economics research.
Where Claims Data Falls Short
The same billing-first architecture that makes claims data so consistent also limits what it can tell researchers. A few limitations that frequently trip up life sciences teams:
- Diagnosis codes aren’t confirmatory. A provider may submit a diagnosis code during a workup, not after a definitive diagnosis – meaning the code reflects a possibility, not a confirmed condition.
- No clinical context. Lab values, vital signs, disease severity, and physician rationale for treatment decisions are simply absent.
- Medication adherence is inferred, not observed. A filled prescription doesn’t confirm a patient took the medication as directed.
- Adjudication lag. Claims data is often delayed weeks or months due to billing cycles, making it a poor source for real-time insights.
- Population bias. Uninsured patients – roughly 8% of the U.S. population – and those who pay out of pocket don’t appear in claims at all.
Relying solely on claims can lead teams to misread critical market dynamics, particularly when patient behavior or clinical nuance is central to the research question.
EHR Data: Clinical Depth Over Breadth
The Clinical Context Claims Can’t Provide
Electronic Health Records are built to support patient care, not billing. That difference in purpose produces a fundamentally richer dataset. EHR records document the full clinical picture of a patient encounter: vital signs, lab and pathology results, radiology imaging, medication orders with dosing instructions, physician progress notes, discharge summaries, and documented comorbidities.
That clinical depth opens research possibilities that claims simply can’t support:
- Identifying patient cohorts based on specific lab thresholds, biomarker status, or disease severity
- Tracking longitudinal disease progression and treatment response within a clinical narrative
- Applying natural language processing (NLP) to unstructured physician notes to extract data not captured in structured fields
- Linking lab results, pathology reports, and imaging to create richer patient phenotypes
EHR data also enables retrospective cohort identification with a level of clinical specificity that billing codes cannot replicate – a capability that directly improves the rigor and efficiency of study design. AI-powered cohorting tools, including large language models and Retrieval-Augmented Generation (RAG) systems, have made it possible to identify complex patient populations from EHR data at scale and with greater precision than manual chart review.
EHR Limitations Life Sciences Teams Overlook
EHR data’s clinical richness comes with real structural challenges that are easy to underestimate:
- Fragmentation across networks. If a patient receives care at multiple unlinked health systems, their full clinical picture may be split across disconnected EHR databases – or missing entirely from any single source.
- Unstructured data burden. Up to 80% of EHR data exists in unstructured form – written notes, free-text fields – requiring NLP and significant data science infrastructure to make usable.
- Inconsistent documentation standards. There are no national standards governing how EHR fields are used, meaning the same data element may be recorded differently across Epic, Cerner, and other systems.
- Smaller, geographically constrained samples. Unlike national claims databases, most EHR datasets are anchored to specific health systems and may not generalize to broader populations.
- Absence of data is ambiguous. A gap in the EHR record may mean a patient didn’t seek care, or simply that they sought it outside the network – there’s often no way to distinguish the two.
Head-to-Head: When to Choose Which
Use Cases That Favor Claims Data
Claims data is the stronger choice when the research question centers on population scale, cost, or utilization patterns across diverse care settings.
- Epidemiological studies – incidence rates, prevalence, and risk factor analyses across large populations
- Treatment pattern research – how drugs are being prescribed, switched, or discontinued across payer types in the real world
- Medication adherence – pharmacy claims provide reliable fill dates and refill patterns over extended periods
- Health economics and outcomes research (HEOR) – total cost of care, resource utilization, and payer-specific cost comparisons
- Post-market safety surveillance – longitudinal cross-network tracking of adverse events and healthcare utilization following exposure
Use Cases That Favor EHR Data
EHR data earns its advantage when clinical granularity, disease biology, or treatment outcomes are central to the question.
- Cohort identification based on clinical criteria – specific lab values, biomarker status, disease staging, or comorbidity combinations
- Comparative effectiveness research – understanding how a treatment performs across patient subgroups defined by clinical, not administrative, characteristics
- Safety signal investigation – cases where understanding the clinical context around an adverse event matters as much as the event itself
- Clinical trial optimization – using retrospective EHR data to refine eligibility criteria, estimate enrollment feasibility, and define endpoints before a trial launches
Integration: The Strongest RWE Strategy
How Linked Data Unlocks Deeper Insights
Neither dataset alone tells the complete story of a patient’s healthcare journey. Claims data traces the path across the system; EHR data illuminates what was happening clinically at each stop along that path. Together, they close the gaps each leaves open.
Linked claims-EHR approaches are being used in practice today. A study examining metastatic breast cancer patients used linked data from Flatiron Health and Komodo Health to understand how comorbid conditions affected survival outcomes – a question claims codes could flag but not clinically characterize, and one EHR data alone couldn’t answer at sufficient scale.
The most common integration methods include:
- Patient-level linkage via tokenized or probabilistic matching across datasets
- Using claims to build initial patient cohorts, then pulling EHR records for deep clinical review
- Combining pharmacy fill data from claims with EHR medication orders to assess real-world adherence
- Pairing claims-based cost data with EHR clinical outcomes for HEOR analyses
The key operational challenge is data quality. Most EHR data requires substantial cleansing, harmonization, and NLP processing before it can be meaningfully analyzed alongside structured claims. Life sciences teams that underestimate this preparation burden often find their timelines extended and their insights delayed.
RWE Strategy: A Cornerstone for Market Access Success
Payers and HTA bodies are asking harder questions than they were five years ago. Demonstrating that a therapy performs effectively in real-world patient populations – not just in a controlled trial – has become a core requirement for formulary access, reimbursement negotiations, and value-based contracting.
RWE built on well-chosen, well-integrated data supports value propositions that speak directly to payer priorities: cost-effectiveness, quality of life improvements, and reduced healthcare utilization over time. It also supports regulatory submissions, label expansions, and safety monitoring programs that sustain a product’s commercial life well past initial approval.
The teams that build durable market access positions aren’t choosing between claims and EHR– they’re designing programs that use each source where it’s strongest, integrate them where warranted, and match their data strategy to specific evidence gaps rather than defaulting to what’s most familiar or most available.
MEDDDICAL
Aptos 221
Edificio D2C
Sotogrande
Cadiz
11310
Spain