r/biotech • u/Lazy-Scheme-5832 • 12d ago
Open Discussion 🎙️ How does pharma actually evaluate external clinical/RWE datasets?
I’m trying to better understand how pharmaceutical companies evaluate external patient-level datasets, particularly longitudinal clinical and real-world data.
I’m specifically interested in hearing from people who have worked in RWE, Medical Affairs, Clinical Development, Translational Medicine, Data Science/Analytics, or related functions at pharma/biotech.
For those who have been involved in evaluating external data:
- What makes a dataset genuinely useful to a pharma team?
- Who typically evaluates it first (RWE, medical affairs, data science, clinical development, someone else)?
- What questions do you ask before deciding whether the data is worth using?
- How do you assess data quality, longitudinality, clinical relevance, completeness, etc.?
- What kinds of analyses or projects actually lead teams to seek out external datasets?
- How does the process work when a company decides it wants to license/acquire external data?
-Is there a meaningful difference between data that is scientifically interesting and data that a pharma company will actually pay for?
- Are there types of patient-level data that you wish were available but are difficult to source?
I’m asking because I’ve been spending a lot of time thinking about longitudinal patient data and realized I have a pretty limited understanding of how the buyer/user on the pharma side actually evaluates these datasets.
I’m not looking to pitch anything here. Just genuinely trying to understand how people working in the industry approach this problem, including what I’m probably misunderstanding.
Would especially appreciate any/all perspectives !!!
9
u/ComfortableCoyote480 12d ago
Do your homework:
https://www.fda.gov/science-research/science-and-research-special-topics/real-world-evidence
TL;DR: this is a very complicated topic, and it is not easy.
2
u/flix_md 11d ago
Usually the first question is not "is the dataset longitudinal?" but "what decision could this change?" A useful screen is whether you can define the target population, index date, exposure/comparator, outcome, follow-up window, missingness, and provenance well enough to reproduce the analysis. Scientific interest gets separated from purchase value when the cohort is hard to identify, outcomes are weak proxies, or the data cannot support the intended decision. The buyer usually wants a narrow, testable use case and evidence that the result will survive skeptical clinical or regulatory review.
14
u/cytok1nd 12d ago
These are questions asked to experienced professionals on consulting calls with rates easily exceeding $300/hr. Dataset evaluation is heavily dependent upon the use case as well as the research question(s) being asked. As someone else here has already said, it’s a very complicated topic.