Faculty Spotlight: Fatema Shafie Khorassani, PhD

Dr. Fatema Shafie Khorassani is an Assistant Professor of Biostatistics at the Boston University School of Public Health. Dr. Shafie Khorassani’s work focuses on data integration methods for time-to-event outcomes, and statistical methods for surrogate endpoints. Her other research interests include statistical methods for data integration, survival analysis, and causal inference for observational dataHer applied work uses biostatistics to study health disparities and cancer using complex data sources.  

Read on to learn more about Fatema and her work!

Can you tell us about your academic and professional background, and how it led you to where you are now?

I majored in math in undergrad at Wayne State University in Detroit, Michigan. After graduating, I taught math for a while, and worked as a medical assistant before going back to Wayne State for my MPH, with a concentration in biostatistics. After working as an applied biostatistician for a while, I realized I missed teaching, and wanted to have more time for independent research, so I decided to get a PhD in Biostatistics, which I completed at the University of Michigan in 2023, before joining the biostatistics faculty at the Boston University School of Public Health.

Some of your research focuses on data integration methods. Can you explain what that means and how you’ve applied it to health research?

Data integration is a statistical process for combining information from multiple data sources, without necessarily linking individual level data. This shows up in a lot of different areas of public health research. Some examples include: (1) combining small studies with rich outcome or covariate information with a larger dataset that does not include all of those variables but can improve the efficiency and precision of parameter estimates, (2) adjusting an analysis for unmeasured confounders that are available in an alternate data source, and (3) incorporating summary level information from multiple external data sources that have disparate covariate information. In my research I have developed data integration approaches for time-to-event outcomes motivated by studying racial disparities in cancer mortality adjusted for confounders that are only collected in an external data source. The main benefit of data integration methods is that they leverage existing data sources to answer research questions that we would not have otherwise been able to answer.

What advice would you give to students and trainees interested in the field of health data science?

Spend time understanding both the theoretical underpinnings of the methods you’re using, and the literature in the health area you’re applying it to. A theoretical method is most impactful when it is motivated by an important research question.

What is something about you that others would be surprised to learn?

I have a serious book hoarding problem!