News

BMSIP Projects 2026

Project PI Student Keywords
Image-to-Transcriptome: Predicting Neuroinflammatory Signatures from Brain Histology in Chronic Traumatic Encephalopathy Dr. Adam Labadorf & Dr. Jon Cherry Reem Rasmy Whole Slide imaging, snRNAseq, Deep Learning, Chronic Traumatic Encephalopathy (CTE)
Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas Dr. Andrey Sharov Nanyu Huang Bulk RNAseq, Epigenetics, Workflow Development, Assembly
Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas Dr. Andrey Sharov Krishen Patel Bulk RNAseq, Epigenetics, Workflow Development, Assembly
Digital Pathology for Kidney Disease Diagnosis Dr. Chao Zhang Qiuguo Tang Transmission Electron Microscopy, Image Analysis, ML models, Kidney Disease
Studying activity dependent transcriptome in ACC and LPFC cortical regions of rhesus monkeys Dr. Ella Zeldich & Dr. Maya Medalla Lauren Anderson Transcriptomics, scRNAseq, Neuroscience
Deep learning in studying cancer epigenomics Dr. Ignaty Leshchiner Dennis Godin Long Read Sequencing, Nanopore, epigenetics cancer
Anthracotic Pigment as a Tissue-Embedded Exposure Footprint in Lung Cancer Dr. Jennifer Beane Nitachakan Obma RNAseq, Spatial Transcriptomics, Whole Slide Imaging, Cancer
Synergistic Effects of Genetic Risk and Social Determinants of Health on Cognitive Decline Associated with Alzheimer’s Disease Dr. Jinying Chen Tungalan Ganbaatar ML Models, Deep Learning, Data Visualization, Alzheimer’s Disease
Immune changes in tumor tissue and lymph nodes associated with aggressive non-small cell lung cancer (NSCLC) Dr. Josh Campbell & Dr. Sarah Mazzilli Vaidehi Gupta Imaging Mass Cytometry, Single Cell and Image Analysis, Cancer
Immune changes in tumor tissue and lymph nodes associated with aggressive non-small cell lung cancer (NSCLC) Dr. Josh Campbell & Dr. Sarah Mazzilli Tazein Shah Imaging Mass Cytometry, Single Cell and Image Analysis, Cancer
Computational Identification of Airway Epithelial Surface Targets from Lung Single-Cell Atlases Dr. Lei Hou Iris Lee scRNAseq, Protein Databases, Interaction and Signaling Databases, Lung Biology
Genomics of mosquito virus small RNAs Dr. Nelson Lau Emily Dunlop RNAseq, Small RNAs, Package / Workflow Development, Mosquito Biology
Mapping cross-disorder genetic risk in psychiatry to molecular changes in the human postmortem brain. Dr. Nikolaos Daskalakis Amalya Murrill Multi-trait GWAS, Polygenic Risk Score Models, Multi-omics Integration, Psychiatric Disorders
Machine learning and modeling of steroid metabolic and catabolic pathways and enzymes Dr. Pinghua Liu Xiaohe Jin Predictive Modeling, Enzyme Catalysis, Human Microbiome
Machine learning and modeling of steroid metabolic and catabolic pathways and enzymes Dr. Pinghua Liu Shu-Wen Yu Predictive Modeling, Enzyme Catalysis, Human Microbiome
Deep-learning-based image feature extraction and integration for spatial transcriptomics in R Dr. Ruben Dries Penelope Varela Dye Data Analysis, Package / Software Development, Spatial Transcriptomics, Image Analysis
Benchmarking high-resolution spatial transcriptomics re-segmentation tools Dr. Ruben Dries Ryleigh Jerome Spatial Transcriptomics, Benchmarking, Package Development
Decoding chromatin accessibility and lineage plasticity in pancreatic neuroendocrine tumors via spatial atac-seq Dr. Ruben Dries Leah Morzenti ATACseq, Software / Package Development, Spatial Analyses, Cancer
Structural-Based Discovery and In Silico Engineering of Large Serine Recombinases for Precision Genome Editing Dr. Samagya Banskota Cam Nowack Genome Editing, Large Serine Recombinases, Structural Homology Workflows (AlphaFold, Foldseek), GPU Computing
Transcriptional programs of APOE–receptor interactions Dr. Uwe Beffert Christine Snow Differential Expression, Public Data Integration, Bulk RNAseq, Deconvolution, Alzheimer’s Disease
Alternative Lengthening of Telomeres (ALT) Genetic Mutations in Pediatric Osteosarcoma Dr. Rachel Flynn Mohammad Gharandouq Genome Sequencing, Phylogenetics, Mutational profile, Machine Learning, Random Forest, Cancer

 

top

Image-to-Transcriptome: Predicting Neuroinflammatory Signatures from Brain Histology in Chronic Traumatic Encephalopathy

PI: Dr. Adam Labadorf & Dr. Jon Cherry top

Biological Background

Chronic traumatic encephalopathy (CTE) is a neurodegenerative disease caused by
repetitive head impacts, most studied in contact sport athletes. CTE is
characterized by distinctive deposits of phosphorylated tau protein around blood
vessels at the depths of cortical sulci, accompanied by neuroinflammation driven
largely by microglia — the brain’s resident immune cells. Recent single-nucleus
RNA sequencing (snRNA-seq) of post-mortem brain tissue from young athletes has
revealed that repetitive head impacts trigger neuronal loss and microglial
activation even before clinical symptoms appear, identifying novel inflammatory
microglial populations associated with injury. However, snRNA-seq is expensive
and destructive, while stained histology slides are routinely generated during
neuropathological evaluation. If molecular signatures could be predicted
directly from histology images, it would dramatically expand our ability to
study CTE biology across large archival brain bank cohorts where sequencing data
will never be available.

Problem Statement

We aim to determine whether deep learning models can predict transcriptomic
features — specifically cell-type proportions and neuroinflammatory pathway
activity — from stained brain tissue images in CTE. This approach is
well-validated in cancer pathology but has never been applied to
neurodegenerative brain tissue, representing a novel application domain. A key
design constraint is that histology images and snRNA-seq data come from opposite
brain hemispheres of the same donor, precluding spatial registration. We will
therefore use sample-level prediction, mapping whole slide images to donor-level
molecular profiles. The 10-week scope targets a proof-of-concept demonstrating
feasibility and identifying which molecular features are most predictable from
brain morphology, providing the foundation for a federal grant proposal to scale
and extend this work.

Data Types and Analytical Methods

The dataset comprises ~240 donors with paired whole slide images (WSIs) and
snRNA-seq profiles. WSIs include multiple immunohistochemical stains: AT8
(phospho-tau), Iba1 and CD68 (microglia), P2RY12 and TMEM119 (microglial
homeostatic markers), and Nissl/LFB-CV (neuronal architecture). Prediction
targets derived from snRNA-seq include pseudo-bulk expression profiles,
cell-type proportions (via deconvolution), and pathway activity scores (via
ssGSEA) for neuroinflammation-related gene sets. The analytical pipeline uses a
pre-trained pathology vision foundation model (CONCH or UNI2) to extract
tile-level feature embeddings from WSIs, followed by attention-based multiple
instance learning (ABMIL/CLAM) with regression heads to predict continuous
molecular targets. Performance is evaluated by Pearson correlation and R² under
5-fold donor-stratified cross-validation. Attention heatmaps provide
interpretability by highlighting tissue regions driving predictions.

top

Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas

PI: Dr. Andrey Sharov top

Biological Background

The naked mole rat Heterocephalus glaber is a unique mammalian model
characterized by exceptional longevity, resistance to cancer, and preserved
tissue homeostasis throughout life. Skin in naked mole rats displays unusual
properties related to epidermal differentiation, wound healing, and tumor
resistance, making it an attractive system to study protective epigenetic
mechanisms. However, unlike mouse or human, the naked mole rat remains a
non-model organism with incomplete genome annotation, limiting the application
of modern epigenomic approaches such as chromatin accessibility profiling and 3D
genome analysis. This project focuses on skin biology and genome regulation,
laying the groundwork necessary to enable future studies of epigenetic control
of development, regeneration, and cancer resistance in this species.

Problem Statement

Large-scale epigenomic analyses require a well-curated reference genome and
transcriptome, which are currently lacking for naked mole rat skin. The absence
of standardized genome builds, refined gene models, and cross-species
annotations represents a major barrier to interpreting RNA-seq, ATAC-seq, and
chromatin conformation data. This internship will address this foundational
problem by evaluating existing naked mole rat genomic resources and generating
an analysis-ready genome and transcriptome framework. The resulting reference
will support the creation of a future naked mole rat skin epigenomic atlas and
enable meaningful comparisons with mouse and human skin.

Data Types and Analytical Methods

The intern will work primarily with bulk RNA-seq data generated from naked mole
rat skin. Analytical methods will include RNA-seq quality control, alignment,
transcript quantification, and identification of expressed genes. The student
will evaluate available naked mole rat genome assemblies and annotations, select
a primary reference, and generate standardized FASTA and GTF files suitable for
downstream epigenomic analyses. Additional analyses will include basic
transcriptome annotation refinement, assessment of repeat content and
mappability, and ortholog mapping between naked mole rat, mouse, and human
genes. Emphasis will be placed on reproducible workflows and biological
interpretation.

top

Digital Pathology for Kidney Disease Diagnosis

PI: Chao Zhang top

Biological Background

Chronic kidney diseases (CKD) affect 13% of the population and costs the
US at least $50 billion annually. Transmission electron microscopy (TEM) remains
gold standard for renal histopathological diagnoses, especially for proteinuric
kidney diseases. Currently, measurements of podocyte foot process width (FPW)
and glomerular basement membrane (GBM) width in TEM images are performed
manually, which limits the accuracy and efficacy of ultrastructural analysis.
Therefore, we propose to develop an AI computational digital biopsy platform to
measure podocyte FPW and GBM width of healthy and pathological kidney specimens
automatically using TEM images from pre-clinical podocytopathy animal models and
clinical kidney biopsy samples from patients with podocytopathy diagnoses.
Specifically, this product will manifest in the form of a web application that
can be accessed online and is categorized as an image analysis software. It
could enhance treatment decisions and expedite drug development in pre-clinical
models and clinical trials, advancing solutions for proteinuric kidney diseases.

Problem Statement

Collaborating with several different labs, we have collected 60000+ patients’
TEM images. The previous model works very well on mouse and rat data but is less
accurate on human data due to a lack of labels. The potential projects during
the summer internship could be: 1) Assisting in labeling, organizing human data,
and improving the deep learning model to achieve better accuracy. 2)
Implementing a Mamba-based model to integrate the TEM images and pathological
reports. This could involve generating pathological report templates from images
to reduce the writing time for pathologists or building new classification
models for disease and drug response prediction. 3) Utilizing data from The
Kidney Precision Medicine Project to explore the potential of multimodal
integration for TEM images, histology images, and omics data Data Types and
Analytical Methods What types of biological data will be used? How will it be
analyzed? The intern must possess strong Python programming skills and
experience in medical image analysis, particularly with libraries such as
OpenCV, scikit-image, numpy and pandas Knowledge of and experience with
analyzing pathological Images is highly desirable. The intern’s primary
responsibility will be to develop and implement algorithms for the processing
and analysis of images, focusing on image segmentation, feature extraction, and
machine learning-based classification. Collaboration with a multidisciplinary
team, including medical professionals, to refine analysis techniques and ensure
clinical relevance. Evaluation of existing analytical methods and adaptation or
development of new approaches to improve accuracy and efficiency. Documentation
of methodologies, code, and analysis procedures for future reference and use by
the research team.

top

Studying activity dependent transcriptome in ACC and LPFC cortical regions of rhesus monkeys

PI: Dr. Maya Medalla and Dr. Ella Zeldich top

Biological Background

In our previous work we profiled transcriptionally and functionally the
differences between the two important brain regions. In the current project, we
will be specifically focusing on the changes related to neuronal activity
induces by cognitive tests. We will look at the population of the cfos-positive
neurons and the changes within this population. The obtained results will be
integrated with our previous data.

Problem Statement

How the regional differences between ACC and LPFC pertain to cognition-related
neuronal activity

Data Types and Analytical Methods

scRNA-seq

top

Computational Identification of Airway Epithelial Surface Targets from Lung Single-Cell Atlases
PI: Lei Hou top 

Biological Background

Targeted delivery of gene-editing or RNA therapeutics to airway epithelial cells
remains a major challenge in cystic fibrosis (CF) and other lung diseases.
Current delivery systems often rely on broadly expressed receptors, resulting in
inefficient uptake by the desired cell types and unintended interactions with
immune or stromal cells. The airway epithelium itself is heterogeneous,
consisting of basal, secretory, ciliated, and rare specialized cell populations,
each with distinct molecular programs and surface protein composition. Recent
advances in single-cell RNA sequencing (scRNA-seq) have generated comprehensive
atlases of the human lung across healthy and disease contexts. These datasets
provide an opportunity to systematically identify cell-type–specific surface
proteins that could serve as molecular “entry points” for targeted delivery.
However, gene expression alone is insufficient to determine whether a candidate
receptor is biologically active or functionally relevant. Integrating additional
layers of information—such as cell–cell communication and regulatory network
context—can help identify receptors that are not only expressed but also
actively involved in epithelial signaling programs.

Problem Statement

Although lung single-cell datasets have revealed extensive epithelial
heterogeneity, there is currently no systematic framework to identify and
prioritize airway epithelial surface targets for therapeutic delivery. In
particular, three key questions remain unresolved: Which receptors and membrane
proteins are robustly and specifically expressed in airway epithelial cell types
across datasets and disease conditions? Among these proteins, which are actively
engaged in cell–cell communication, suggesting extracellular accessibility and
functional relevance? Which receptors are embedded in regulatory programs
linking receptor activation to downstream transcriptional responses, indicating
biological importance within epithelial signaling networks? Addressing these
questions requires integrating single-cell transcriptomics with protein
annotations and network-based inference methods to produce a ranked set of
candidate surface targets.

Data Types and Analytical Methods

Single-cell transcriptomics Public lung scRNA-seq datasets including: CF and
control airway epithelium datasets (e.g., Nat Med 2021) Airway epithelial
differentiation datasets (e.g., Shah et al., Nat Commun 2025) Lung atlases
such as LungMAP and the Human Lung Cell Atlas (HLCA) These datasets provide
cell-type resolution across epithelial subpopulations.

Protein and functional annotation resources Gene Ontology (GO) membrane
annotations UniProt subcellular localization Cell Surface Protein Atlas (CSPA)
Human Protein Atlas (HPA) These resources help translate gene expression signals
into surface protein candidates. Interaction and signaling databases
Ligand–receptor interaction databases (e.g., CellPhoneDB, NicheNet) Protein
interaction databases (e.g., STRING, BioGRID) These provide biological context
for receptor activity.

The project will consist of three major analytical components.

1. Identification of Epithelial Surface Proteins from Lung scRNA-seq We will
integrate multiple lung scRNA-seq datasets to identify receptors and membrane
proteins enriched in airway epithelial populations. Key steps include: Dataset
harmonization and batch correction Identification of epithelial cell types and
subtypes Differential expression analysis to identify epithelial-enriched genes
Filtering candidate genes based on plasma membrane annotations and surface
protein databases This analysis will produce a catalog of epithelial surface
proteins, annotated by functional class such as receptors, transporters,
adhesion molecules, and ion channels.

2. Inference of Receptor Activity through Cell–Cell Communication Analysis To
evaluate whether candidate receptors participate in extracellular signaling, we
will infer ligand–receptor interactions between epithelial cells and neighboring
cell populations. Methods include: Ligand–receptor inference frameworks such as
CellPhoneDB or NicheNet Quantification of interaction strength between
epithelial cells and surrounding immune or stromal populations This step helps
identify receptors that are actively involved in extracellular signaling
networks, supporting their accessibility and biological relevance.

3. Inference of Receptor Function through Receptor–TF–Gene Regulatory Networks
To assess downstream signaling consequences, we will construct receptor-centered
regulatory networks linking receptors to transcription factors and target genes.
This analysis will involve: Mapping candidate receptors to known signaling
pathways Integrating protein interaction networks to connect receptors with
transcription factors Evaluating downstream transcriptional programs associated
with receptor activity Receptors embedded in coherent receptor–TF–gene
regulatory modules will be prioritized, as they are more likely to represent
functionally active signaling nodes in epithelial cells.

top

Deep learning in studying cancer epigenomics

PI: Dr.Ignaty Leshchiner top

Biological Background

Computational biology and cancer bioinformatics. Research in the lab is focused on applying new genomic technologies, computational analysis and AI methods on data from patients’ tumors to understand the biology behind tumor development, treatment evasion, and progression to metastasis. We are developing and applying tools for simultaneous analysis of multiple samples from the same patient, clonal structure, integration of single cell genomics and transcriptomics, reconstruction of cell subpopulations, their growth kinetics and expression, tumor micro-environment effects, estimation of order of events (“timing”) during tumor development and progression. We work with pre- and post- treatment samples, autopsies and longitudinal blood biopsies in solid and blood malignancies.

Problem Statement

There is an ever-growing body of literature which show how modifications to the DNA, both genetically and epigenetically, as well as modifications to proteins are distinct in the setting of cancer. Changes to RNA, especially those epitranscriptomic in nature, have not yet extensively been studied. This is because while the technologies to accurately assess the base-pair changes to DNA are well established, the ability to detect native epitranscriptomic changes in DNA and RNA is not yet a robust technology. We are developing deep learning methods that enable calling of methylation modifications from both RNA and DNA with higher accuracy from Native Nanopore based sequencing and identify tumor type specific DNA/RNA modifications. The project will involve analyzing and training the models to improve call accuracy and detect biological changes within samples.

Data Types and Analytical Methods

For this project we will use raw signal processing of Nanopore hdf5 files and converting the current signal into accurate base space and methylation calls. We train the models on standard data we generate in the lab both on cancer samples and cancer/normal cell line models.

top

Anthracotic Pigment as a Tissue-Embedded Exposure Footprint in Lung Cancer

PI: Dr. Jennifer Beane top

Biological Background

This study involves examining the cells that comprise the lung and how they
respond to chronic inhaled pollutants and how this exposure increases the risk
of developing lung cancer and other lung diseases.

Problem Statement

Chronic inhaled exposures, including cigarette smoke and air pollution, promote
lung cancer risk by causing cumulative genomic damage and remodeling the lung
microenvironment. Anthracotic pigment is visible carbonaceous particulate matter
in lung tissue that can bind carcinogens and may serve as a tissue-embedded
exposure footprint, but its spatial patterns and associated molecular programs
across non-cancerous and cancerous tissues remain poorly defined. We will
leverage large public cohorts and AI-based methods for whole slide images to
quantify pigment and define lung cancer-associated pigment subtypes and spatial
niches, enabling development of exposure-informed biomarkers that improve risk
stratification beyond self-reported smoking history.

Data Types and Analytical Methods

The types of data used in the project are bulk RNA sequencing data, H&E-stained
whole slide images (WSIs), and spatial transcriptomics data. The project
involves applying state-of-the art cell segmentation and classification methods
to the WSIs and analyzed both the bulk and spatial transcriptomics data to
identify both local, distance-dependent effects and global tumor
microenvironment shifts associated with WSI phenotypes.

top

Synergistic Effects of Genetic Risk and Social Determinants of Health on Cognitive Decline Associated with Alzheimer’s Disease

PI: Dr. Jinying Chen top

Project Summary

Alzheimer’s disease (AD) develops over decades, with amyloid and tau pathology
accumulating before clinical symptoms emerge. During this preclinical stage,
trajectories of cognitive decline vary widely, suggesting that biological
susceptibility and modifiable contextual factors jointly shape risk. Polygenic
risk scores (PRS) capture inherited liability to AD and have been associated
with earlier onset and steeper cognitive decline, yet PRS alone provides an
incomplete account of individual outcomes. Social determinants of health
(SDOH)—including socioeconomic resources, neighborhood context, education, and
social support—may influence cognitive reserve, health behaviors, comorbidity
burden, and access to care, thereby modifying the clinical expression of
underlying AD pathology. However, the combined and potentially interactive
contributions of AD polygenic risk and SDOH to cognitive decline in
biomarker-defined preclinical AD remain insufficiently characterized.

Problem Statement

Understanding how genetic susceptibility and SDOH jointly relate to early
cognitive changes could improve risk stratification and inform more equitable
prevention strategies. This study aims to assess the interaction effects between
genetic risk factors and SDOH on cognitive decline associated with AD by using
statistical methods, data visualization, and machine learning.

Data Types and Analytical Methods Data (e.g., demographics,
genetic, SDOH, clinical, AD biomarkers) from national and international AD
cohorts (e.g., Health and Retirement Study, Alzheimer’s Disease Data Initiative)
will be used for this study. Cognitive decline will be characterized by
statistical methods (e.g., latent trajectory modeling and/or other regression
models) and data visualization methods, and predicted by machine learning (e.g.,
Lasso regression, Random Forest, deep neural networks, SHAP analysis). The
student will be supervised by Dr. Chen and is expected to conduct a significant
piece of work (data preprocessing and analysis using either machine learning or
statistical methods or both) independently for this project. This position
requires strong programming skills in Python and/or R, and prior experience in
developing and evaluating machine learning models and/or conducting statistical
analysis. Knowledge and experience with developing deep learning models is a
plus. The data analysis will be conducted on BU Shared Computing Cluster (SCC)
and/or the computational platform of Alzheimer’s Disease Data Initiative, with
Python and/or R.

top

Immune changes in tumor tissue and lymph nodes associated with aggressive non-small cell lung cancer (NSCLC)

PI: Dr. Josh Campbell & Dr. Sarah Mazilli top

Project Summary

Regional lymph nodes (LNs) in the thoracic cavity serve as essential
immunological hubs that coordinate humoral and cell-mediated responses against
the development and progression of non-small cell lung cancer (NSCLC). To
investigate immune dysregulation in the non-metastatic regional LNs of patients
with aggressive NSCLC, we performed multimodal profiling on 36 LNs from 11
patients undergoing curative-intent resection including CITE-seq, scRNA-seq, and
Imaging Mass Cytometry (IMC). Regional N1 LNs from patients with more aggressive
disease (stage IB–IIIA) exhibited a significant enrichment of dysfunctional CD8⁺
T cells and regulatory T cells (Tregs) compared to N2 LNs and LNs from patients
with less aggressive disease (stage IA). These immune subsets were spatially
co-localized with mature regulatory dendritic cells (mregDCs; CD1c⁺, TIM3⁺,
LAMP3⁺), forming an immunosuppressive niche uniquely enriched in the N1 LNs of
higher-stage patients. Concurrently, higher-stage N1 LNs contained larger number
of “decorticated” B-cell follicles characterized by decreased encapsulation of
the mantle zone layer surrounding the germinal centers. This mantle zone
disorganization was associated with increased spatial niches involving Tregs,
CD68+ CD163⁺ TIM3⁺ Macrophages, CD163⁺ TIM3dim Monocytic-Myeloid Derived
Suppressor Cells (M-MDSC), plasma B cells, and a decrease in spatial niches
involving CD4⁺ T helper cells and fibroblastic reticular cells (FRCs). Together,
our findings reveal parallel alterations in humoral and cell-mediated immunity
within the regional LNs of patients with aggressive NSCLC.

Problem Statement

While we have analyzed the IMC and single cell data from LNs, we have also
generated IMC data from multiple regions of tumor tissue and adjacent normal in
each patient. We need to have this data analyzed to identify immune populations
and cellular niches in tumors. These findings will be correlated with those we
previously found in the regional LNs. The goal will be to see what aberrant
immune populations in the aggressive tumors are also observed in LNs or are
specific to the tumors.

Data Types and Analytical Methods

Imaging mass cytometry (IMC) which measure levels of cell surface proteins on
tissue slides to maintain spatial architecture. We use computational pipelines
for single cell analysis, image analysis, niche identification, and statistical
methods for associating cell populations and niches with clinical phenotypes.

top

Genomics of Mosquito Virus small RNAs

PI: Dr. Nelson Lau top

Problem Statement
The mosquito Aedes aegypti is the major vector for arbovirus pathogens like
Dengue and Zika viruses. We are also analyzing mosquito tombus viruses that may
infect and compete against Dengue and Zika viruses to generate a small RNA/ RNAi
response. The Lau lab is looking for a BMSIP intern to work on continuing the
analysis of the small RNA responses in mosquitoes and mosquito cells subjected
to infection by tombus viruses. We will be seeing if other mosquito genes are
being affected in RNAi mutants that may lose the capacity to keep mosquito
viruses in check. The intern will learn how to parse the outputs from our
Mosquito Small RNA Genomics pipeline.

Data Types and Analytical Methods

There are many RNAseq and small RNA libraries already generated and sequenced,
and the project will involve working closely with the Lau lab team to organize
and conduct differential expression analysis on mosquito libraries to look for
genes, transposons and virus expression changes. The goal will be to see if the
mosquito virus and RNAi pathways are impacting gene expression in the Aedes
aegypti mosquitoes.

top 

Mapping cross-disorder genetic risk in psychiatry to molecular changes in the human postmortem brain

PI: Dr. Nikolaos Daskalakis top 

Biological Background

This project investigates the functional genomics of psychiatric pleiotropy in
humans – the phenomenon where single genetic variants influence multiple
distinct disorders. Traditionally, disorders like PTSD, MDD, and substance use
disorders (SUD) have been studied in diagnostic silos. However, shared genetic
architecture suggests common underlying biological systems, such as dysregulated
neuroendocrine signaling, synaptic plasticity, and neuroinflammatory pathways.
By leveraging deep-phenotyping data from the human postmortem brain, we are
looking at the ‘molecular endophenotypes’ of these disorders. We are
specifically interested in how genetic risk manifests across different brain
regions and cell types (e.g., excitatory neurons vs. glia). This work moves
beyond the ‘one gene, one disease’ model to explore how broad genetic factors –
in line with notions of psychiatric comorbidities and multimorbidities – shape
the biological landscape of the human brain before and after the onset of
clinical symptoms.

Problem Statement

It is well-established that psychiatric disorders share e.g. trauma, mood and
substance use disorders share not only clinical outcomes (phenotype), but also
genomic architecture. However, the translation of this shared genetic risk into
molecular changes within the brain remains a critical knowledge gap. The primary
challenge in precision psychiatry is the ‘missing link’ between an individual’s
genetic liability (what one is born with) against their clinical state (the
actual disease outcome). While we can identify genetic risk through GWAS, we
often do not know if the molecular changes we see in a patient’s brain are the
cause of the disorder or a consequence of living with it (e.g., due to chronic
stress or medication). This internship addresses this by comparing Polygenic
Risk Scores (PRS) – a genetically -derived ‘biomarker’ of disease risk, against
postmortem molecular data. By doing so, we aim to validate the biological
reality of cross-disorder genetics. We are asking: does a high genetic risk for
internalizing disorders create a specific, observable molecular signature in the
brain, regardless of a patient’s clinical diagnosis? Solving this helps us move
toward a biologically-defined classification of mental health disorders rather
than one based purely on clinical phenotypes. As such, the proposed project
bridges genetics and multi-omics by leveraging the latest cross-disorder GWAS
summary statistics (Grotzinger et al., 2025, Nature) to conduct a
training-to-validation study in a postmortem brain cohort (Daskalakis et al.,
2024, Science).

Data Types and Analytical Methods

We will adopt a testing-to-validation computational framework. In the training
stage, multi-trait GWAS summary statistics (Grotzinger et al., 2025)
corresponding to shared genomic factors e.g. (i) internalizing factor (PTSD,
MDD, anxiety disorder) and (ii) substance use factor (cannabis use disorder,
alcohol use disorder, nicotine dependence, opioid use disorder) will be used to
derive polygenic risk scores (PRS) models. In the validation stage, the trained
models will then be applied to the postmortem brain cohort (Daskalakis et al.,
2024), for PRS estimation (on the aforementioned factors) and downstream
association testing. This is a well characterized cohort of N=304 donors across
three brain regions (medial prefrontal cortex, dentate gyrus and central
amygdala) that is heavily examined within the Lab. We hypothesize that these PRS
encoding for shared genetic risk will map to broad, transdiagnostic clinical
phenotypes (e.g. symptom clusters, risk factors, non-psychiatric health
comorbidities) rather than isolated diagnostic categories. Next, we will examine
whether the shared genetic burden also translate to multi-omic (transcriptomic,
methylomic, proteomic) and single-cell molecular changes within the brain. This
will enable comparison of genetic risk vs clinical outcomes on the
molecular/biological effects on the brain. Softwares include PRScs for PRS model
training, PLINK for PRS score estimation, R programming for statistical testing
and differential expression analysis using linear approaches (LIMMA).

top 

Machine learning and modeling of steroid metabolic and catabolic pathways and enzymes

PI: Dr. Pinghua Liu top 

Biological Background

The importance of bile acids in human health has been known for decades. In
recent years, it comes to the realization that the diverse biological roles for
bile acids have been closely related to the human microbiota, especially the
steroid metabolic and catabolic enzymes. In these reactions, hydroxylation and
dihydroxylation reactions, especially the site- and stereo-specificities are
known to be important for their biological activities. In this project, we would
like to systematically evaluate literature to develop a predictive model for gut
bacteria steroid metabolic and catabolic pathways, which will then be linked to
the function of gut microbiota to human health.

Problem Statement

We aim at develop predictive model for steroid metabolic and catabolic enzyme
functions.

Data Types and Analytical Methods

Literature information on various enzymes. We have develop overexpression and
enzymatic catalytic assays for these enzymatic reactions already.

top 

Deep-learning-based image feature extraction and integration for spatial transcriptomics in R

PI: Dr. Ruben Dries top 

Biological Background

Spatial transcriptomics technologies are increasingly paired with
high-resolution histological images, most commonly hematoxylin and eosin
(H&E)–stained sections. These images capture rich morphological information
about tissue architecture, cellular organization, and pathological features that
are not directly encoded in gene expression data alone. In cancer and other
complex diseases, histological context, such as stromal organization, tumor
boundaries, or immune infiltration, play critical roles in disease progression
and treatment response. Recent advances in deep learning have made it possible
to extract informative image features from histological images at multiple
spatial scales. When integrated with spatial transcriptomics data, these
features offer the potential to link morphology with molecular states, enabling
more comprehensive models of tissue organization and function.

Problem Statement

While spatial transcriptomics datasets routinely include associated
histological images, these images are often underutilized in downstream
analysis. Existing methods for image feature extraction and multimodal
integration are fragmented, difficult to reproduce, or require leaving the R
ecosystem entirely. Indeed, most existing approaches rely on Python-based
pipelines and are not yet natively accessible within R-based spatial analysis
frameworks commonly used by biologists. As a result, many spatial studies miss
the opportunity to systematically incorporate morphological context.

This internship aims to explore R-native implementations for deep-learning and
ultimately create a blueprint for establishing a robust and extensible baseline
for image-based analysis within the Giotto Suite. Specifically, the project will
explore how different deep-learning models and image tiling strategies affect
downstream biological interpretation, and how image-derived features can be
meaningfully integrated with gene expression data. The goal is not to build a
single “best” model, but rather to define practical, well-documented utilities
that enable users to explore image–expression relationships in a standardized
and reproducible way.

Data Types and Analytical Methods

The project will primarily use spatial transcriptomics datasets (e.g., Xenium,
MERFISH, Visium HD) with associated H&E images, drawn from public datasets and
internal examples. Key analytical components include: Image feature extraction:
Testing and extending existing Giotto utilities to support multiple pretrained
deep-learning models, and evaluating whether model choice materially affects
downstream analyses. Multi-scale tiling strategies: Exploring approaches that
combine large, medium, and small image tiles to capture tissue-, neighborhood-,
and cell-level morphology. Multimodal integration: Evaluating methods to
integrate image features with gene expression, including joint latent spaces and
correlated feature representations. Documentation and tutorials: Creating a
user-facing tutorial demonstrating image-based workflows within Giotto Suite.
All development and analysis will be performed primarily in R, with a focus on
making advanced image analysis natively available to Giotto users. Ideal
candidates have some previous ML/DL experience and are interested to work at the
interface of data analysis and package/software development.

top 

Benchmarking high-resolution spatial transcriptomics re-segmentation tools

PI: Dr. Ruben Dries top

Biological Background

High-resolution spatial transcriptomics (ST) platforms, such as Xenium and
MERFISH, have revolutionized our ability to study tissue biology by providing
subcellular localization of RNA molecules. By mapping transcripts within their
native architecture, researchers can identify distinct cellular niches,
metabolic gradients, and complex cell-cell communication networks. These
technologies are critical for understanding how the spatial organization of
tissues—such as the tumor microenvironment or the layered structure of the
brain—dictates biological function and disease progression.

Problem Statement

The accuracy of all downstream spatial analysis is dependent on cell
segmentation: the process of defining cellular boundaries. However, current
image-based segmentation often suffers from “transcript leakage,” where RNAs
from one cell are incorrectly assigned to a neighbor. This noise confounds
differential expression analysis, masks rare cell states, and generates false
positives in ligand-receptor signaling studies. While new algorithms (e.g.
RNA2seg & others) offer post-hoc refinement using transcript density or
statistical demultiplexing strategies, they are currently fragmented, not
implemented in downstream analysis pipelines, or tested on real-world datasets.
This internship addresses the lack of a standardized, user-friendly pipeline to
correct these errors, which currently limits the biological reliability of
high-resolution ST data.

Data Types and Analytical Methods

This project will utilize high-resolution ST datasets (Xenium, MERFISH, and
CosMx) from both public repositories and in-house experiments. The primary
analytical goal is to operationalize and quantitatively benchmark
re-segmentation workflows within the Giotto Suite, an R-based framework
developed in the Dries lab. Methodologically, the project involves: Software
Integration: Developing scripts to connect existing transcript-based refinement
and statistical deconvolution methods into existing Giotto pipelines.
Benchmarking: Applying quantitative metrics (e.g., silhouette scores, cluster
purity) to compare standard segmentation against refined outputs. Validation:
Evaluating the “signal-to-noise” recovery in downstream tasks, specifically
focusing on the clarity of cell-type markers and the accuracy of spatial
neighborhood analyses.

top 

Decoding chromatin accessibility and lineage plasticity in pancreatic neuroendocrine tumors via spatial atac-seq

PI: Dr. Ruben Dries top

Biological Background

Pancreatic neuroendocrine tumors (PanNETs) are characterized by significant
clinical heterogeneity and the ability of tumor cells to shift between different
phenotypic states, a process known as lineage plasticity. While transcriptomic
technologies have highlighted these states, the spatial context of the
underlying chromatin landscape, including the regulatory regions such as
promoters and enhancers, remains largely unexplored. Spatial ATAC-seq (Assay for
Transposase-Accessible Chromatin using sequencing) allows for the mapping of
open chromatin regions directly within the tissue architecture. Through a
collaboration with AtlasXomics, the Heaphy (BU) and Singhi (UPMC) labs, we have
applied this for the first time to Formalin-Fixed Paraffin-Embedded (FFPE)
tissue, the gold standard for clinical samples. This approach enables us to
study high-quality patient specimens through a new epigenomic lens. By capturing
the epigenetic state in situ, we can observe how the spatial organization of the
tumor microenvironment influences gene regulation and facilitates the transition
between cell lineages.

Problem Statement

The primary challenge in understanding PanNET progression is identifying the
specific epigenetic drivers that trigger lineage transitions. Current analysis
pipelines for spatial chromatin data are still in their infancy and often fail
to link spatial accessibility directly to local and spatially-informed Gene
Regulatory Networks (GRNs). This internship addresses this gap by focusing on:
Regulatory Mapping: Identifying changes in promoter and enhancer accessibility
that correlate with lineage plasticity. GRN Modeling: Building spatially aware
gene regulatory networks to identify the transcription factors driving
phenotypic switches. Tool Development: Implementing these spatial chromatin
analysis workflows within the Giotto Suite framework (www.giottosuite.com).

Data Types and Analytical Methods

The project will utilize high-resolution spatial ATAC-seq data from human PanNET
FFPE samples. Methodologically, the project involves: Giotto Integration:
Developing and benchmarking new modules, inspired by ArchR, within the Giotto
Suite (an R-based framework developed in the Dries lab) specifically for spatial
chromatin data. Regulatory Analysis: Analyzing differential accessibility at
promoters and distal enhancers to identify regulatory elements associated with
specific cellular niches. Network Inference: Using accessibility patterns and TF
motif enrichment to reconstruct GRNs that define lineage plasticity. Secondary
Analysis (Optional): Testing computational methods to infer CNV profiles to
track clonal evolution across the tissue. Ideal candidates should have at least
minimal hands-on experience in sequence-level genomics, such as analyzing
(single-cell) ATAC-seq, ChIP-seq, CUT&RUN, GRO-seq, TT-seq, etc.

top

Structural-Based Discovery and In Silico Engineering of Large Serine Recombinases for Precision Genome Editing

PI: Dr. Samagya Banskota top

Biological Background

Large Serine Recombinases (LSRs) are a class of enzymes derived from
bacteriophages and other mobile genetic elements (MGEs) that facilitate
site-specific integration of large DNA payloads into host genomes. Unlike
CRISPR-Cas systems, which create double-stranded breaks, LSRs catalyze highly
efficient, one-way recombination between distinct attachment sites. This makes
them indispensable tools for genome engineering, particularly for therapeutic
gene insertion in mammalian cells. This project shifts the focus from
traditional sequence-based identification toward structural conservation and
biophysical properties of these enzymes in order to identify improved protein
scaffolds. Protein engineering can further improve these novel scaffolds to
address a need for higher specificity and efficiency in current genetic therapy
approaches.

Problem Statement

Past LSR discovery relied heavily on sequence-based homology (HMMs), which often
fail to identify highly divergent LSRs that maintain a similar functional
structure. This creates a “blind spot” in the known landscape of recombinases.
Furthermore, many naturally occurring LSRs lack the necessary specificity or
efficiency for clinical use, or they target “safe harbor” sites that are not
therapeutically relevant. This internship will address these limitations by
advancing existing enzyme discovery pipelines with a structure-based homology
search and attempt in silico protein engineering to refine feature spaces for
improved efficiency, specificity, and the re-targeting of LSRs to
therapeutically

Data Types and Analytical Methods

The project will utilize several gigabytes of publicly available whole-genome assembly data from
NCBI, the Sequence Read Archive (SRA) and protein structure data from the ESM Metagenomic
Atlas. Analytical methods will transition from HMMER-based sequence searches to structural
homology workflows using FoldSeek and AlphaFold/ESMFold to identify candidates based on 3Di
embedding architecture rather than nucleotide sequence.

Key analytical steps include:

1. Seed-Based Refinement: Using high-confidence, improved LSR structures as seeds to
iteratively refine the structural feature space of our search.

2. Computational Re-targeting: Applying machine learning or physics-based modeling to predict
mutations that allow LSRs to recognize and integrate into specific, non-canonical att sites.

3. Workflow Management: All analysis will be performed using a Snakemake-managed pipeline,
utilizing Python for data wrangling and MAFFT for sequence/structural alignment validation.

 top

Transcriptional programs of APOE–receptor interactions

PI: Dr. Uwe Beffert top

Biological Background

The APOE ε4 allele confers the strongest genetic risk for late-onset Alzheimer’s disease and
modulates neuroinflammatory and neuronal vulnerability pathways. ApoE binds several lipoprotein
receptors including APOER2 (Apoer2 in mouse), which undergoes alternative splicing. In our near-
submission manuscript, we demonstrate that inclusion vs deletion of Apoer2 exon 19 is a key
determinant of APOE-dependent signaling in the brain. Using four mouse lines combining Apoer2
exon 19 inclusion/deletion with humanized APOE3/APOE4 backgrounds, we show that Apoer2
splicing is the primary driver of hippocampal gene expression variance, with APOE genotype exerting
strong context-dependent effects. RNA-seq and pathway analyses highlight
inflammatory/immune/cilia programs in APOE4 contexts and antiviral/RNA quality control programs in
exon 19–deleted contexts. Follow-up imaging suggests ApoE genotype and Apoer2 splicing also
affect glial features and primary neuronal cilia morphology.

Problem Statement

We have already performed major RNA-seq analyses for the manuscript, including principal
component analysis, multiple differential expression contrasts, gene set/pathway enrichment, and
selected follow-up validations (some figures may still contain placeholders pending final data).
However, there are additional analyses that could significantly strengthen the paper and generate
valuable mechanistic hypotheses, particularly by:
1. explicitly modeling this dataset as a 2×2 factorial design (ApoE3 vs ApoE4) × (Apoer2
+Ex19 vs ΔEx19) to identify interaction genes that define context-dependent ApoE effects;
2. inferring which cell types (neurons vs astrocytes vs microglia vs oligodendrocyte lineage vs
endothelial) are most likely responsible for observed expression programs in bulk
hippocampus;
3. extending the work from “lists of DE genes” to network/module-level interpretation; and
4. comparing signatures to publicly available brain/cell-type datasets relevant to AD,
neuroinflammation, interferon/antiviral programs, and cilia biology.

Data Types and Analytical Methods

Data: Bulk hippocampus RNA-seq from four mouse genotypes: humanized ApoE3 or ApoE4 crossed
with Apoer2 exon 19 inclusion (+Ex19) or deletion (ΔEx19). Metadata include genotype labels and,
where available, sex, batch, and sample QC metrics.
Work already completed (foundation): The manuscript currently includes core analyses such as
PCA/variance structure, multiple DE contrasts (including ApoE effects within an Apoer2 splice
background and vice versa), overlap summaries (e.g., Venn-style comparisons), and pathway
enrichment summaries; some validation figures are in progress.
Internship extensions (new deliverables):
1. Factorial DE modeling: DESeq2 (or equivalent) using a full design with main effects and an
interaction term (ApoE × Apoer2 splicing) to identify genes with true context-dependent
regulation.
2. Cell-type inference: Enrichment/deconvolution approaches using public mouse hippocampus
single-cell/single-nucleus references to assign DE/interaction signals to likely cell types.
3. Gene network/module analysis: co-expression modules (e.g., WGCNA) and module–trait
associations (genotype, inferred cell-type scores, and available phenotypes).
4. Public dataset comparison: evaluate concordance with published AD-relevant signatures
(microglial activation/DAM, interferon/antiviral programs, cilia gene sets, astrocyte reactivity),
and summarize which signatures map to which genotype contrasts.
Outputs: a reproducible repo (scripts + README), ranked gene lists for main effects/interaction,
module summaries, and manuscript-ready figures/tables.

 top

Alternative Lengthening of Telomeres (ALT) Genetic Mutations in Pediatric Osteosarcoma

PI: Dr. Rachel Flynn top

Biological Background

This project focuses on telomere maintenance mechanisms in pediatric osteosarcoma, a rare and aggressive primary bone cancer. Specifically, it examines the alternative lengthening of telomeres (ALT) pathway, which accounts for cellular immortalization in approximately 75% of pediatric osteosarcoma cases.

ALT relies on homologous recombination rather than the enzyme telomerase to maintain telomere length. Central to this work is the gene TCAB1 (also known as WRAP53), an RNA chaperone that is essential for telomerase holoenzyme assembly and localization.

TCAB1 partially overlaps with the TP53 tumor suppressor gene on chromosome 17p13.1, and structural variations (SVs) disrupting TP53 in osteosarcoma may simultaneously inactivate TCAB1, thereby abolishing telomerase activity and potentially driving cells toward ALT activation. Genes such as ATRX, DAXX, SMARCAL1, and SLX4IP are also known contributors to ALT and will be examined in the context of tumor evolution.

Problem Statement

The genetic mechanisms underlying ALT activation in osteosarcoma remain incompletely understood. While mutations in chromatin remodeling genes like ATRX and DAXX are known contributors, they alone are insufficient to fully explain ALT induction. Recent work from the Flynn Lab identified structural variations at the TP53/TCAB1 locus in approximately 40% of ALT-positive osteosarcoma tumors, suggesting that functional inactivation of the telomerase holoenzyme may be an early, previously unrecognized step in ALT activation.

This internship will extend that work by:

(1) Analyzing existing RNA-seq data from both the TARGET-OS cohort and the Flynn Lab's GCCRI PDX cohort to investigate differential gene expression, including oxidative stress pathways, between ALT-positive and ALT-negative tumors.

(2) identifying TP53/TCAB1 SVs using WGS data from the GCCRI PDX cohort alongside 23 TARGET-OS samples to expand on the lab's prior SV findings and using TelFusDetector for ALT prediction. We might also use WGS data from Genomics England for analysis.

(3) Implementing Breakend analysis of structural variants in TP53/TCAB1.

(4) Establishing a phylogenetic analysis pipeline using PhylogicNDT, initially will be tested on the 5 available high-coverage samples, in preparation for a larger incoming dataset of 40 samples to investigate the timing of mutational events during osteosarcoma tumorigenesis.

Data Types and Analytical Methods

The project will utilize two primary data types: whole-genome sequencing (WGS) data and RNA sequencing (RNA-seq) data, drawn from two cohorts, the publicly available TARGET-OS project (accessible via NCI's Genomic Data Commons) and the Flynn Lab's existing patient-derived xenograft (PDX) cohort from the Greehey Children's Cancer Research Institute (GCCRI), and Genomics England WGS data.

WGS data will be processed using established SV detection tools, such as DELLY, MANTA, and SvABA, with results merged via SURVIVOR and annotated using AnnotSV, mirroring the pipeline described in the lab's prior work.

Copy number variant analysis will be performed using GATK's somatic CNV workflow. RNA-seq data will be analyzed for differential expression between ALT-positive and ALT-negative samples using DESeq2, with a focus on oxidative stress gene signatures and TCAB1 expression levels.

Breakend analysis of structural variants in TP53/TCAB1 will also be conducted, as well as ALT detection using TelFusDetector.

A preliminary run of the PhylogicNDT phylogenetic inference tool will be conducted on the 5 available high-coverage (60x) WGS samples to validate the pipeline ahead of the full 40-sample dataset.

BMSIP Projects 2025

Project title PI Intern
Alt splicing in AD, PacBio Uwe Beffert Rachel Bozadjian
Ligand-Receptor in AD, AlphaFold Uwe Beffert Riya Jadhav
TCAB1 Mutations in Pediatric Osteosarcoma Rachel Flynn Sydney Sorbello
GPS2-mediated signaling and mitonuclear contact sites Sahana Mitra Tyler Kwok
Diff expression in beetle Lynette Strickland Katherine Kitrick
Measuring kinetic rates of synthetic transcription factors via proSEQ Mo Khalil Nicholas White
Epigenomics underlying microglia cellular states transitions Lei Hou Wenshou He
Multi-omic Approaches to metagenomics Melisa Osborne Benjamin Pfeiffer
Multi-modal omics integration to study HSC engraftment potential. RUBEN DRIES Anuradha Basyal
Cross-species data Integration for Single-Cell Neural Datasets Chao Zhang Jinglin Han
Single-Cell Db for Brain Aging and Neurodegeneration Chao Zhang Sofiya Patra
Building a multimodal model for treatment outcome prediction by integrating histology images and gene expression data Chao Zhang Elaine Huang
Alzheimer’s Disease Risk Modeling Jinying Chen Joshi, Dhruvi
Deep learning to detect base modifications Ignaty Leshchiner Sandilya Bhamidipat
Phylogenetic reconstruction models in tumor biology Ignaty Leshchiner Beatriz Bergamo
GPS2-mediated signaling and mtUPR pathway Valentina Perissi Xinyu Li
scRNA-seq generated from human iPSC-derived organoids Ella Zeldich Shivani Pimparkar
Genomics and transcriptomics of the Betta splendens skin Nelson Lau Hossain, Manseeb
Stem cells in cancer and aging Deborah Lang Jacques Dirabou

More

BMSIP Projects 2024

2024

Project Title PI Intern
Predictive Models for Alzheimer's Disease Risk using Diverse Clinical and Demographic Data Jinying Chen Haochun Huang
Cancer Progression Markers in Melanoma using Multimodal Sequencing Datasets Deborah Lang Shripushkar Ganesh Krishnan
Deep Learning Methods for Cancer Epitranscriptomics Ignaty Leshchiner Xavier Roy
Alcohol Addiction Associated Spatial Gene Expression Characterization in the Brain Phillip Mews Andreea Soica
GPS2-mediated Signaling Crosstalk with Mitochondrial Unfolded Protein Response using ChIPSeq Valentina Perissi Jawahar Mahendran
Characterizing White Adipose Tissue Celltype Lineage Commitment with scRNASeq Nabil Rabhi Akhila Gundavelli
Neuronal Vulnerability in Alzheimer's Disease using snRNASeq Jean-Pierre Roussarie, Anatomy & Neurobiology at BUSM Bhanu Shankar Dhulipalla
Molecular Mechanisms of Aortic Aneurysm using Multimodal Sequencing Datasets Francesca Seta Allison Madsen
Down Syndrome Epigenetics using iPSC-derived Cortical Organoids Ella Zeldich Shreya Nalluri
Bidirectional Encoder Representations Transformer (BERT)-based Microbial Identification Analysis Pipeline Chao Zhang Truman Fogler
Transmission Electron Microsopy Image Curation and Analysis in Chronic Kidney Diseases Chao Zhang Zach Derse
Whole Slide Image Analysis Algorithms in Ovarian and Breast Cancer Chao Zhang Saumya Pothukuchi
Novel Graph-based Model Algorithms for scRNASeq Analysis Chao Zhang Muxi Wang

More

BMSIP Projects 2023

2023

Project Title PI Intern
Data Harmonization with Machine Learning in Alzheimer's Disease Jinying Chen Suraj Prabhu
Machine Learning Algorithms for Alzheimer's Progression Risk Prediction Jinying Chen Vijetha Balakundi
In silico Cell Segmentation of Spatial Transcriptomics in Triple Negative Breast Cancer Tumors Ruben Dries ChihWei Fan
Genetics of Super-resilience in Creutzfeldt-Jakob Disease David Harris/Gustav Mostoslavsky Avarind Sundaravadivelu
Mechanisms of Growth and Differentiation of Melanocytes in Melanoma Deborah Lang Khushi Ahuja
Therapeutic Capacity of Extracellular Vesicles after Cortical Injury Tara Moore Sonal Dinesh Khanna
Molecular Mechanisms sirtuin-1 in Aortic Aneurysm Francesca Seta Pooja Paresh Savla
Molecular Mechanisms of Immune Suppression by AhR in Oral and Lung Cancers David Sherr Vinay Kumar Duggineni
Interactive meta-analysis web tool for gene expression-based contextual classification of cellular phenotypes Tuan Leng Tay Krupa Sampat
Understanding the Epigenetic Landscape of Down Syndrome using Cortical Organoids Ella Zeldich Anna McNiff
Markers of Aging in Mouse and Monkey Models of Alzheimer's Disease Chao Zhang Jou-Hsuan Lee
Computational pipeline for a novel single-cell RNASeq protocol Chao Zhang Zedong Lin

More

BMSIP Projects 2022

2022

Project Title PI Intern
Building a comprehensive soil metagenome database Jenny Bhatnagar, Department of Biology Daniel Golden
Analyses of brain gene expression and DNA methylation in Alzheimer’s disease model mice - effects of perinatal choline nutrition Jan Krzysztof Blusztajn PhD, Department of Pathology and Laboratory Medicine Navin Ramanan
Graph database representations of clinical and 'omics data Adam Labadorf/Taylor Falk, Department of Neurology, BUSM and Bioinformatics Program Merai Dandouch
scRNAseq analysis of mouse and human neurons to understand early Alzheimer's disease related pathogenesis in the entorhinal cortex Jean-Pierre Roussarie, Anatomy & Neurobiology at BUSM Manas Dhanuka
RNA sequencing and proteomics analysis to identify molecular pathways implicated in the development of aortic aneurysms Dr. Francesca Seta, Vascular Biology Section at the Boston University School of Medicine Kyra Griffin-Mitchell
Characterizing epigenetic changes to study the mechanisms of stem cell driven epithelial tissue regeneration using RNAseq, ATACseq, HiC Dr. Andrey Sharov, Dermatology at the Boston University School of Medicine Go Ogata
Using scRNAseq to investigate the use of mesenchymal stromal-cell derived extracellular vesicles in mitigating pathogenic signaling in human oligocortical spheroids Dr. Ella Zeldich, Anatomy & Neurobiology at Boston University School of Medicine Raghad Yamani
Long non-coding RNA expression in post mortem brains in alcohol use disorder Huiping Zhang, Departments of Psychiatry and Medicine, Section of Biomedical Genetics Janvee Patel

More