{"id":234,"date":"2026-07-30T06:51:36","date_gmt":"2026-07-30T10:51:36","guid":{"rendered":"https:\/\/sites.bu.edu\/bmsip\/?p=234"},"modified":"2026-08-03T02:01:10","modified_gmt":"2026-08-03T06:01:10","slug":"bmsip-projects-2026","status":"publish","type":"post","link":"https:\/\/sites.bu.edu\/bmsip\/2026\/07\/30\/bmsip-projects-2026\/","title":{"rendered":"BMSIP Projects 2026"},"content":{"rendered":"<p><google-sheets-html-origin><\/google-sheets-html-origin><\/p>\n<style type=\"text\/css\"><!--td {border: 1px solid #cccccc;}br {mso-data-placement:same-cell;}--><\/style>\n<table xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\" cellspacing=\"0\" cellpadding=\"0\" dir=\"ltr\" border=\"1\" data-sheets-root=\"1\" data-sheets-baot=\"1\">\n<colgroup>\n<col width=\"292\" \/>\n<col width=\"98\" \/>\n<col width=\"100\" \/>\n<col width=\"271\" \/><\/colgroup>\n<tbody>\n<tr>\n<td>Project<\/td>\n<td>PI<\/td>\n<td>Student<\/td>\n<td>Keywords<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Adam_Labadorf_Jon_Cherry\" target=\"_blank\" rel=\"noopener\">Image-to-Transcriptome: Predicting Neuroinflammatory Signatures from Brain Histology in Chronic Traumatic Encephalopathy<\/a><\/td>\n<td>Dr. Adam Labadorf &amp; Dr. Jon Cherry<\/td>\n<td>Reem Rasmy<\/td>\n<td>Whole Slide imaging, snRNAseq, Deep Learning, Chronic Traumatic Encephalopathy (CTE)<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Andrey_Sharov\" target=\"_blank\" rel=\"noopener\">Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas<\/a><\/td>\n<td>Dr. Andrey Sharov<\/td>\n<td>Nanyu Huang<\/td>\n<td>Bulk RNAseq, Epigenetics, Workflow Development, Assembly<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Andrey_Sharov\" target=\"_blank\" rel=\"noopener\">Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas<\/a><\/td>\n<td>Dr. Andrey Sharov<\/td>\n<td>Krishen Patel<\/td>\n<td>Bulk RNAseq, Epigenetics, Workflow Development, Assembly<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Chao_Zhang\" target=\"_blank\" rel=\"noopener\">Digital Pathology for Kidney Disease Diagnosis<\/a><\/td>\n<td>Dr. Chao Zhang<\/td>\n<td>Qiuguo Tang<\/td>\n<td>Transmission Electron Microscopy, Image Analysis, ML models, Kidney Disease<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Ella_Zeldich_Maya_Medalla\" target=\"_blank\" rel=\"noopener\">Studying activity dependent transcriptome in ACC and LPFC cortical regions of rhesus monkeys<\/a><\/td>\n<td>Dr. Ella Zeldich &amp; Dr. Maya Medalla<\/td>\n<td>Lauren Anderson<\/td>\n<td>Transcriptomics, scRNAseq, Neuroscience<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Ignaty_Leshchiner\" target=\"_blank\" rel=\"noopener\">Deep learning in studying cancer epigenomics<\/a><\/td>\n<td>Dr. Ignaty Leshchiner<\/td>\n<td>Dennis Godin<\/td>\n<td>Long Read Sequencing, Nanopore, epigenetics cancer<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Jennifer_Beane\" target=\"_blank\" rel=\"noopener\">Anthracotic Pigment as a Tissue-Embedded Exposure Footprint in Lung Cancer<\/a><\/td>\n<td>Dr. Jennifer Beane<\/td>\n<td>Nitachakan Obma<\/td>\n<td>RNAseq, Spatial Transcriptomics, Whole Slide Imaging, Cancer<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Jinying_Chen\" target=\"_blank\" rel=\"noopener\">Synergistic Effects of Genetic Risk and Social Determinants of Health on Cognitive Decline Associated with Alzheimer\u2019s Disease<\/a><\/td>\n<td>Dr. Jinying Chen<\/td>\n<td>Tungalan Ganbaatar<\/td>\n<td>ML Models, Deep Learning, Data Visualization, Alzheimer\u2019s Disease<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Josh_Campbell_Sarah_Mazzilli\" target=\"_blank\" rel=\"noopener\">Immune changes in tumor tissue and lymph nodes associated with aggressive non-small cell lung cancer (NSCLC)<\/a><\/td>\n<td>Dr. Josh Campbell &amp; Dr. Sarah Mazzilli<\/td>\n<td>Vaidehi Gupta<\/td>\n<td>Imaging Mass Cytometry, Single Cell and Image Analysis, Cancer<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Josh_Campbell_Sarah_Mazzilli\" target=\"_blank\" rel=\"noopener\">Immune changes in tumor tissue and lymph nodes associated with aggressive non-small cell lung cancer (NSCLC)<\/a><\/td>\n<td>Dr. Josh Campbell &amp; Dr. Sarah Mazzilli<\/td>\n<td>Tazein Shah<\/td>\n<td>Imaging Mass Cytometry, Single Cell and Image Analysis, Cancer<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Lei_Hou\" target=\"_blank\" rel=\"noopener\">Computational Identification of Airway Epithelial Surface Targets from Lung Single-Cell Atlases<\/a><\/td>\n<td>Dr. Lei Hou<\/td>\n<td>Iris Lee<\/td>\n<td>scRNAseq, Protein Databases, Interaction and Signaling Databases, Lung Biology<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Nelson_Lau\" target=\"_blank\" rel=\"noopener\">Genomics of mosquito virus small RNAs<\/a><\/td>\n<td>Dr. Nelson Lau<\/td>\n<td>Emily Dunlop<\/td>\n<td>RNAseq, Small RNAs, Package \/ Workflow Development, Mosquito Biology<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Nikolaos_Daskalakis\" target=\"_blank\" rel=\"noopener\">Mapping cross-disorder genetic risk in psychiatry to molecular changes in the human postmortem brain.<\/a><\/td>\n<td>Dr. Nikolaos Daskalakis<\/td>\n<td>Amalya Murrill<\/td>\n<td>Multi-trait GWAS, Polygenic Risk Score Models, Multi-omics Integration, Psychiatric Disorders<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Pinghua_Liu\" target=\"_blank\" rel=\"noopener\">Machine learning and modeling of steroid metabolic and catabolic pathways and enzymes<\/a><\/td>\n<td>Dr. Pinghua Liu<\/td>\n<td>Xiaohe Jin<\/td>\n<td>Predictive Modeling, Enzyme Catalysis, Human Microbiome<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Pinghua_Liu\" target=\"_blank\" rel=\"noopener\">Machine learning and modeling of steroid metabolic and catabolic pathways and enzymes<\/a><\/td>\n<td>Dr. Pinghua Liu<\/td>\n<td>Shu-Wen Yu<\/td>\n<td>Predictive Modeling, Enzyme Catalysis, Human Microbiome<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Ruben_Dries1\" target=\"_blank\" rel=\"noopener\">Deep-learning-based image feature extraction and integration for spatial transcriptomics in R<\/a><\/td>\n<td>Dr. Ruben Dries<\/td>\n<td>Penelope Varela Dye<\/td>\n<td>Data Analysis, Package \/ Software Development, Spatial Transcriptomics, Image Analysis<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Ruben_Dries2\" target=\"_blank\" rel=\"noopener\">Benchmarking high-resolution spatial transcriptomics re-segmentation tools<\/a><\/td>\n<td>Dr. Ruben Dries<\/td>\n<td>Ryleigh Jerome<\/td>\n<td>Spatial Transcriptomics, Benchmarking, Package Development<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Ruben_Dries3\" target=\"_blank\" rel=\"noopener\">Decoding chromatin accessibility and lineage plasticity in pancreatic neuroendocrine tumors via spatial atac-seq<\/a><\/td>\n<td>Dr. Ruben Dries<\/td>\n<td>Leah Morzenti<\/td>\n<td>ATACseq, Software \/ Package Development, Spatial Analyses, Cancer<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Samagya_Banskota\" target=\"_blank\" rel=\"noopener\">Structural-Based Discovery and In Silico Engineering of Large Serine Recombinases for Precision Genome Editing<\/a><\/td>\n<td>Dr. Samagya Banskota<\/td>\n<td>Cam Nowack<\/td>\n<td>Genome Editing, Large Serine Recombinases, Structural Homology Workflows (AlphaFold, Foldseek), GPU Computing<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Uwe_Beffert\" target=\"_blank\" rel=\"noopener\">Transcriptional programs of APOE\u2013receptor interactions<\/a><\/td>\n<td>Dr. Uwe Beffert<\/td>\n<td>Christine Snow<\/td>\n<td>Differential Expression, Public Data Integration, Bulk RNAseq, Deconvolution, Alzheimer\u2019s Disease<\/td>\n<\/tr>\n<tr>\n<td><a class=\"in-cell-link\" href=\"#Rachel_Flynn\" target=\"_blank\" rel=\"noopener\">Alternative Lengthening of Telomeres (ALT) Genetic Mutations in Pediatric Osteosarcoma<\/a><\/td>\n<td>Dr. Rachel Flynn<\/td>\n<td>Mohammad Gharandouq<\/td>\n<td>Genome Sequencing, Phylogenetics, Mutational profile, Machine Learning, Random Forest, Cancer<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><a id=\"Adam_Labadorf_Jon_Cherry\"><\/a><a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Image-to-Transcriptome: Predicting Neuroinflammatory Signatures from Brain Histology in Chronic Traumatic Encephalopathy<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr. Adam Labadorf &amp; Dr. Jon Cherry <a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Chronic traumatic encephalopathy (CTE) is a neurodegenerative disease caused by<br \/>\nrepetitive head impacts, most studied in contact sport athletes. CTE is<br \/>\ncharacterized by distinctive deposits of phosphorylated tau protein around blood<br \/>\nvessels at the depths of cortical sulci, accompanied by neuroinflammation driven<br \/>\nlargely by microglia \u2014 the brain\u2019s resident immune cells. Recent single-nucleus<br \/>\nRNA sequencing (snRNA-seq) of post-mortem brain tissue from young athletes has<br \/>\nrevealed that repetitive head impacts trigger neuronal loss and microglial<br \/>\nactivation even before clinical symptoms appear, identifying novel inflammatory<br \/>\nmicroglial populations associated with injury. However, snRNA-seq is expensive<br \/>\nand destructive, while stained histology slides are routinely generated during<br \/>\nneuropathological evaluation. If molecular signatures could be predicted<br \/>\ndirectly from histology images, it would dramatically expand our ability to<br \/>\nstudy CTE biology across large archival brain bank cohorts where sequencing data<br \/>\nwill never be available.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>We aim to determine whether deep learning models can predict transcriptomic<br \/>\nfeatures \u2014 specifically cell-type proportions and neuroinflammatory pathway<br \/>\nactivity \u2014 from stained brain tissue images in CTE. This approach is<br \/>\nwell-validated in cancer pathology but has never been applied to<br \/>\nneurodegenerative brain tissue, representing a novel application domain. A key<br \/>\ndesign constraint is that histology images and snRNA-seq data come from opposite<br \/>\nbrain hemispheres of the same donor, precluding spatial registration. We will<br \/>\ntherefore use sample-level prediction, mapping whole slide images to donor-level<br \/>\nmolecular profiles. The 10-week scope targets a proof-of-concept demonstrating<br \/>\nfeasibility and identifying which molecular features are most predictable from<br \/>\nbrain morphology, providing the foundation for a federal grant proposal to scale<br \/>\nand extend this work.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>The dataset comprises ~240 donors with paired whole slide images (WSIs) and<br \/>\nsnRNA-seq profiles. WSIs include multiple immunohistochemical stains: AT8<br \/>\n(phospho-tau), Iba1 and CD68 (microglia), P2RY12 and TMEM119 (microglial<br \/>\nhomeostatic markers), and Nissl\/LFB-CV (neuronal architecture). Prediction<br \/>\ntargets derived from snRNA-seq include pseudo-bulk expression profiles,<br \/>\ncell-type proportions (via deconvolution), and pathway activity scores (via<br \/>\nssGSEA) for neuroinflammation-related gene sets. The analytical pipeline uses a<br \/>\npre-trained pathology vision foundation model (CONCH or UNI2) to extract<br \/>\ntile-level feature embeddings from WSIs, followed by attention-based multiple<br \/>\ninstance learning (ABMIL\/CLAM) with regression heads to predict continuous<br \/>\nmolecular targets. Performance is evaluated by Pearson correlation and R\u00b2 under<br \/>\n5-fold donor-stratified cross-validation. Attention heatmaps provide<br \/>\ninterpretability by highlighting tissue regions driving predictions.<\/p>\n<p><a id=\"Andrey_Sharov\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr. Andrey Sharov <a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>The naked mole rat<span>\u00a0<\/span><i>Heterocephalus glaber<\/i><span>\u00a0<\/span>is a unique mammalian model<br \/>\ncharacterized by exceptional longevity, resistance to cancer, and preserved<br \/>\ntissue homeostasis throughout life. Skin in naked mole rats displays unusual<br \/>\nproperties related to epidermal differentiation, wound healing, and tumor<br \/>\nresistance, making it an attractive system to study protective epigenetic<br \/>\nmechanisms. However, unlike mouse or human, the naked mole rat remains a<br \/>\nnon-model organism with incomplete genome annotation, limiting the application<br \/>\nof modern epigenomic approaches such as chromatin accessibility profiling and 3D<br \/>\ngenome analysis. This project focuses on skin biology and genome regulation,<br \/>\nlaying the groundwork necessary to enable future studies of epigenetic control<br \/>\nof development, regeneration, and cancer resistance in this species.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>Large-scale epigenomic analyses require a well-curated reference genome and<br \/>\ntranscriptome, which are currently lacking for naked mole rat skin. The absence<br \/>\nof standardized genome builds, refined gene models, and cross-species<br \/>\nannotations represents a major barrier to interpreting RNA-seq, ATAC-seq, and<br \/>\nchromatin conformation data. This internship will address this foundational<br \/>\nproblem by evaluating existing naked mole rat genomic resources and generating<br \/>\nan analysis-ready genome and transcriptome framework. The resulting reference<br \/>\nwill support the creation of a future naked mole rat skin epigenomic atlas and<br \/>\nenable meaningful comparisons with mouse and human skin.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>The intern will work primarily with bulk RNA-seq data generated from naked mole<br \/>\nrat skin. Analytical methods will include RNA-seq quality control, alignment,<br \/>\ntranscript quantification, and identification of expressed genes. The student<br \/>\nwill evaluate available naked mole rat genome assemblies and annotations, select<br \/>\na primary reference, and generate standardized FASTA and GTF files suitable for<br \/>\ndownstream epigenomic analyses. Additional analyses will include basic<br \/>\ntranscriptome annotation refinement, assessment of repeat content and<br \/>\nmappability, and ortholog mapping between naked mole rat, mouse, and human<br \/>\ngenes. Emphasis will be placed on reproducible workflows and biological<br \/>\ninterpretation.<\/p>\n<p><a id=\"Chao_Zhang\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Digital Pathology for Kidney Disease Diagnosis<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Chao Zhang<a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Chronic kidney diseases (CKD) affect 13% of the population and costs the<br \/>\nUS at least $50 billion annually. Transmission electron microscopy (TEM) remains<br \/>\ngold standard for renal histopathological diagnoses, especially for proteinuric<br \/>\nkidney diseases. Currently, measurements of podocyte foot process width (FPW)<br \/>\nand glomerular basement membrane (GBM) width in TEM images are performed<br \/>\nmanually, which limits the accuracy and efficacy of ultrastructural analysis.<br \/>\nTherefore, we propose to develop an AI computational digital biopsy platform to<br \/>\nmeasure podocyte FPW and GBM width of healthy and pathological kidney specimens<br \/>\nautomatically using TEM images from pre-clinical podocytopathy animal models and<br \/>\nclinical kidney biopsy samples from patients with podocytopathy diagnoses.<br \/>\nSpecifically, this product will manifest in the form of a web application that<br \/>\ncan be accessed online and is categorized as an image analysis software. It<br \/>\ncould enhance treatment decisions and expedite drug development in pre-clinical<br \/>\nmodels and clinical trials, advancing solutions for proteinuric kidney diseases.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>Collaborating with several different labs, we have collected 60000+ patients\u2019<br \/>\nTEM images. The previous model works very well on mouse and rat data but is less<br \/>\naccurate on human data due to a lack of labels. The potential projects during<br \/>\nthe summer internship could be: 1) Assisting in labeling, organizing human data,<br \/>\nand improving the deep learning model to achieve better accuracy. 2)<br \/>\nImplementing a Mamba-based model to integrate the TEM images and pathological<br \/>\nreports. This could involve generating pathological report templates from images<br \/>\nto reduce the writing time for pathologists or building new classification<br \/>\nmodels for disease and drug response prediction. 3) Utilizing data from The<br \/>\nKidney Precision Medicine Project to explore the potential of multimodal<br \/>\nintegration for TEM images, histology images, and omics data Data Types and<br \/>\nAnalytical Methods What types of biological data will be used? How will it be<br \/>\nanalyzed? The intern must possess strong Python programming skills and<br \/>\nexperience in medical image analysis, particularly with libraries such as<br \/>\nOpenCV, scikit-image, numpy and pandas Knowledge of and experience with<br \/>\nanalyzing pathological Images is highly desirable. The intern\u2019s primary<br \/>\nresponsibility will be to develop and implement algorithms for the processing<br \/>\nand analysis of images, focusing on image segmentation, feature extraction, and<br \/>\nmachine learning-based classification. Collaboration with a multidisciplinary<br \/>\nteam, including medical professionals, to refine analysis techniques and ensure<br \/>\nclinical relevance. Evaluation of existing analytical methods and adaptation or<br \/>\ndevelopment of new approaches to improve accuracy and efficiency. Documentation<br \/>\nof methodologies, code, and analysis procedures for future reference and use by<br \/>\nthe research team.<\/p>\n<p><a id=\"Ella_Zeldich_Maya_Medalla\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Studying activity dependent transcriptome in ACC and LPFC cortical regions of rhesus monkeys<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr. Maya Medalla and Dr. Ella Zeldich<a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>In our previous work we profiled transcriptionally and functionally the<br \/>\ndifferences between the two important brain regions. In the current project, we<br \/>\nwill be specifically focusing on the changes related to neuronal activity<br \/>\ninduces by cognitive tests. We will look at the population of the cfos-positive<br \/>\nneurons and the changes within this population. The obtained results will be<br \/>\nintegrated with our previous data.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>How the regional differences between ACC and LPFC pertain to cognition-related<br \/>\nneuronal activity<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>scRNA-seq<\/p>\n<p><a id=\"Lei_Hou\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Computational Identification of Airway Epithelial Surface Targets from Lung Single-Cell Atlases<\/strong><br \/>\n<strong>PI:<\/strong><span>\u00a0<\/span>Lei Hou <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Targeted delivery of gene-editing or RNA therapeutics to airway epithelial cells<br \/>\nremains a major challenge in cystic fibrosis (CF) and other lung diseases.<br \/>\nCurrent delivery systems often rely on broadly expressed receptors, resulting in<br \/>\ninefficient uptake by the desired cell types and unintended interactions with<br \/>\nimmune or stromal cells. The airway epithelium itself is heterogeneous,<br \/>\nconsisting of basal, secretory, ciliated, and rare specialized cell populations,<br \/>\neach with distinct molecular programs and surface protein composition. Recent<br \/>\nadvances in single-cell RNA sequencing (scRNA-seq) have generated comprehensive<br \/>\natlases of the human lung across healthy and disease contexts. These datasets<br \/>\nprovide an opportunity to systematically identify cell-type\u2013specific surface<br \/>\nproteins that could serve as molecular \u201centry points\u201d for targeted delivery.<br \/>\nHowever, gene expression alone is insufficient to determine whether a candidate<br \/>\nreceptor is biologically active or functionally relevant. Integrating additional<br \/>\nlayers of information\u2014such as cell\u2013cell communication and regulatory network<br \/>\ncontext\u2014can help identify receptors that are not only expressed but also<br \/>\nactively involved in epithelial signaling programs.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>Although lung single-cell datasets have revealed extensive epithelial<br \/>\nheterogeneity, there is currently no systematic framework to identify and<br \/>\nprioritize airway epithelial surface targets for therapeutic delivery. In<br \/>\nparticular, three key questions remain unresolved: Which receptors and membrane<br \/>\nproteins are robustly and specifically expressed in airway epithelial cell types<br \/>\nacross datasets and disease conditions? Among these proteins, which are actively<br \/>\nengaged in cell\u2013cell communication, suggesting extracellular accessibility and<br \/>\nfunctional relevance? Which receptors are embedded in regulatory programs<br \/>\nlinking receptor activation to downstream transcriptional responses, indicating<br \/>\nbiological importance within epithelial signaling networks? Addressing these<br \/>\nquestions requires integrating single-cell transcriptomics with protein<br \/>\nannotations and network-based inference methods to produce a ranked set of<br \/>\ncandidate surface targets.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>Single-cell transcriptomics Public lung scRNA-seq datasets including: CF and<br \/>\ncontrol airway epithelium datasets (e.g., Nat Med 2021) Airway epithelial<br \/>\ndifferentiation datasets (e.g., Shah et al., Nat Commun 2025) Lung atlases<br \/>\nsuch as LungMAP and the Human Lung Cell Atlas (HLCA) These datasets provide<br \/>\ncell-type resolution across epithelial subpopulations.<\/p>\n<p>Protein and functional annotation resources Gene Ontology (GO) membrane<br \/>\nannotations UniProt subcellular localization Cell Surface Protein Atlas (CSPA)<br \/>\nHuman Protein Atlas (HPA) These resources help translate gene expression signals<br \/>\ninto surface protein candidates. Interaction and signaling databases<br \/>\nLigand\u2013receptor interaction databases (e.g., CellPhoneDB, NicheNet) Protein<br \/>\ninteraction databases (e.g., STRING, BioGRID) These provide biological context<br \/>\nfor receptor activity.<\/p>\n<p>The project will consist of three major analytical components.<\/p>\n<p>1. Identification of Epithelial Surface Proteins from Lung scRNA-seq We will<br \/>\nintegrate multiple lung scRNA-seq datasets to identify receptors and membrane<br \/>\nproteins enriched in airway epithelial populations. Key steps include: Dataset<br \/>\nharmonization and batch correction Identification of epithelial cell types and<br \/>\nsubtypes Differential expression analysis to identify epithelial-enriched genes<br \/>\nFiltering candidate genes based on plasma membrane annotations and surface<br \/>\nprotein databases This analysis will produce a catalog of epithelial surface<br \/>\nproteins, annotated by functional class such as receptors, transporters,<br \/>\nadhesion molecules, and ion channels.<\/p>\n<p>2. Inference of Receptor Activity through Cell\u2013Cell Communication Analysis To<br \/>\nevaluate whether candidate receptors participate in extracellular signaling, we<br \/>\nwill infer ligand\u2013receptor interactions between epithelial cells and neighboring<br \/>\ncell populations. Methods include: Ligand\u2013receptor inference frameworks such as<br \/>\nCellPhoneDB or NicheNet Quantification of interaction strength between<br \/>\nepithelial cells and surrounding immune or stromal populations This step helps<br \/>\nidentify receptors that are actively involved in extracellular signaling<br \/>\nnetworks, supporting their accessibility and biological relevance.<\/p>\n<p>3. Inference of Receptor Function through Receptor\u2013TF\u2013Gene Regulatory Networks<br \/>\nTo assess downstream signaling consequences, we will construct receptor-centered<br \/>\nregulatory networks linking receptors to transcription factors and target genes.<br \/>\nThis analysis will involve: Mapping candidate receptors to known signaling<br \/>\npathways Integrating protein interaction networks to connect receptors with<br \/>\ntranscription factors Evaluating downstream transcriptional programs associated<br \/>\nwith receptor activity Receptors embedded in coherent receptor\u2013TF\u2013gene<br \/>\nregulatory modules will be prioritized, as they are more likely to represent<br \/>\nfunctionally active signaling nodes in epithelial cells.<\/p>\n<p><a id=\"Ignaty_Leshchiner\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Deep learning in studying cancer epigenomics<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr.Ignaty Leshchiner<a href=\"#\" style=\"font-size: small;\">\u00a0top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Computational biology and cancer bioinformatics. Research in the lab is focused on applying new genomic technologies, computational analysis and AI methods on data from patients\u2019 tumors to understand the biology behind tumor development, treatment evasion, and progression to metastasis. We are developing and applying tools for simultaneous analysis of multiple samples from the same patient, clonal structure, integration of single cell genomics and transcriptomics, reconstruction of cell subpopulations, their growth kinetics and expression, tumor micro-environment effects, estimation of order of events (\u201ctiming\u201d) during tumor development and progression. We work with pre- and post- treatment samples, autopsies and longitudinal blood biopsies in solid and blood malignancies.<\/p>\n<p class=\"p1\"><b>Problem Statement <\/b><\/p>\n<p class=\"p1\">There is an ever-growing body of literature which show how modifications to the DNA, both genetically and epigenetically, as well as modifications to proteins are distinct in the setting of cancer. Changes to RNA, especially those epitranscriptomic in nature, have not yet extensively been studied. This is because while the technologies to accurately assess the base-pair changes to DNA are well established, the ability to detect native epitranscriptomic changes in DNA and RNA is not yet a robust technology. We are developing deep learning methods that enable calling of methylation modifications from both RNA and DNA with higher accuracy from Native Nanopore based sequencing and identify tumor type specific DNA\/RNA modifications. The project will involve analyzing and training the models to improve call accuracy and detect biological changes within samples.<\/p>\n<p class=\"p1\"><b>Data Types and Analytical Methods<\/b><span class=\"s1\"> <\/span><\/p>\n<p class=\"p1\">For this project we will use raw signal processing of Nanopore hdf5 files and converting the current signal into accurate base space and methylation calls. We train the models on standard data we generate in the lab both on cancer samples and cancer\/normal cell line models.<\/p>\n<p><a id=\"Jennifer_Beane\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Anthracotic Pigment as a Tissue-Embedded Exposure Footprint in Lung Cancer<\/h4>\n<p><strong>PI:<\/strong><span> Dr. <\/span>Jennifer Beane <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>This study involves examining the cells that comprise the lung and how they<br \/>\nrespond to chronic inhaled pollutants and how this exposure increases the risk<br \/>\nof developing lung cancer and other lung diseases.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>Chronic inhaled exposures, including cigarette smoke and air pollution, promote<br \/>\nlung cancer risk by causing cumulative genomic damage and remodeling the lung<br \/>\nmicroenvironment. Anthracotic pigment is visible carbonaceous particulate matter<br \/>\nin lung tissue that can bind carcinogens and may serve as a tissue-embedded<br \/>\nexposure footprint, but its spatial patterns and associated molecular programs<br \/>\nacross non-cancerous and cancerous tissues remain poorly defined. We will<br \/>\nleverage large public cohorts and AI-based methods for whole slide images to<br \/>\nquantify pigment and define lung cancer-associated pigment subtypes and spatial<br \/>\nniches, enabling development of exposure-informed biomarkers that improve risk<br \/>\nstratification beyond self-reported smoking history.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>The types of data used in the project are bulk RNA sequencing data, H&amp;E-stained<br \/>\nwhole slide images (WSIs), and spatial transcriptomics data. The project<br \/>\ninvolves applying state-of-the art cell segmentation and classification methods<br \/>\nto the WSIs and analyzed both the bulk and spatial transcriptomics data to<br \/>\nidentify both local, distance-dependent effects and global tumor<br \/>\nmicroenvironment shifts associated with WSI phenotypes.<\/p>\n<p><a id=\"Jinying_Chen\"><\/a> <a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<h4 class=\"page-title\">Synergistic Effects of Genetic Risk and Social Determinants of Health on Cognitive Decline Associated with Alzheimer\u2019s Disease<\/h4>\n<p><strong>PI<\/strong>: Dr. Jinying Chen <a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<p><strong>Project Summary<\/strong><\/p>\n<p>Alzheimer\u2019s disease (AD) develops over decades, with amyloid and tau pathology<br \/>\naccumulating before clinical symptoms emerge. During this preclinical stage,<br \/>\ntrajectories of cognitive decline vary widely, suggesting that biological<br \/>\nsusceptibility and modifiable contextual factors jointly shape risk. Polygenic<br \/>\nrisk scores (PRS) capture inherited liability to AD and have been associated<br \/>\nwith earlier onset and steeper cognitive decline, yet PRS alone provides an<br \/>\nincomplete account of individual outcomes. Social determinants of health<br \/>\n(SDOH)\u2014including socioeconomic resources, neighborhood context, education, and<br \/>\nsocial support\u2014may influence cognitive reserve, health behaviors, comorbidity<br \/>\nburden, and access to care, thereby modifying the clinical expression of<br \/>\nunderlying AD pathology. However, the combined and potentially interactive<br \/>\ncontributions of AD polygenic risk and SDOH to cognitive decline in<br \/>\nbiomarker-defined preclinical AD remain insufficiently characterized.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>Understanding how genetic susceptibility and SDOH jointly relate to early<br \/>\ncognitive changes could improve risk stratification and inform more equitable<br \/>\nprevention strategies. This study aims to assess the interaction effects between<br \/>\ngenetic risk factors and SDOH on cognitive decline associated with AD by using<br \/>\nstatistical methods, data visualization, and machine learning.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><span>\u00a0<\/span>Data (e.g., demographics,<br \/>\ngenetic, SDOH, clinical, AD biomarkers) from national and international AD<br \/>\ncohorts (e.g., Health and Retirement Study, Alzheimer\u2019s Disease Data Initiative)<br \/>\nwill be used for this study. Cognitive decline will be characterized by<br \/>\nstatistical methods (e.g., latent trajectory modeling and\/or other regression<br \/>\nmodels) and data visualization methods, and predicted by machine learning (e.g.,<br \/>\nLasso regression, Random Forest, deep neural networks, SHAP analysis). The<br \/>\nstudent will be supervised by Dr. Chen and is expected to conduct a significant<br \/>\npiece of work (data preprocessing and analysis using either machine learning or<br \/>\nstatistical methods or both) independently for this project. This position<br \/>\nrequires strong programming skills in Python and\/or R, and prior experience in<br \/>\ndeveloping and evaluating machine learning models and\/or conducting statistical<br \/>\nanalysis. Knowledge and experience with developing deep learning models is a<br \/>\nplus. The data analysis will be conducted on BU Shared Computing Cluster (SCC)<br \/>\nand\/or the computational platform of Alzheimer\u2019s Disease Data Initiative, with<br \/>\nPython and\/or R.<\/p>\n<p><a id=\"Josh_Campbell_Sarah_Mazzilli\"><\/a> <a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<h4 class=\"page-title\">Immune changes in tumor tissue and lymph nodes associated with aggressive non-small cell lung cancer (NSCLC)<\/h4>\n<p><strong>PI<\/strong>: Dr. Josh Campbell &amp; Dr. Sarah Mazilli <a href=\"#\" style=\"font-size: small;\"> top<\/a><\/p>\n<p><strong>Project Summary<\/strong><\/p>\n<p>Regional lymph nodes (LNs) in the thoracic cavity serve as essential<br \/>\nimmunological hubs that coordinate humoral and cell-mediated responses against<br \/>\nthe development and progression of non-small cell lung cancer (NSCLC). To<br \/>\ninvestigate immune dysregulation in the non-metastatic regional LNs of patients<br \/>\nwith aggressive NSCLC, we performed multimodal profiling on 36 LNs from 11<br \/>\npatients undergoing curative-intent resection including CITE-seq, scRNA-seq, and<br \/>\nImaging Mass Cytometry (IMC). Regional N1 LNs from patients with more aggressive<br \/>\ndisease (stage IB\u2013IIIA) exhibited a significant enrichment of dysfunctional CD8\u207a<br \/>\nT cells and regulatory T cells (Tregs) compared to N2 LNs and LNs from patients<br \/>\nwith less aggressive disease (stage IA). These immune subsets were spatially<br \/>\nco-localized with mature regulatory dendritic cells (mregDCs; CD1c\u207a, TIM3\u207a,<br \/>\nLAMP3\u207a), forming an immunosuppressive niche uniquely enriched in the N1 LNs of<br \/>\nhigher-stage patients. Concurrently, higher-stage N1 LNs contained larger number<br \/>\nof \u201cdecorticated\u201d B-cell follicles characterized by decreased encapsulation of<br \/>\nthe mantle zone layer surrounding the germinal centers. This mantle zone<br \/>\ndisorganization was associated with increased spatial niches involving Tregs,<br \/>\nCD68+ CD163\u207a TIM3\u207a Macrophages, CD163\u207a TIM3dim Monocytic-Myeloid Derived<br \/>\nSuppressor Cells (M-MDSC), plasma B cells, and a decrease in spatial niches<br \/>\ninvolving CD4\u207a T helper cells and fibroblastic reticular cells (FRCs). Together,<br \/>\nour findings reveal parallel alterations in humoral and cell-mediated immunity<br \/>\nwithin the regional LNs of patients with aggressive NSCLC.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>While we have analyzed the IMC and single cell data from LNs, we have also<br \/>\ngenerated IMC data from multiple regions of tumor tissue and adjacent normal in<br \/>\neach patient. We need to have this data analyzed to identify immune populations<br \/>\nand cellular niches in tumors. These findings will be correlated with those we<br \/>\npreviously found in the regional LNs. The goal will be to see what aberrant<br \/>\nimmune populations in the aggressive tumors are also observed in LNs or are<br \/>\nspecific to the tumors.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>Imaging mass cytometry (IMC) which measure levels of cell surface proteins on<br \/>\ntissue slides to maintain spatial architecture. We use computational pipelines<br \/>\nfor single cell analysis, image analysis, niche identification, and statistical<br \/>\nmethods for associating cell populations and niches with clinical phenotypes.<\/p>\n<p><a id=\"Nelson_Lau\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Genomics of Mosquito Virus small RNAs<\/h4>\n<p><strong>PI:<\/strong><span> Dr. <\/span>Nelson Lau <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Problem Statement<\/strong><br \/>\nThe mosquito Aedes aegypti is the major vector for arbovirus pathogens like<br \/>\nDengue and Zika viruses. We are also analyzing mosquito tombus viruses that may<br \/>\ninfect and compete against Dengue and Zika viruses to generate a small RNA\/ RNAi<br \/>\nresponse. The Lau lab is looking for a BMSIP intern to work on continuing the<br \/>\nanalysis of the small RNA responses in mosquitoes and mosquito cells subjected<br \/>\nto infection by tombus viruses. We will be seeing if other mosquito genes are<br \/>\nbeing affected in RNAi mutants that may lose the capacity to keep mosquito<br \/>\nviruses in check. The intern will learn how to parse the outputs from our<br \/>\nMosquito Small RNA Genomics pipeline.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>There are many RNAseq and small RNA libraries already generated and sequenced,<br \/>\nand the project will involve working closely with the Lau lab team to organize<br \/>\nand conduct differential expression analysis on mosquito libraries to look for<br \/>\ngenes, transposons and virus expression changes. The goal will be to see if the<br \/>\nmosquito virus and RNAi pathways are impacting gene expression in the Aedes<br \/>\naegypti mosquitoes.<\/p>\n<p><a id=\"Nikolaos_Daskalakis\"><\/a> <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<h4 class=\"page-title\">Mapping cross-disorder genetic risk in psychiatry to molecular changes in the human postmortem brain<\/h4>\n<p><strong>PI:<\/strong><span> Dr. <\/span>Nikolaos Daskalakis <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>This project investigates the functional genomics of psychiatric pleiotropy in<br \/>\nhumans \u2013 the phenomenon where single genetic variants influence multiple<br \/>\ndistinct disorders. Traditionally, disorders like PTSD, MDD, and substance use<br \/>\ndisorders (SUD) have been studied in diagnostic silos. However, shared genetic<br \/>\narchitecture suggests common underlying biological systems, such as dysregulated<br \/>\nneuroendocrine signaling, synaptic plasticity, and neuroinflammatory pathways.<br \/>\nBy leveraging deep-phenotyping data from the human postmortem brain, we are<br \/>\nlooking at the \u2018molecular endophenotypes\u2019 of these disorders. We are<br \/>\nspecifically interested in how genetic risk manifests across different brain<br \/>\nregions and cell types (e.g., excitatory neurons vs. glia). This work moves<br \/>\nbeyond the \u2018one gene, one disease\u2019 model to explore how broad genetic factors \u2013<br \/>\nin line with notions of psychiatric comorbidities and multimorbidities \u2013 shape<br \/>\nthe biological landscape of the human brain before and after the onset of<br \/>\nclinical symptoms.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>It is well-established that psychiatric disorders share e.g. trauma, mood and<br \/>\nsubstance use disorders share not only clinical outcomes (phenotype), but also<br \/>\ngenomic architecture. However, the translation of this shared genetic risk into<br \/>\nmolecular changes within the brain remains a critical knowledge gap. The primary<br \/>\nchallenge in precision psychiatry is the \u2018missing link\u2019 between an individual\u2019s<br \/>\ngenetic liability (what one is born with) against their clinical state (the<br \/>\nactual disease outcome). While we can identify genetic risk through GWAS, we<br \/>\noften do not know if the molecular changes we see in a patient\u2019s brain are the<br \/>\ncause of the disorder or a consequence of living with it (e.g., due to chronic<br \/>\nstress or medication). This internship addresses this by comparing Polygenic<br \/>\nRisk Scores (PRS) \u2013 a genetically -derived \u2018biomarker\u2019 of disease risk, against<br \/>\npostmortem molecular data. By doing so, we aim to validate the biological<br \/>\nreality of cross-disorder genetics. We are asking: does a high genetic risk for<br \/>\ninternalizing disorders create a specific, observable molecular signature in the<br \/>\nbrain, regardless of a patient\u2019s clinical diagnosis? Solving this helps us move<br \/>\ntoward a biologically-defined classification of mental health disorders rather<br \/>\nthan one based purely on clinical phenotypes. As such, the proposed project<br \/>\nbridges genetics and multi-omics by leveraging the latest cross-disorder GWAS<br \/>\nsummary statistics (Grotzinger et al., 2025, Nature) to conduct a<br \/>\ntraining-to-validation study in a postmortem brain cohort (Daskalakis et al.,<br \/>\n2024, Science).<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>We will adopt a testing-to-validation computational framework. In the training<br \/>\nstage, multi-trait GWAS summary statistics (Grotzinger et al., 2025)<br \/>\ncorresponding to shared genomic factors e.g. (i) internalizing factor (PTSD,<br \/>\nMDD, anxiety disorder) and (ii) substance use factor (cannabis use disorder,<br \/>\nalcohol use disorder, nicotine dependence, opioid use disorder) will be used to<br \/>\nderive polygenic risk scores (PRS) models. In the validation stage, the trained<br \/>\nmodels will then be applied to the postmortem brain cohort (Daskalakis et al.,<br \/>\n2024), for PRS estimation (on the aforementioned factors) and downstream<br \/>\nassociation testing. This is a well characterized cohort of N=304 donors across<br \/>\nthree brain regions (medial prefrontal cortex, dentate gyrus and central<br \/>\namygdala) that is heavily examined within the Lab. We hypothesize that these PRS<br \/>\nencoding for shared genetic risk will map to broad, transdiagnostic clinical<br \/>\nphenotypes (e.g. symptom clusters, risk factors, non-psychiatric health<br \/>\ncomorbidities) rather than isolated diagnostic categories. Next, we will examine<br \/>\nwhether the shared genetic burden also translate to multi-omic (transcriptomic,<br \/>\nmethylomic, proteomic) and single-cell molecular changes within the brain. This<br \/>\nwill enable comparison of genetic risk vs clinical outcomes on the<br \/>\nmolecular\/biological effects on the brain. Softwares include PRScs for PRS model<br \/>\ntraining, PLINK for PRS score estimation, R programming for statistical testing<br \/>\nand differential expression analysis using linear approaches (LIMMA).<\/p>\n<p><a id=\"Pinghua_Liu\"><\/a> <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<h4 class=\"page-title\">Machine learning and modeling of steroid metabolic and catabolic pathways and enzymes<\/h4>\n<p><strong>PI:<\/strong><span> Dr. <\/span>Pinghua Liu <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>The importance of bile acids in human health has been known for decades. In<br \/>\nrecent years, it comes to the realization that the diverse biological roles for<br \/>\nbile acids have been closely related to the human microbiota, especially the<br \/>\nsteroid metabolic and catabolic enzymes. In these reactions, hydroxylation and<br \/>\ndihydroxylation reactions, especially the site- and stereo-specificities are<br \/>\nknown to be important for their biological activities. In this project, we would<br \/>\nlike to systematically evaluate literature to develop a predictive model for gut<br \/>\nbacteria steroid metabolic and catabolic pathways, which will then be linked to<br \/>\nthe function of gut microbiota to human health.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>We aim at develop predictive model for steroid metabolic and catabolic enzyme<br \/>\nfunctions.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>Literature information on various enzymes. We have develop overexpression and<br \/>\nenzymatic catalytic assays for these enzymatic reactions already.<\/p>\n<p><a id=\"Ruben_Dries1\"><\/a> <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<h4 class=\"page-title\">Deep-learning-based image feature extraction and integration for spatial transcriptomics in R<\/h4>\n<p><strong>PI<\/strong>: Dr. Ruben Dries <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Spatial transcriptomics technologies are increasingly paired with<br \/>\nhigh-resolution histological images, most commonly hematoxylin and eosin<br \/>\n(H&amp;E)\u2013stained sections. These images capture rich morphological information<br \/>\nabout tissue architecture, cellular organization, and pathological features that<br \/>\nare not directly encoded in gene expression data alone. In cancer and other<br \/>\ncomplex diseases, histological context, such as stromal organization, tumor<br \/>\nboundaries, or immune infiltration, play critical roles in disease progression<br \/>\nand treatment response. Recent advances in deep learning have made it possible<br \/>\nto extract informative image features from histological images at multiple<br \/>\nspatial scales. When integrated with spatial transcriptomics data, these<br \/>\nfeatures offer the potential to link morphology with molecular states, enabling<br \/>\nmore comprehensive models of tissue organization and function.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>While spatial transcriptomics datasets routinely include associated<br \/>\nhistological images, these images are often underutilized in downstream<br \/>\nanalysis. Existing methods for image feature extraction and multimodal<br \/>\nintegration are fragmented, difficult to reproduce, or require leaving the R<br \/>\necosystem entirely. Indeed, most existing approaches rely on Python-based<br \/>\npipelines and are not yet natively accessible within R-based spatial analysis<br \/>\nframeworks commonly used by biologists. As a result, many spatial studies miss<br \/>\nthe opportunity to systematically incorporate morphological context.<\/p>\n<p>This internship aims to explore R-native implementations for deep-learning and<br \/>\nultimately create a blueprint for establishing a robust and extensible baseline<br \/>\nfor image-based analysis within the Giotto Suite. Specifically, the project will<br \/>\nexplore how different deep-learning models and image tiling strategies affect<br \/>\ndownstream biological interpretation, and how image-derived features can be<br \/>\nmeaningfully integrated with gene expression data. The goal is not to build a<br \/>\nsingle \u201cbest\u201d model, but rather to define practical, well-documented utilities<br \/>\nthat enable users to explore image\u2013expression relationships in a standardized<br \/>\nand reproducible way.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>The project will primarily use spatial transcriptomics datasets (e.g., Xenium,<br \/>\nMERFISH, Visium HD) with associated H&amp;E images, drawn from public datasets and<br \/>\ninternal examples. Key analytical components include: Image feature extraction:<br \/>\nTesting and extending existing Giotto utilities to support multiple pretrained<br \/>\ndeep-learning models, and evaluating whether model choice materially affects<br \/>\ndownstream analyses. Multi-scale tiling strategies: Exploring approaches that<br \/>\ncombine large, medium, and small image tiles to capture tissue-, neighborhood-,<br \/>\nand cell-level morphology. Multimodal integration: Evaluating methods to<br \/>\nintegrate image features with gene expression, including joint latent spaces and<br \/>\ncorrelated feature representations. Documentation and tutorials: Creating a<br \/>\nuser-facing tutorial demonstrating image-based workflows within Giotto Suite.<br \/>\nAll development and analysis will be performed primarily in R, with a focus on<br \/>\nmaking advanced image analysis natively available to Giotto users. Ideal<br \/>\ncandidates have some previous ML\/DL experience and are interested to work at the<br \/>\ninterface of data analysis and package\/software development.<\/p>\n<p><a id=\"Ruben_Dries2\"><\/a> <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<h4 class=\"page-title\">Benchmarking high-resolution spatial transcriptomics re-segmentation tools<\/h4>\n<p><strong>PI:<\/strong><span> Dr. <\/span>Ruben Dries <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>High-resolution spatial transcriptomics (ST) platforms, such as Xenium and<br \/>\nMERFISH, have revolutionized our ability to study tissue biology by providing<br \/>\nsubcellular localization of RNA molecules. By mapping transcripts within their<br \/>\nnative architecture, researchers can identify distinct cellular niches,<br \/>\nmetabolic gradients, and complex cell-cell communication networks. These<br \/>\ntechnologies are critical for understanding how the spatial organization of<br \/>\ntissues\u2014such as the tumor microenvironment or the layered structure of the<br \/>\nbrain\u2014dictates biological function and disease progression.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>The accuracy of all downstream spatial analysis is dependent on cell<br \/>\nsegmentation: the process of defining cellular boundaries. However, current<br \/>\nimage-based segmentation often suffers from \u201ctranscript leakage,\u201d where RNAs<br \/>\nfrom one cell are incorrectly assigned to a neighbor. This noise confounds<br \/>\ndifferential expression analysis, masks rare cell states, and generates false<br \/>\npositives in ligand-receptor signaling studies. While new algorithms (e.g.<br \/>\nRNA2seg &amp; others) offer post-hoc refinement using transcript density or<br \/>\nstatistical demultiplexing strategies, they are currently fragmented, not<br \/>\nimplemented in downstream analysis pipelines, or tested on real-world datasets.<br \/>\nThis internship addresses the lack of a standardized, user-friendly pipeline to<br \/>\ncorrect these errors, which currently limits the biological reliability of<br \/>\nhigh-resolution ST data.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>This project will utilize high-resolution ST datasets (Xenium, MERFISH, and<br \/>\nCosMx) from both public repositories and in-house experiments. The primary<br \/>\nanalytical goal is to operationalize and quantitatively benchmark<br \/>\nre-segmentation workflows within the Giotto Suite, an R-based framework<br \/>\ndeveloped in the Dries lab. Methodologically, the project involves: Software<br \/>\nIntegration: Developing scripts to connect existing transcript-based refinement<br \/>\nand statistical deconvolution methods into existing Giotto pipelines.<br \/>\nBenchmarking: Applying quantitative metrics (e.g., silhouette scores, cluster<br \/>\npurity) to compare standard segmentation against refined outputs. Validation:<br \/>\nEvaluating the \u201csignal-to-noise\u201d recovery in downstream tasks, specifically<br \/>\nfocusing on the clarity of cell-type markers and the accuracy of spatial<br \/>\nneighborhood analyses.<\/p>\n<p><a id=\"Ruben_Dries3\"><\/a> <a href=\"#\" style=\"font-size: small;\">top\u00a0<\/a><\/p>\n<h4 class=\"page-title\">Decoding chromatin accessibility and lineage plasticity in pancreatic neuroendocrine tumors via spatial atac-seq<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr. Ruben Dries <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Pancreatic neuroendocrine tumors (PanNETs) are characterized by significant<br \/>\nclinical heterogeneity and the ability of tumor cells to shift between different<br \/>\nphenotypic states, a process known as lineage plasticity. While transcriptomic<br \/>\ntechnologies have highlighted these states, the spatial context of the<br \/>\nunderlying chromatin landscape, including the regulatory regions such as<br \/>\npromoters and enhancers, remains largely unexplored. Spatial ATAC-seq (Assay for<br \/>\nTransposase-Accessible Chromatin using sequencing) allows for the mapping of<br \/>\nopen chromatin regions directly within the tissue architecture. Through a<br \/>\ncollaboration with AtlasXomics, the Heaphy (BU) and Singhi (UPMC) labs, we have<br \/>\napplied this for the first time to Formalin-Fixed Paraffin-Embedded (FFPE)<br \/>\ntissue, the gold standard for clinical samples. This approach enables us to<br \/>\nstudy high-quality patient specimens through a new epigenomic lens. By capturing<br \/>\nthe epigenetic state in situ, we can observe how the spatial organization of the<br \/>\ntumor microenvironment influences gene regulation and facilitates the transition<br \/>\nbetween cell lineages.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>The primary challenge in understanding PanNET progression is identifying the<br \/>\nspecific epigenetic drivers that trigger lineage transitions. Current analysis<br \/>\npipelines for spatial chromatin data are still in their infancy and often fail<br \/>\nto link spatial accessibility directly to local and spatially-informed Gene<br \/>\nRegulatory Networks (GRNs). This internship addresses this gap by focusing on:<br \/>\nRegulatory Mapping: Identifying changes in promoter and enhancer accessibility<br \/>\nthat correlate with lineage plasticity. GRN Modeling: Building spatially aware<br \/>\ngene regulatory networks to identify the transcription factors driving<br \/>\nphenotypic switches. Tool Development: Implementing these spatial chromatin<br \/>\nanalysis workflows within the Giotto Suite framework (www.giottosuite.com).<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>The project will utilize high-resolution spatial ATAC-seq data from human PanNET<br \/>\nFFPE samples. Methodologically, the project involves: Giotto Integration:<br \/>\nDeveloping and benchmarking new modules, inspired by ArchR, within the Giotto<br \/>\nSuite (an R-based framework developed in the Dries lab) specifically for spatial<br \/>\nchromatin data. Regulatory Analysis: Analyzing differential accessibility at<br \/>\npromoters and distal enhancers to identify regulatory elements associated with<br \/>\nspecific cellular niches. Network Inference: Using accessibility patterns and TF<br \/>\nmotif enrichment to reconstruct GRNs that define lineage plasticity. Secondary<br \/>\nAnalysis (Optional): Testing computational methods to infer CNV profiles to<br \/>\ntrack clonal evolution across the tissue. Ideal candidates should have at least<br \/>\nminimal hands-on experience in sequence-level genomics, such as analyzing<br \/>\n(single-cell) ATAC-seq, ChIP-seq, CUT&amp;RUN, GRO-seq, TT-seq, etc.<\/p>\n<p><a id=\"Samagya_Banskota\"><\/a> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Structural-Based Discovery and In Silico Engineering of Large Serine Recombinases for Precision Genome Editing<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr. Samagya Banskota <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>Large Serine Recombinases (LSRs) are a class of enzymes derived from<br \/>\nbacteriophages and other mobile genetic elements (MGEs) that facilitate<br \/>\nsite-specific integration of large DNA payloads into host genomes. Unlike<br \/>\nCRISPR-Cas systems, which create double-stranded breaks, LSRs catalyze highly<br \/>\nefficient, one-way recombination between distinct attachment sites. This makes<br \/>\nthem indispensable tools for genome engineering, particularly for therapeutic<br \/>\ngene insertion in mammalian cells. This project shifts the focus from<br \/>\ntraditional sequence-based identification toward structural conservation and<br \/>\nbiophysical properties of these enzymes in order to identify improved protein<br \/>\nscaffolds. Protein engineering can further improve these novel scaffolds to<br \/>\naddress a need for higher specificity and efficiency in current genetic therapy<br \/>\napproaches.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>Past LSR discovery relied heavily on sequence-based homology (HMMs), which often<br \/>\nfail to identify highly divergent LSRs that maintain a similar functional<br \/>\nstructure. This creates a \u201cblind spot\u201d in the known landscape of recombinases.<br \/>\nFurthermore, many naturally occurring LSRs lack the necessary specificity or<br \/>\nefficiency for clinical use, or they target \u201csafe harbor\u201d sites that are not<br \/>\ntherapeutically relevant. This internship will address these limitations by<br \/>\nadvancing existing enzyme discovery pipelines with a structure-based homology<br \/>\nsearch and attempt in silico protein engineering to refine feature spaces for<br \/>\nimproved efficiency, specificity, and the re-targeting of LSRs to<br \/>\ntherapeutically<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>The project will utilize several gigabytes of publicly available whole-genome assembly data from<br \/>\nNCBI, the Sequence Read Archive (SRA) and protein structure data from the ESM Metagenomic<br \/>\nAtlas. Analytical methods will transition from HMMER-based sequence searches to structural<br \/>\nhomology workflows using FoldSeek and AlphaFold\/ESMFold to identify candidates based on 3Di<br \/>\nembedding architecture rather than nucleotide sequence.<\/p>\n<p>Key analytical steps include:<\/p>\n<p>1. Seed-Based Refinement: Using high-confidence, improved LSR structures as seeds to<br \/>\niteratively refine the structural feature space of our search.<\/p>\n<p>2. Computational Re-targeting: Applying machine learning or physics-based modeling to predict<br \/>\nmutations that allow LSRs to recognize and integrate into specific, non-canonical att sites.<\/p>\n<p>3. Workflow Management: All analysis will be performed using a Snakemake-managed pipeline,<br \/>\nutilizing Python for data wrangling and MAFFT for sequence\/structural alignment validation.<\/p>\n<p><a id=\"Uwe_Beffert\"><\/a>\u00a0<a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Transcriptional programs of APOE\u2013receptor interactions<\/h4>\n<p><strong>PI:<\/strong><span>\u00a0<\/span>Dr. <span data-sheets-root=\"1\">Uwe Beffert<\/span> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p><strong>Biological Background<\/strong><\/p>\n<p>The APOE \u03b54 allele confers the strongest genetic risk for late-onset Alzheimer\u2019s disease and<br \/>\nmodulates neuroinflammatory and neuronal vulnerability pathways. ApoE binds several lipoprotein<br \/>\nreceptors including APOER2 (Apoer2 in mouse), which undergoes alternative splicing. In our near-<br \/>\nsubmission manuscript, we demonstrate that inclusion vs deletion of Apoer2 exon 19 is a key<br \/>\ndeterminant of APOE-dependent signaling in the brain. Using four mouse lines combining Apoer2<br \/>\nexon 19 inclusion\/deletion with humanized APOE3\/APOE4 backgrounds, we show that Apoer2<br \/>\nsplicing is the primary driver of hippocampal gene expression variance, with APOE genotype exerting<br \/>\nstrong context-dependent effects. RNA-seq and pathway analyses highlight<br \/>\ninflammatory\/immune\/cilia programs in APOE4 contexts and antiviral\/RNA quality control programs in<br \/>\nexon 19\u2013deleted contexts. Follow-up imaging suggests ApoE genotype and Apoer2 splicing also<br \/>\naffect glial features and primary neuronal cilia morphology.<\/p>\n<p><strong>Problem Statement<\/strong><\/p>\n<p>We have already performed major RNA-seq analyses for the manuscript, including principal<br \/>\ncomponent analysis, multiple differential expression contrasts, gene set\/pathway enrichment, and<br \/>\nselected follow-up validations (some figures may still contain placeholders pending final data).<br \/>\nHowever, there are additional analyses that could significantly strengthen the paper and generate<br \/>\nvaluable mechanistic hypotheses, particularly by:<br \/>\n1. explicitly modeling this dataset as a 2\u00d72 factorial design (ApoE3 vs ApoE4) \u00d7 (Apoer2<br \/>\n+Ex19 vs \u0394Ex19) to identify interaction genes that define context-dependent ApoE effects;<br \/>\n2. inferring which cell types (neurons vs astrocytes vs microglia vs oligodendrocyte lineage vs<br \/>\nendothelial) are most likely responsible for observed expression programs in bulk<br \/>\nhippocampus;<br \/>\n3. extending the work from \u201clists of DE genes\u201d to network\/module-level interpretation; and<br \/>\n4. comparing signatures to publicly available brain\/cell-type datasets relevant to AD,<br \/>\nneuroinflammation, interferon\/antiviral programs, and cilia biology.<\/p>\n<p><strong>Data Types and Analytical Methods<\/strong><\/p>\n<p>Data: Bulk hippocampus RNA-seq from four mouse genotypes: humanized ApoE3 or ApoE4 crossed<br \/>\nwith Apoer2 exon 19 inclusion (+Ex19) or deletion (\u0394Ex19). Metadata include genotype labels and,<br \/>\nwhere available, sex, batch, and sample QC metrics.<br \/>\nWork already completed (foundation): The manuscript currently includes core analyses such as<br \/>\nPCA\/variance structure, multiple DE contrasts (including ApoE effects within an Apoer2 splice<br \/>\nbackground and vice versa), overlap summaries (e.g., Venn-style comparisons), and pathway<br \/>\nenrichment summaries; some validation figures are in progress.<br \/>\nInternship extensions (new deliverables):<br \/>\n1. Factorial DE modeling: DESeq2 (or equivalent) using a full design with main effects and an<br \/>\ninteraction term (ApoE \u00d7 Apoer2 splicing) to identify genes with true context-dependent<br \/>\nregulation.<br \/>\n2. Cell-type inference: Enrichment\/deconvolution approaches using public mouse hippocampus<br \/>\nsingle-cell\/single-nucleus references to assign DE\/interaction signals to likely cell types.<br \/>\n3. Gene network\/module analysis: co-expression modules (e.g., WGCNA) and module\u2013trait<br \/>\nassociations (genotype, inferred cell-type scores, and available phenotypes).<br \/>\n4. Public dataset comparison: evaluate concordance with published AD-relevant signatures<br \/>\n(microglial activation\/DAM, interferon\/antiviral programs, cilia gene sets, astrocyte reactivity),<br \/>\nand summarize which signatures map to which genotype contrasts.<br \/>\nOutputs: a reproducible repo (scripts + README), ranked gene lists for main effects\/interaction,<br \/>\nmodule summaries, and manuscript-ready figures\/tables.<\/p>\n<p><a id=\"Rachel_Flynn\"><\/a>\u00a0<a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<h4 class=\"page-title\">Alternative Lengthening of Telomeres (ALT) Genetic Mutations in Pediatric Osteosarcoma<\/h4>\n<p><strong>PI:<\/strong><span> <\/span><span data-sheets-root=\"1\">Dr. Rachel Flynn<\/span> <a href=\"#\" style=\"font-size: small;\">top<\/a><\/p>\n<p class=\"p1\"><b>Biological Background<\/b><\/p>\n<p class=\"p1\">This project focuses on telomere maintenance mechanisms in pediatric osteosarcoma, a rare and aggressive primary bone cancer. Specifically, it examines the alternative lengthening of telomeres (ALT) pathway, which accounts for cellular immortalization in approximately 75% of pediatric osteosarcoma cases.<\/p>\n<p class=\"p1\">ALT relies on homologous recombination rather than the enzyme telomerase to maintain telomere length. Central to this work is the gene TCAB1 (also known as WRAP53), an RNA chaperone that is essential for telomerase holoenzyme assembly and localization.<\/p>\n<p class=\"p1\">TCAB1 partially overlaps with the TP53 tumor suppressor gene on chromosome 17p13.1, and structural variations (SVs) disrupting TP53 in osteosarcoma may simultaneously inactivate TCAB1, thereby abolishing telomerase activity and potentially driving cells toward ALT activation. Genes such as ATRX, DAXX, SMARCAL1, and SLX4IP are also known contributors to ALT and will be examined in the context of tumor evolution.<\/p>\n<p class=\"p1\"><b>Problem Statement<\/b><\/p>\n<p class=\"p1\">The genetic mechanisms underlying ALT activation in osteosarcoma remain incompletely understood. While mutations in chromatin remodeling genes like ATRX and DAXX are known contributors, they alone are insufficient to fully explain ALT induction. Recent work from the Flynn Lab identified structural variations at the TP53\/TCAB1 locus in approximately 40% of ALT-positive osteosarcoma tumors, suggesting that functional inactivation of the telomerase holoenzyme may be an early, previously unrecognized step in ALT activation.<\/p>\n<p class=\"p1\">This internship will extend that work by:<\/p>\n<p class=\"p1\">(1) Analyzing existing RNA-seq data from both the TARGET-OS cohort and the Flynn Lab&#8217;s GCCRI PDX cohort to investigate differential gene expression, including oxidative stress pathways, between ALT-positive and ALT-negative tumors.<\/p>\n<p class=\"p1\">(2) identifying TP53\/TCAB1 SVs using WGS data from the GCCRI PDX cohort alongside 23 TARGET-OS samples to expand on the lab&#8217;s prior SV findings and using TelFusDetector for ALT prediction. We might also use WGS data from Genomics England for analysis.<\/p>\n<p class=\"p1\">(3) Implementing Breakend analysis of structural variants in TP53\/TCAB1.<\/p>\n<p class=\"p1\">(4) Establishing a phylogenetic analysis pipeline using PhylogicNDT, initially will be tested on the 5 available high-coverage samples, in preparation for a larger incoming dataset of 40 samples to investigate the timing of mutational events during osteosarcoma tumorigenesis.<\/p>\n<p class=\"p1\"><b>Data Types and Analytical Methods<\/b><\/p>\n<p class=\"p1\">The project will utilize two primary data types: whole-genome sequencing (WGS) data and RNA sequencing (RNA-seq) data, drawn from two cohorts, the publicly available TARGET-OS project (accessible via NCI&#8217;s Genomic Data Commons) and the Flynn Lab&#8217;s existing patient-derived xenograft (PDX) cohort from the Greehey Children&#8217;s Cancer Research Institute (GCCRI), and Genomics England WGS data.<\/p>\n<p class=\"p1\">WGS data will be processed using established SV detection tools, such as DELLY, MANTA, and SvABA, with results merged via SURVIVOR and annotated using AnnotSV, mirroring the pipeline described in the lab&#8217;s prior work.<\/p>\n<p class=\"p1\">Copy number variant analysis will be performed using GATK&#8217;s somatic CNV workflow. RNA-seq data will be analyzed for differential expression between ALT-positive and ALT-negative samples using DESeq2, with a focus on oxidative stress gene signatures and TCAB1 expression levels.<\/p>\n<p class=\"p1\">Breakend analysis of structural variants in TP53\/TCAB1 will also be conducted, as well as ALT detection using TelFusDetector.<\/p>\n<p class=\"p1\">A preliminary run of the PhylogicNDT phylogenetic inference tool will be conducted on the 5 available high-coverage (60x) WGS samples to validate the pipeline ahead of the full 40-sample dataset.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Project PI Student Keywords Image-to-Transcriptome: Predicting Neuroinflammatory Signatures from Brain Histology in Chronic Traumatic Encephalopathy Dr. Adam Labadorf &amp; Dr. Jon Cherry Reem Rasmy Whole Slide imaging, snRNAseq, Deep Learning, Chronic Traumatic Encephalopathy (CTE) Establishing a Genome and Transcriptome-based Framework for a Naked Mole Rat Skin Epigenomic atlas Dr. Andrey Sharov Nanyu Huang Bulk RNAseq, [&hellip;]<\/p>\n","protected":false},"author":24347,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[1],"tags":[],"_links":{"self":[{"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/posts\/234"}],"collection":[{"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/users\/24347"}],"replies":[{"embeddable":true,"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/comments?post=234"}],"version-history":[{"count":47,"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/posts\/234\/revisions"}],"predecessor-version":[{"id":273,"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/posts\/234\/revisions\/273"}],"wp:attachment":[{"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/media?parent=234"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/categories?post=234"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/sites.bu.edu\/bmsip\/wp-json\/wp\/v2\/tags?post=234"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}