Dr. Shelley Bull
Lunenfeld-Tanenbaum Research Institute
Statistical genomic methods for population health studies
Our research program aims to develop statistical methods for basic and translational research in population genomic studies: modelling outcomes and associated covariates, and encompassing both discovery and knowledge translation. Technologies of high-dimensional genomic association, genome sequencing and high-throughput proteomics allow interrogation of molecular and genetic variation. Our focus is flexible models and robust inference that integrate multiple sources of trait and genomic data, taking into account complex dependencies and sample characteristics. Our long-term objective is to provide insight into biological mechanisms underlying complex diseases and traits in heterogeneous populations through engagement in multi-disciplinary team collaborations.
Our approach to develop, evaluate and apply innovative analytic methodologies to learn from existing and emerging biomedical study data encompasses fundamental work in statistical theory and methods, evaluation of validity and discovery power of research study designs and statistical technologies, and computational analytic approaches to address gaps in current clinical and population investigations.
Applications in motivating studies of molecular genomic data from national biobanks and multi-disciplinary collaborations are integrated with simulation studies to evaluate estimation accuracy & efficiency and performance of hypothesis tests in datasets generated under complex model scenarios based on study data; evaluations compare novel and existing methods, and seek practical recommendations for research. To this end, we develop software to make analytic and computational technologies accessible to investigators across multiple disciplines.
Email: [email protected]
Floor 5, 60 Murray Street
Toronto, M5T 3L9
Publications: PubMed
Google Scholar: Shelley Bull
ORCID: 0000-0002-3280-7154
- 2026–present; Leadership Team, NSERC CANSSI Genomic Data Science CREATE (Collaborative Research and Training Experience) Program - “STAGE National”
- 2010–present; Associate Editor, Statistics in Medicine
- 2001–present; Professor of Biostatistics, Dalla Lana School of Public Health, University of Toronto, Toronto
- 1992–present; Full Member, Graduate Department of Public Health Sciences, University of Toronto, Toronto
- 1991–present; Senior Investigator, Lunenfeld-Tanenbaum Research Institute, Sinai Health, Toronto
Former appointments
- 2023–2026; Member, Board of Governors, Canadian Statistical Sciences Institute
- 2019–2026; Co-Director, CANSSI Ontario STAGE Program in Genetic Epidemiology & Statistical Genetics, University of Toronto, Toronto
- 2012–2015; Research Committee, Statistical Society of Canada
- 2010–2019; Co-Director, STAGE, CIHR Strategic Training Program in Health Research, University of Toronto, Toronto
- 2005–2008; Board of Directors, International Genetic Epidemiology Society
- 1993–2001; Associate Professor, Department of Public Health Sciences, University of Toronto, Toronto
- 1986–1993; Assistant Professor, Department of Preventive Medicine & Biostatistics, University of Toronto, Toronto
- 1986–1991; Scientist, Clinical Epidemiology, Lunenfeld-Tanenbaum Research Institute, Sinai Health, Toronto
- 1985–1986; Assistant Professor, Department of Epidemiology and Biostatistics, University of Western Ontario, London, Canada
- Ontario Ministry of Health Postdoctoral Fellow, University of Western Ontario, London, Canada; 1983–1985
- PhD, Epidemiology & Biostatistics, University of Western Ontario, London, Canada; 1978–1983
- MMath, Statistics, University of Waterloo, Waterloo, Canada; 1976–1978
- BMath, Statistics, University of Waterloo, Waterloo, Canada; 1972–1976
- 2015 – Statistical Society of Canada Impact Award: Contributions to research in medicine, public health, genetics & epidemiology; development of statistical methodology; leadership, supervision, & mentorship in the Canadian statistical genetics/genetic epidemiology communities
- 2012 – International Genetic Epidemiology Leadership Award in recognition of outstanding leadership through research, teaching or service to the International Genetic Epidemiology Society
- 2002–2007 – Senior Investigator Award, Canadian Institutes of Health Research
- 2001 – Anthony Miller Award for Excellence in Research in Public Health to recognize the outstanding contributions to research of faculty in the Graduate Department of Public Health Sciences
- 1989–1994 – National Health Research Scholar, National Health Research Development Program (First award)
- 1994–1999 – National Health Research Scholar (Second award)
- 1983–1985 – Ontario Ministry of Health Post-doctoral Fellowship
- 1979–1983 – National Health Research Development Program, Doctoral Training Award
Multi-variant region-level methods for data mining in Genome-wide Association Studies
Genetic architecture of complex traits is characterized by genomic polygenicity where multiple genetic variants, each contributing a modest effect, collectively influence the trait, and possibly interact with other variants and/or environmental factors. Genome-wide Association Studies (GWAS) take a comprehensive approach to discover variants that are associated with complex traits and characterize variation in variant effects across populations and environments. Allele frequencies and linkage disequilibrium (LD) among variants differ among diverse populations and across regions of the genome, yielding variation in local genetic architecture.
The emergence of large-scale population cohort studies that comprehensively phenotype participants offer unprecedented opportunities to discover, localize, and fine-map complex trait associations using dense genotyping and/or sequencing data. Analysis in which multiple variants are analyzed together include sets defined by genes, chromosome regions, and/or by variant function.
We aim to advance and evaluate analytic approaches for trait association that integrate region-level genetic association across the genomic spectrum of rare and common variants, using multi-variant methods to reduce multiple testing burden and improve detection of association when there is heterogeneity. Regions are natural subunits in which to refine variant sets, and can be combined to help resolve complex and long-range LD patterns related to trait variation and disease expression.
To inform conditional analysis aimed at identifying biologically relevant variants in regions that have exceeded genome-wide significance thresholds, we are investigating region-level signal-mapping scores designed to prioritize sets of variants and potential biological pathways for experimental studies leading to therapeutic insight.
Improved inference in mixture cure models for prognosis
Survival outcomes arise in association studies in which individuals are followed through time and the occurrence of an event is assessed longitudinally. When the population of interest consists of a mixture of susceptible and non-susceptible individuals, and a significant proportion of individuals will never experience the event (ie. are said to be “cured”), conventional proportional hazards models are inappropriate because they assume that all individuals continue to be at risk.
Our research has been motivated by statistical challenges that emerged in analyses of long-term followup in a molecular prognostic cohort of women diagnosed with early breast cancer. One challenge is to recover information lost in biological sample assay studies when biomarker data are incomplete. Another is improving precision in assessment of association between prognostic factors and clinical outcomes when the sample size is limited.
We are developing novel methods to address these challenges including:
- Penalized maximum likelihood to produce estimates less susceptible to estimation bias, and more robust statistics for hypothesis testing.
- Incorporation of penalization into multiple imputation procedures to augment incomplete data values.
- Implementation of procedures to ensure compatibility between imputation and analysis models used to investigate prognostic biomarkers.
- Profile likelihood methods for bi-parameter hypothesis testing and construction of an associated confidence region of plausible values for relationships of biomarkers with short-term and long-term survival in clinical settings.
Genomic studies of repeated measures of quantitative traits and time-to-event outcomes
When quantitative traits (QTs) are risk factors for disease onset or progression, joint analysis with time-to-event traits (TTEs) can effectively identify association of genetic variants directly with time-to-event and/or indirectly through association with a QT. Although a fully joint likelihood model is feasible for a QT-TTE trait pair, it is computationally challenging for multiple QTs and TTEs. In data-driven simulation studies, we found joint modeling of multiple longitudinal risk factors and TTE outcomes that accounts for within-person QT variability over time and QT-TTE dependence is more informative than methods that ignore these factors.
In on-going work to further improve an existing two-stage approach, we aim to integrate:
- Multivariate linear mixed models describing trajectories of multiple longitudinal traits as a function of time, multiple variant effects, and subject-specific random effects.
- A frailty Cox survival model that depends on SNPs, longitudinal trajectory effects, and subject-specific frailty, accounting for multiple TTE dependence.
Because joint analyses have been limited largely to longitudinal cohort studies with pre-specified visit schedules and long-term follow-up, we are addressing new challenges raised by emergence of population-based biobanks with high-dimensional multi-omic data, short-term follow-up and irregularly collected/incomplete complex trait data.
Family-based studies: Trait ascertainment and selection on polygenic scores in genetic association analysis
Family-based designs involve ascertainment of affected family members: eg. sibling pairs with early age at disease onset, or parent-child trios ascertained on childhood disease. Because ascertainment depends on family member traits, followed by genomic sequencing, we model individual genotypes as an outcome conditional on observed traits. When disease susceptibility depends on combined effects of rare germline variants and polygenic background, ascertainment affects signal detection in complicated ways with implications for both design and analysis.
Motivated by studies of breast cancer and autism spectrum disorder, we designed numerical simulation studies of family ascertainment from a large population under various individual-level disease susceptibility assumptions to compare performance of statistical tests, and quantify effects of alternative ascertainment schema.
For sib-pair studies in which rare variant counts are stratified by local identical-by-descent sharing, power to detect association decreases dramatically when disease susceptibility also depends on polygenic background; efficiency is improved by excluding sib-pairs with high polygenic scores and/or by setting stricter age-at-onset criteria.
For parent-child trios, validity and efficiency is improved by modelling polygenic score as an effect modifier in conditional logistic regression of parent-child allelic transmissions. Ascertainment on individual and familial characteristics can enrich a sample for rare variants of moderate to high penetrance, generating a signal beyond that achievable in an unselected population design.
As a natural extension of on-going work, we are interested in evaluation of methods to integrate family studies with general population-based studies.
We are always looking for motivated researchers to join our team.
Postdocs
Our research group is interested in recruiting highly motivated postdoctoral fellows with a publication record in statistical methods and statistical genomic applications. Please forward your CV, references and research interests to Shelley Bull.
Graduate students
Our research group is part of the Division of Biostatistics in the Dalla Lana School of Public Health, at the University of Toronto, which has a central admission committee. Graduate students interested in doing a PhD in the group must first be accepted in the Biostatistics Program in the Dalla Lana School of Public Health. This also requires that a supervisor be identified before admission to the graduate program.
Summer students
Summer students are selected from successful applicants to the Research Training Center (RTC) at the Lunenfeld-Tanenbaum Research Institute. Applications are available online and need to be filled by February 28th of each year.
Notable publications
Statistics in Medicine, 2026
Bioinformatics Advances, 2025
PLOS Genetics, 2024
Genetics, 2023
Genetic Epidemiology, 2020
Join our team
Visit our job board to see research positions.