Date: 2026-09-25
Time: 15:30-16:30 (Montreal time)
Location: In person, Burnside 1104
https://mcgill.zoom.us/j/89492447923
Meeting ID: 894 9244 7923
Abstract:
While classical statistical methods are designed for testing hypotheses about pre-specified models, the reality of modern science is that analysts often explore their data before coming up with models and hypotheses of interest. We refer to the practice of using the same data to generate and then test a hypothesis as double dipping. Classical statistical hypothesis tests will fail to control the Type 1 error rate in settings that involve double dipping. Often, we avoid double dipping by splitting our observations into a training set and a test set. While this sample splitting approach is straightforward and easy to understand, it is generally inapplicable in unsupervised settings. Motivated by unsupervised problems that arise in the analysis of single-cell RNA sequencing data, we first propose data thinning, an alternative to sample splitting that splits each observation in a dataset into two independent pieces. We show that this method provides an elegant solution to our motivating problems under distributional assumptions. We then show how the data fission framework of Leiner et al. (2025) can be fruitfully applied to circumvent the distributional assumptions of data thinning. We conclude by discussing the promise and the challenges of data fission in the context of inference after clustering and inference after variable selection in logistic regression.
Speaker
Anna Neufeld is an Assistant Professor of Statistics at Williams College, where her research focuses on selective inference and the analysis of genomic data. She received her PhD in Statistics from the University of Washington in 2023 and subsequently completed postdoctoral training at the Fred Hutchinson Cancer Center before joining the faculty at Williams in 2024.