/tags/2026-fall/index.xml 2026 Fall - McGill Statistics Seminars
  • Avoiding double dipping through data thinning and data fission

    Date: 2026-09-25

    Time: 15:30-16:30 (Montreal time)

    Location: In person, Burnside 1104

    https://mcgill.zoom.us/j/89492447923

    Meeting ID: 894 9244 7923

    Abstract:

    While classical statistical methods are designed for testing hypotheses about pre-specified models, the reality of modern science is that analysts often explore their data before coming up with models and hypotheses of interest. We refer to the practice of using the same data to generate and then test a hypothesis as double dipping. Classical statistical hypothesis tests will fail to control the Type 1 error rate in settings that involve double dipping. Often, we avoid double dipping by splitting our observations into a training set and a test set. While this sample splitting approach is straightforward and easy to understand, it is generally inapplicable in unsupervised settings. Motivated by unsupervised problems that arise in the analysis of single-cell RNA sequencing data, we first propose data thinning, an alternative to sample splitting that splits each observation in a dataset into two independent pieces. We show that this method provides an elegant solution to our motivating problems under distributional assumptions. We then show how the data fission framework of Leiner et al. (2025) can be fruitfully applied to circumvent the distributional assumptions of data thinning. We conclude by discussing the promise and the challenges of data fission in the context of inference after clustering and inference after variable selection in logistic regression.