Avoiding double dipping through data thinning and data fission
Anna Neufeld · Sep 25, 2026
Date: 2026-09-25
Time: 15:30-16:30 (Montreal time)
Location: In person, Burnside 1104
https://mcgill.zoom.us/j/89492447923
Meeting ID: 894 9244 7923
Abstract:
While classical statistical methods are designed for testing hypotheses about pre-specified models, the reality of modern science is that analysts often explore their data before coming up with models and hypotheses of interest. We refer to the practice of using the same data to generate and then test a hypothesis as double dipping. Classical statistical hypothesis tests will fail to control the Type 1 error rate in settings that involve double dipping. Often, we avoid double dipping by splitting our observations into a training set and a test set. While this sample splitting approach is straightforward and easy to understand, it is generally inapplicable in unsupervised settings. Motivated by unsupervised problems that arise in the analysis of single-cell RNA sequencing data, we first propose data thinning, an alternative to sample splitting that splits each observation in a dataset into two independent pieces. We show that this method provides an elegant solution to our motivating problems under distributional assumptions. We then show how the data fission framework of Leiner et al. (2025) can be fruitfully applied to circumvent the distributional assumptions of data thinning. We conclude by discussing the promise and the challenges of data fission in the context of inference after clustering and inference after variable selection in logistic regression.