/post/index.xml Past Seminar Series - McGill Statistics Seminars
  • Enriched post-selection models for high dimensional data

    Date: 2022-04-08

    Time: 15:35-16:35 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    High dimensional data are rapidly growing in many domains, for example, in microarray gene expression studies, fMRI data analysis, large-scale healthcare analytics, text/image analysis, natural language processing and astronomy, to name but a few. In the last two decades regularisation approaches have become the methods of choice for analysing high dimensional data. However, obtaining accurate estimates and predictions as well as valid statistical inference remains a major challenge in high dimensional situations. In this talk, we present enriched post-selection models that aim to improve parameter estimation and prediction, and to facilitate statistical inferences in high dimensional regression models. The enriched post-selection method enables us to construct valid post-selection inference for regression parameters in high dimensions. We discuss the empirical and asymptotic properties of the enriched post-selection method.

  • Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control

    Date: 2022-04-01

    Time: 15:35-16:35 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    We introduce Learn then Test, a framework for calibrating machine learning models so that their predictions satisfy explicit, finite-sample statistical guarantees regardless of the underlying model and (unknown) data-generating distribution. The framework addresses, among other examples, false discovery rate control in multi-label classification, intersection-over-union control in instance segmentation, and the simultaneous control of the type-1 error of outlier detection and confidence set coverage in classification or regression. To accomplish this, we solve a key technical challenge: the control of arbitrary risks that are not necessarily monotonic. Our main insight is to reframe the risk-control problem as multiple hypothesis testing, enabling techniques and mathematical arguments different from those in the previous literature. We use our framework to provide new calibration methods for several core machine learning tasks with detailed worked examples in computer vision.

  • Distribution-​free inference for regression: discrete, continuous, and in between

    Date: 2022-03-25

    Time: 15:35-16:35 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    In data analysis problems where we are not able to rely on distributional assumptions, what types of inference guarantees can still be obtained? Many popular methods, such as holdout methods, cross-validation methods, and conformal prediction, are able to provide distribution-free guarantees for predictive inference, but the problem of providing inference for the underlying regression function (for example, inference on the conditional mean E[Y|X]) is more challenging. If X takes only a small number of possible values, then inference on E[Y|X] is trivial to achieve. At the other extreme, if the features X are continuously distributed, we show that any confidence interval for E[Y|X] must have non-vanishing width, even as sample size tends to infinity - this is true regardless of smoothness properties or other desirable features of the underlying distribution. In between these two extremes, we find several distinct regimes - in particular, it is possible for distribution-free confidence intervals to have vanishing width if and only if the effective support size of the distribution ofXis smaller than the square of the sample size.

  • New Approaches for Inference on Optimal Treatment Regimes

    Date: 2022-03-11

    Time: 15:30-16:30 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    Finding the optimal treatment regime (or a series of sequential treatment regimes) based on individual characteristics has important applications in precision medicine. We propose two new approaches to quantify uncertainty in optimal treatment regime estimation. First, we consider inference in the model-free setting, which does not require specifying an outcome regression model. Existing model-free estimators for optimal treatment regimes are usually not suitable for the purpose of inference, because they either have nonstandard asymptotic distributions or do not necessarily guarantee consistent estimation of the parameter indexing the Bayes rule due to the use of surrogate loss. We study a smoothed robust estimator that directly targets the parameter corresponding to the Bayes decision rule for optimal treatment regimes estimation. We verify that a resampling procedure provides asymptotically accurate inference for both the parameter indexing the optimal treatment regime and the optimal value function. Next, we consider the high-dimensional setting and propose a semiparametric model-assisted approach for simultaneous inference. Simulation results and real data examples are used for illustration.

  • Structure learning for extremal graphical models

    Date: 2022-02-18

    Time: 15:30-16:30 (Montreal time)

    https://umontreal.zoom.us/j/85105423917?pwd=enM3MGpFNkZKU2daMjRITmo0N0JUUT09

    Meeting ID: 851 0542 3917

    Passcode: 403790

    Abstract:

    Extremal graphical models are sparse statistical models for multivariate extreme events. The underlying graph encodes conditional independencies and enables a visual interpretation of the complex extremal dependence structure. For the important case of tree models, we provide a data-driven methodology for learning the graphical structure. We show that sample versions of the extremal correlation and a new summary statistic, which we call the extremal variogram, can be used as weights for a minimum spanning tree to consistently recover the true underlying tree. Remarkably, this implies that extremal tree models can be learned in a completely non-parametric fashion by using simple summary statistics and without the need to assume discrete distributions, existence of densities, or parametric models for marginal or bivariate distributions. Extensions to more general graphs are also discussed.

  • Integration of multi-omics data for the discovery of novel regulators that modulate biological processes

    Date: 2022-02-11

    Time: 15:30-16:30 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    The cellular states in various biological processes such as cell differentiation, disease progression, and treatment response are often enormously complex and thus hard to be profiled with unimodal profiling (e.g., transcriptome). Although those unimodal measurements had brought success for studies in a large variety of studies, the incomplete (and often misleading) unimodal cellular profiling could lead to
    biased and inaccurate conclusions. With the development of biotechnologies, the availability of multi-omics data (bulk or single-cell) is ever-increasing. The rapid-accumulating multi-omics data offers unprecedented opportunities to accurately decode the cellular states in biological process and thus could derive a deep understanding of the change of the cellular states, crucial for finding biomarkers and therapeutic intervention strategies. In this talk, we will discuss a few multimodal methods that we developed to integrate multi-omics data for the discovery of novel regulators for multiple biological processes. Many of the novel predictions from the multimodal methods were experimentally validated and had brought new understandings of the underlying mechanisms for several diseases. I will also discuss how a potential novel COVID19 drug is discovered from such a multi-omics data integration analysis.

  • Off-Policy Confidence Interval Estimation with Confounded Markov Decision Process

    Date: 2022-02-04

    Time: 15:30-16:30 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    In this talk, we consider constructing a confidence interval for a target policy’s value offline based on pre-collected observational data in infinite horizon settings. Most of the existing works assume no unmeasured variables exist that confound the observed actions. This assumption, however, is likely to be violated in real applications such as healthcare and technological industries. We show that with some auxiliary variables that mediate the effect of actions on the system dynamics, the target policy’s value is identifiable in a confounded Markov decision process. Based on this result, we develop an efficient off-policy value estimator that is robust to potential model misspecification and provides rigorous uncertainty quantification.

  • Risk assessment, heavy tails, and asymmetric least squares techniques

    Date: 2022-01-28

    Time: 15:30-16:30 (Montreal time)

    https://umontreal.zoom.us/j/93983313215?pwd=clB6cUNsSjAvRmFMME1PblhkTUtsQT09

    Meeting ID: 939 8331 3215

    Passcode: 096952

    Abstract:

    Statistical risk assessment, in particular in finance and insurance, requires estimating simple indicators to summarize the risk incurred in a given situation. Of most interest is to infer extreme levels of risk so as to be able to manage high-impact rare events such as extreme climate episodes or stock market crashes. A standard procedure in this context, whether in the academic, industrial or regulatory circles, is to estimate a well-chosen single quantile (or Value-at-Risk). One drawback of quantiles is that they only take into account the frequency of an extreme event, and in particular do not give an idea of what the typical magnitude of such an event would be. Another issue is that they do not induce a coherent risk measure, which is a serious concern in actuarial and financial applications. In this talk, after giving a leisurely tour of extreme quantile estimation, I will explain how, starting from the formulation of a quantile as the solution of an optimization problem, one may come up with two alternative families of risk measures, called expectiles and extremiles, in order to address these two drawbacks. I will give a broad overview of their properties, as well as of their estimation at extreme levels in heavy-tailed models, and explain why they constitute sensible alternatives for risk assessment using real data applications. This is based on joint work with Abdelaati Daouia, Irène Gijbels, Stéphane Girard, Simone Padoan and Antoine Usseglio-Carleve.

  • Change-point analysis for complex data structures

    Date: 2022-01-21

    Time: 15:30-16:30 (Montreal time)

    https://mcgill.zoom.us/j/83436686293?pwd=b0RmWmlXRXE3OWR6NlNIcWF5d0dJQT09

    Meeting ID: 834 3668 6293

    Passcode: 12345

    Abstract:

    The change-point analysis is more than sixty years old. Over this long period, it has been an important subject of interest in many scientific disciplines such as finance and econometrics, bioinformatics and genomics, climatology, engineering, and technology.

    In this talk, I will provide a general overview of the topic alongside some historical notes. I will then review the most recent and transformative advancements on the subject. Finally, I will discuss the change-point methodologies that my research team has developed over the past several years, covering various complex data structures.

  • Adventures with Partial Identifications in Studies of Marked Individuals

    Date: 2021-11-26

    Time: 15:30-16:30 (Montreal time)

    Zoom Link

    Meeting ID: 939 8331 3215

    Passcode: 096952

    Abstract:

    Monitoring marked individuals is a common strategy in studies of wild animals (referred to as mark-recapture or capture-recapture experiments) and hard to track human populations (referred to as multi-list methods or multiple-systems estimation). A standard assumption of these techniques is that individuals can be identified uniquely and without error, but this can be violated in many ways. In some cases, it may not be possible to identify individuals uniquely because of the study design or the choice of marks. Other times, errors may occur so that individuals are incorrectly identified. I will discuss work with my collaborators over the past 10 ye ars developing methods to account for problems that arise when are only individuals are only partially identified. I will present theoretical aspects of this research, including an introduction to the latent multinomial model and algebraic statistics, and also describe applications to studies of species ranging from the golden mantella (an endangered frog endemic to Madagascar measuring only 20 mm) to the whale shark (the largest known species of sh measuring up to 19 m).