ISBA Young Researchers Day / YRD
18/09/2026 - 09:00 - ISBA C115 -
9h00 Introduction
9h10 Tommaso Martini
"Smoothing Copulas Without Smoothing Away Tail Dependence"
Abstract:
Smooth nonparametric copula estimators have been found to improve finite-sample performance over the empirical copula in a range of settings. A prominent example is the empirical beta copula introduced by Segers, Sibuya, and Tsukahara. At any fixed sample size, however, the empirical beta copula has zero lower- and upper-tail dependence coefficients. This motivates the study of smoothing mechanisms that retain the benefits of smoothing while permitting nontrivial tail dependence. The empirical beta copula and, more generally, the class of smooth nonparametric copula estimators studied by Kojadinovic and Yi admit a common plug-in representation through an operator that averages a starting copula with respect to a family of smoothing distributions. We study the tail behaviour of the copula produced by this operator. Under suitable regularity conditions, we derive an explicit expression for its lower tail copula. The expression shows how the induced lower-tail behaviour is shaped jointly by the starting copula, by the marginal behaviour of the smoothing distributions, and by the lower tail copula of the common survival copula governing their dependence structure. For grid-supported smoothing distributions, the general result reduces to finite-sum expressions for the lower tail copula and its associated lower-tail dependence coefficients, extending composite Bernstein-type constructions beyond binomial smoothing. These results also suggest a class of rank-based semiparametric copula estimators in which a nonparametric pilot provides the starting dependence structure, while a parametric survival copula enters through the smoothing mechanism and influences the induced lower-tail behaviour of the smoothed estimator.
9h50 Cyril Ghislain
"Distribution-Free Runs Tests for Directional Serial Dependence via Measure Transportation"
Abstract:
Directional runs statistics provide natural measures of serial dependence, but their null distribution generally depends on the unknown marginal distribution of the observations. This prevents the direct extension of the distribution-free Wald–Wolfowitz test to multivariate directional data. We address this problem by transporting directional signs to a uniform distribution on the sphere before measuring their serial alignment. The resulting empirical procedure is based on an optimal assignment of the observations to a deterministic spherical grid. We establish the asymptotic normality of the resulting statistic and show that estimating the underlying direction is asymptotically costless. Under contiguous Markov alternatives generating weak serial alignment, the lag-one statistic coincides with the central sequence of the limiting experiment and yields a locally asymptotically most powerful test. Simulations support the finite-sample accuracy of the asymptotic theory.
10h20 Lise Léonard
"Asymptotic Inference for High-Dimensional Linear Regression via Model Averaging Debiased SLOPE"
Abstract:
We consider the problem of estimation and inference in high-dimensional linear regression models where the number of covariates exceeds the sample size. While penalized estimators such as the Lasso provide effective variable selection, their use for statistical inference is hindered by two key limitations: they are sparse and depend on a regularization parameter that is unknown in practice.
The SLOPE estimator, recently proposed in the literature, is a sorted $\ell_1$-penalized regression method that assigns larger penalties to larger coefficients. It can be viewed as a generalization of the Lasso and enjoys improved oracle properties, including minimax-optimal estimation rates.
Leveraging these advantages, we address both challenges simultaneously by proposing MADSlope, which combines a debiasing procedure with model averaging over a grid of regularization parameters.
We establish that MADSlope is asymptotically Gaussian, even when the averaging weights are data-driven and thus random, and we prove that the weighting scheme is asymptotically optimal in terms of prediction loss. Moreover, the resulting estimator is tuning-free, eliminating the need to select a single regularization parameter.
Our theoretical guarantees are supported by simulation studies and an application to a real high-dimensional dataset.
10h50 Break
11h10 Philippe Hauchamps
"Deciphering microbial subcommunities of the vaginal microbiota: a statistical modelling approach using topic models"
Abstract:
The vaginal microbiota plays a crucial role in women’s health, with its composition and dynamics influencing susceptibility to infections, adverse pregnancy outcomes, and overall well-being. While mainstream approaches to characterize metagenomic sample composition have focused on dominant taxa, some recent studies highlight the importance of microbial subcommunities and their interactions within this complex ecosystem.
In this talk, we explore the use of statistical topic models as a framework to decipher microbial subcommunities in the vaginal microbiota. Topic models, widely used in text mining, allow for the identification of latent structures within high-dimensional count data. By applying this approach, we aim to provide a data-driven characterization of bacterial subcommunities, with the hope to shed some light on their ecological roles, and their association with health and disease states.
11h50 Kenrick So
"Climate Driven Mortality Forecasting using Deep Learning"
Abstract:
Climate extremes are now important drivers of mortality, producing sudden spikes that traditional models fail to predict. Despite this, most mortality models ignore climate causing life insurers to be exposed to these extreme events. To address this, we propose a two-step framework that combines a regional weekly Lee–Carter baseline with CNN-LSTM and GNN-LSTM models to capture residual climate-mortality patterns. The CNN and GNN components capture spatial effects across regions, while the LSTM models both short- and long-term temporal relationships of climate and mortality. This allows our models to capture the response of mortality to the delayed and nonlinear climate effects. We demonstrate that our model captures the association between extreme temperatures and excess mortality, and generates forecasts that account for both extreme events and forecast uncertainty. As a result, our proposed models are more accurate compared to both the Lee–Carter baseline and a gated recurrent unit-based mortality model. From a risk management perspective, the framework provides a more realistic assessment of extreme climate-driven mortality risk, with important implications for Solvency II capital adequacy evaluation.
12h20 Brief ILV on-the-spot feedback and closing
12h30 Lunch in the cafeteria