Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Date: Wednesday, 18/Mar/2026 | |
| 8:50am - 9:00am | Opening Location: 0.004 |
| 9:00am - 10:00am | Plenary Lecture 1 Location: 0.004 |
|
|
A unified theory of order flow, market impact and volatility Ecole Polytechnique, France We propose a microstructural model for the order flow in financial markets that distinguishes between core orders and reaction flow, both modeled as Hawkes processes. This model has a natural scaling limit that reconciles a number of salient empirical properties: persistent signed order flow, rough trading volume and volatility, and power-law market impact. In our framework, all these quantities are pinned down by a single statistic H_0, which measures the persistence of the core flow. Specifically, the signed flow converges to the sum of a fractional process with Hurst index H_0 and a martingale, while the limiting traded volume is a rough process with Hurst index H_0-1/2. No-arbitrage constraints imply that volatility is rough, with Hurst parameter 2H_0-3/2, and that the price impact of trades follows a power law with exponent 2-2H_0. The analysis of signed order flow data yields an estimate H_0 close to 3/4. This is not only consistent with the square-root law of market impact, but also turns out to match estimates for the roughness of traded volumes and volatilities remarkably well. |
| 10:00am - 10:40am | Coffee break 1 |
| 10:40am - 12:10pm | Statistics in natural sciences and technology Location: 0.001 Session Chair: Gaby Schneider Session Chair: Ansgar Steland |
|
|
Self-Normalization for CUSUM-based Change Detection in Locally Stationary Time Series FH Aachen, Germany
A novel self-normalization procedure for CUSUM-based change detection in the mean of a locally stationary time series is introduced. Classical self-normalization relies on the factorization of a constant long-run variance and a stochastic factor. In this case, the CUSUM statistic can be divided by another statistic proportional to the long-run variance, so that the latter cancels. Thereby, a tedious estimation of the long-run variance can be avoided.
Under local stationarity, the partial sum process converges to $\int_0^t \sigma(x) dBx$ and no such factorization is possible. To overcome this obstacle, a self-normalized test statistic is constructed from a carefully designed bivariate partial-sum process. Weak convergence of the process implies that the resulting self-normalized test attains asymptotic level α under the null hypothesis of no change, while being consistent against a broad class of alternatives. Extensive simulations demonstrate better finite-sample properties compared to existing methods. Applications to real data illustrate the method’s practical effectiveness.
Prior shift estimation for positive unlabeled data through the lens of kernel embedding 1: Warsaw University of Technology, Poland; 2: Institute of Computer Science; 3: Nicolas Copernicus University We study estimation of a class prior for unlabeled target samples which possibly differs from that of source population. Moreover, it is assumed that the source data is partially observable: only samples from the positive class and from the whole population are available (PU learning scenario). We introduce a novel direct estimator of the class prior which avoids estimation of posterior probabilities in both populations and has a simple geometric interpretation. It is based on a distribution matching technique together with kernel embedding in Reproducing Kernel Hilbert Space and is obtained as an explicit solution to an optimisation task. We establish its asymptotic consistency as well as an explicit non-asymptotic bound on its deviation from the unknown prior, which is calculable in practice. We study finite sample behaviour for synthetic and real data and show that the proposal works consistently on par or better than its competitors. Asymptotic studies of adapted threshold detectors based on density processes RWTH Aachen University, Germany Control statistics are widely used to monitor the quality of processes in various fields such as industry, healthcare, and machine learning. These statistics give an alarm when observed data exceed a threshold, traditionally set as a constant value to maintain a desired false alarm rate. Now we want to focus on a new setting: When monitoring a sequence of observations, there may be additional information that potentially affects the law of the observations, and we would like to change the design by using adapted thresholds, which are functions of the additional information. |
| 10:40am - 12:10pm | Discrete time series Location: 0.002 Session Chair: Christian H. Weiß |
|
|
Overview of the STINARMA Class of Models and its STINAR and STINMA Subclasses 1: Institute of Electronics and Informatics Engineering of Aveiro (IEETA) and Department of Electronics, Telecommunications and Informatics (DETI), University of Aveiro, Aveiro, Portugal; Intelligent Systems Associate Laboratory (LASI), University of Aveiro, Portugal.; 2: Center for Computational and Stochastic Mathematics (CEMAT), Department of Mathematics, IST, University of Lisbon, Lisbon, Portugal; 3: Department of Mathematics and Statistics, Helmut Schmidt University, Hamburg, Germany Spatio-temporal count data arise in many applied fields, where observations are collected over time across multiple spatial units. In these settings, it is crucial to jointly capture temporal and spatial dynamics. The spatio-temporal integer-valued autoregressive and moving average (STINARMA) class of models provides a flexible framework to address these challenges within the class of integer-valued processes. This work presents an overview of the STINARMA class of models, together with its main subclasses, those of the STINAR and STINMA models.The STINARMA can be viewed as the natural spatio-temporal extension of univariate INARMA models. Moreover, they are the integer counterpart of the continuous STARMA models, which is achieved by replacing the multiplication operator with the matrix binomial thinning operator and by considering component-wise independent discrete innovations. The general class of STINARMA models is introduced, followed by a discussion of its autoregressive and moving average subclasses. Key probabilistic properties are briefly presented through first- and second-order moments. Estimation approaches based on the method of moments, conditional least squares and conditional maximum likelihood are also outlined. The practical relevance of the STINARMA class is illustrated using spatio-temporal health data from Portugal and Germany, and its performance is compared with multivariate models that do not explicitly account for spatial dependence. References Martins, A., Scotto, M. G., Weiß, C. H., Gouveia, S. Space-time integer-valued ARMA modelling for time series of counts, Electronic Journal of Statistics, 17 (2), (2023), 3472-3511. Franke, J. Subba Rao, T. Multivariate First-Order Integer-Valued Autoregressions, Technical Report, University of Kaiserslaute, (1993). Pfeifer P. E., Deutsch S. J., A Three-Stage Iterative Procedure for Space-Time Modeling, Technometrics, 22 (1), (1980), 35-47. Steutel, F. W., Van Harn, K., Discrete Analogues of Self-Decomposability and Stability, The Annals of Probability, 7 (5), (1979), 893-899 Integer-valued random field models Helmut-Schmidt-Universität, Germany Ghodsi et al. (2012) have introduced the first-order integer-valued autoregressive model for count random fields as a planar analogue of the classical INAR(1) model, designed for count data observed on a regular lattice. We extend this framework to higher-order dependence structures and derive key stochastic properties of the resulting models. Building on this approach, we further propose two additional count random field models: the CINAR random field model and the INMA random field model. For each model, we investigate fundamental properties and provide a comparative analysis highlighting their respective strengths and limitations. Ghodsi, A., Shitan, M., & Bakouch, H. S. (2012). A first-order spatial integer-valued autoregressive SINAR (1, 1) model. Communications in Statistics-Theory and Methods, 41(15), 2773-2787. Influence network reconstruction from discrete time-series of count data modelled by multidimensional Hawkes processes University of Surrey, United Kingdom Identifying key influencers from time series data without a known prior network structure is a challenging problem in various applications, from crime analysis to social media. While much work has focused on event-based time series (timestamp) data, fewer methods address count data, where event counts are recorded in fixed intervals. We develop network inference methods for both batched and sequential count data. Here the strong network connection represents the key influences among the nodes. We introduce an ensemble-based algorithm, rooted in the expectation-maximization (EM) framework, and demonstrate its utility to identify node dynamics and connections through a discrete-time Cox or Hawkes process. For the linear multidimensional Hawkes model, we employ a minimization-majorization (MM) approach, allowing for parallelized inference of networks. For sequential inference, we use a second-order approximation of the Bayesian inference problem. Under certain assumptions, a rank-1 update for the covariance matrix reduces computational costs. We validate our methods on synthetic data and real-world datasets, including email communications within European academic communities. Our approach effectively reconstructs underlying networks, accounting for both excitation and diffusion influences. This work advances network reconstruction from count data in real-world scenarios. |
| 10:40am - 12:10pm | Multivariate Statistics and Copulas Location: 0.004 Session Chair: Sebastian Fuchs |
|
|
Measures and Models of Non-Monotonic Dependence 1: University of York, United Kingdom; 2: McGill University, Montreal, Canada; 3: University College Dublin, Ireland We propose a margin-free measure of bivariate association generalizing Spearman’s rho to the case of non- monotonic dependence that is defined in terms of two square integrable functions on the unit interval. We investigate properties of generalized Spearman correlation when the functions are piecewise continuous and strictly monotonic, with particular focus on the special cases where the functions are drawn from orthonormal bases defined by Legendre polynomials and cosine functions. For continuous random variables, generalized Spearman correlation is treated as a copula-based measure and shown to depend on a pair of uniform-distribution-preserving (udp) transformations determined by the underlying functions. We derive bounds for generalized Spearman correlation and we use a novel technique that we refer to as stochastic inversion of udp transformations to construct singular copulas that attain the bounds and parametric copulas with densities that interpolate between the bounds and model different degrees of non-monotonic dependence. We also propose sample analogues of generalized Spearman correlation and investigate their asymptotic and small-sample properties. Potential applications of the theory are demonstrated including: exploratory analyses of the dependence structures of datasets and their symmetries; elicitation of functions maximizing generalized Spearman correlation via expansions in orthonormal basis functions; and construction of tractable probability densities to model a wide variety of non-monotonic dependencies. Multivariate tail dependence: further insights with an application to the Spanish banking sector 1: Università del Salento, Italy; 2: Universidad de Valladolid, Spain Extending bivariate dependence concepts to higher dimensions is a challenging but essential task for a comprehensive understanding of multivariate dependence. Moreover, measuring overall dependence based on averages across the full domain of the joint distribution may fail to discern changes in dependence across different segments of the distribution, especially in the tails. In order to incorporate these features, we present the multivariate tail concentration function (TCF) as a graphical tool to assess both global and tail dependence. We show that this tool allows to represent multivariate dependence in a 2D plot regardless of the number of dimensions, it quantifies both lower and upper tail dependence at a finite scale, and it relates to multivariate Blomqvist’s beta. We propose to estimate the TCF non-parametrically using two methods and we compare their finite sample performance through a simulation study. To illustrate its practical application, we use the TCF to evaluate co-movements among the six Spanish banks included in the IBEX35 stock index. Multivariate Kendall regression coefficients University of Applied Sciences Merseburg, Germany In multivariate regression analysis, the multiple linear correlation coefficient is a commonly used association measure. This measure focuses on a linear relationship between a response variable and predictor variables. When moving away from the linearity of the functional relationship, then we arrive at Kendall's tau and multivariate versions, among others. In an earlier paper by the author (2021), the Kendall regression coefficient was introduced. Here, we extend the coefficient to vector responses Y and discuss properties of it. The coefficient we introduce describes to what degree the response variable Y can be approximated by a monotonous function of the regressors. These regressors are combined in a random vector. One advantage of this approach is that the association measure is based only on the copula (does not depend on marginal distributions), and is hence robust against outliers. |
| 10:40am - 12:10pm | Data Science Perspectives from Industry Location: 1.002 Session Chair: Rainer Göb |
|
|
Deploying Deep Learning for Real-Time Optical Sorting: A Case Study in Hazelnut Quality Control 1: prognostica GmbH; 2: IFSYS Integrated Feeding Systems GmbH Optical sorting is widely used in industrial quality control, yet conventional rule-based vision systems often struggle when quality cues are subtle, heterogeneous, or hard to formalize. We present an industry data science case study on deploying deep learning for real-time optical sorting of hazelnuts, driven by the practical need to grade product quality from fine-grained appearance characteristics under strict throughput and latency constraints. The talk traces the path from an early prototype to an industrialized system that has been transferred into a market-ready product and is operated in practice. We summarize the end-to-end solution: multi-camera image acquisition, a supervised learning pipeline built on a representative labeled dataset, domain-specific preprocessing and targeted data augmentation, and a neural image classifier designed for on-premise inference. We emphasize industrial aspects that proved central for making the system operational: formalizing expert grading into maintainable classes, managing imbalance and borderline cases during data preparation, data labeling and training, and setting decision thresholds based on acceptance criteria. We then cover deployment realities for industrial environments, e.g. latency, throughput, robustness, and the interface between the ML component and machine control. Finally, we describe how the solution was productized and extended beyond hazelnuts to additional crops, enabling new application scenarios and market opportunities for the customer. We conclude with practical considerations for lifecycle management and periodic re-calibration. Bridging the Gap: Operational Realities and Emerging Trends in Supply Chain Forecasting prognostica GmbH, Germany While forecasting remains a cornerstone of strategic decision-making, its industrial application involves challenges that extend beyond model accuracy. In the context of supply chain management, a forecast must not only be precise but also interpretable and actionable within specific operational constraints. This talk provides insights into how practitioners bridge the gap between theoretical models and business requirements, focusing on the following key areas:
The presentation demonstrates that the value of Generative AI in forecasting lies not only in potential accuracy gains but also in its capacity to handle unstructured context and significantly improve interactability with the forecasts. By highlighting these real-world requirements and current technical frontiers, the talk seeks to provide practical impulses and identify open questions for further academic research in the field of applied AI and time series analysis. |
| 10:40am - 12:10pm | High-dimensional statistics and learning Location: 1.012 Session Chair: Martin Wahl |
|
|
Supervised classification for Ornstein-Uhlenbeck diffusions with separation condition Humboldt University of Berlin, Germany We study binary supervised classification based on repeated independent observations of continuous sample paths. Our focus is a diffusion classification model in which the features follow an Ornstein-Uhlenbeck process with class-dependent drifts. We consider plug-in classifiers constructed from drift estimators and analyze the performance via the excess risk. Under a separation condition on the drift parameters, we establish upper bounds of the excess risk, which are explicitly parametrized by the separation distance quantifying the difficulty of the problem. Specifically, when the drift distance is bounded away from zero, the plug-in classifiers achieve a fast convergence rate of order n-1 (up to logarithmic factors) in the constant drift scenario. Furthermore, we discuss extensions of this framework to time-inhomogeneous drift functions. The theoretical approach utilizes the Wiener chaos representation and spectral theory to characterize the log-likelihood ratio as a quadratic form of Gaussian random variables, enabling a precise analysis of margin properties and concentration results. This extends the fast-rate results from classification problems with linear and Gaussian white noise models to dynamical diffusion systems with Gaussian structure under separation conditions. Asymptotic Bounds and Online Algorithms for Average-Case Matrix Discrepancy 1: Johns Hopkins University, USA; 2: FAU Erlangen-Nürnberg, Germany; 3: Yale University, USA
We study the matrix discrepancy problem in the average-case setting. Given a sequence of $m \times m$ symmetric matrices $A_1,\ldots,A_n$, its discrepancy is defined as the minimal spectral norm over all signed sums $\sum_{i=1}^n x_iA_i$ with $x_1,\ldots,x_n \in \{\pm1\}$. Our contributions are twofold. First, we study the asymptotic discrepancy of random matrices. When the matrices belong to the Gaussian orthogonal ensemble, we provide a sharp characterization of the asymptotic discrepancy and show that the limiting distribution is concentrated around $\Theta(\sqrt{nm}4^{-(1 + o(1))n/m^2})$, under the assumption $m^2 \ll n/\log{n}$. We observe that the trivial bound $O(\sqrt{nm})$ cannot be improved when $n \ll m^2$ and show that this phenomenon occurs for a broad class of random matrices. In the case $n = \Omega(m^2)$, we provide a matching upper bound. Second, we analyse the matrix hyperbolic cosine algorithm, an online algorithm for matrix discrepancy minimization due to Zouzias~(2011), in the average-case setting. We show that the algorithm achieves with high probability a discrepancy of $O(m\log{m})$ for a broad class of random matrices, including Wigner matrices with entries satisfying a hypercontractive inequality and Gaussian Wishart matrices.
Asymptotic confidence bands for centered purely random forests Karlsruhe Institute of Technology, Germany In this talk we will study asymptotic uniform confidence bands for centered purely random forests in a multivariate nonparametric regression setting. The most popular example in this class of random forests, namely the uniformly centered purely random forests, is well known to suffer from suboptimal rates. Therefore, a new type of purely random forests, called the Ehrenfest centered purely random forests, is proposed which achieves minimax optimal rates. Our main confidence band theorem applies to both random forests. The proof is based on an interpretation of random forests as generalized U-Statistics together with a Gaussian approximation of the supremum of empirical processes. |
| 12:10pm - 1:30pm | Lunch break 1 |
| 1:30pm - 3:30pm | New developments in nonparametric classification and estimation based on the nearest neighbor method Location: 0.001 Session Chair: Hajo Holzmann |
|
|
Chatterjee's graph correlation University of Washington, United States of America This talk will survey recent advances in understanding Chatterjee's nearest neighbor graph-based correlation coefficient. I will introduce, for the first time, a comprehensive theoretical framework for statistical inference based on this coefficient. The framework involves results on asymptotic normality, bias correction, and the (in)consistency of bootstrap methods. Nearest Neighbor Estimates for Dependent Data University of Manitoba, Canada This paper considers the nonparametric estimation problem for a class of nonlinear time series Nearest Neighbor matching: from Average Treatment Effects to Transfer Learning ENSAI-CREST, France Estimating some mathematical expectations from partially observed data and in particular missing outcomes is a central problem encountered in numerous fields such as transfer learning, counterfactual analysis or causal inference. Matching estimators, estimators based on k-nearest neighbors, are widely used in this context. Under suitable regularity conditions, one can show that the variance of such estimators can converge to zero at a parametric rate. However their bias can have a slower rate when the dimension of the covariates is larger than 2. This makes analysis of this bias particularly important. In this paper, we provide higher order properties of the bias. In contrast to the existing literature on this topic, we do not assume that the support of the target distribution of the covariates is strictly included in that of the source, and we discuss two geometric conditions on the support that prevent boundary bias issues. We show that these conditions are much more general than the usual convex support assumption, leading to an improvement of existing results. Furthermore, we show that the matching estimator studied by Abadie and Imbens (2006) for the average treatment effect can be asymptotically efficient when the dimension of the covariates is less than 4, a result only known in dimension 1. Multivariate Root-N-Consistent Smoothing Parameter Free Matching Estimators and Estimators of Inverse Density Weighted Expectations 1: Universität Rostock, Germany; 2: Philipps-Universität Marburg, Germany Expected values weighted by the inverse of a multivariate density or, equivalently, Lebesgue integrals of regression functions with multivariate regressors occur in various areas of applications, including estimating average treatment effects, nonparametric estimators in random coefficient regression models or deconvolution estimators in Berkson errors-in-variables models. The frequently used nearest-neighbor and matching estimators suffer from bias problems in multiple dimensions. By using polynomial least squares fits on each cell of the Kth-order Voronoi tessellation for sufficiently large K, we develop novel modifications of nearest-neighbor and matching estimators which again converge at the parametric root-n-rate under mild smoothness assumptions on the unknown regression function and without any smoothness conditions on the unknown density of the covariates. We stress that in contrast to competing methods for correcting for the bias of matching estimators, our estimators do not involve nonparametric function estimators and in particular do not rely on sample-size dependent smoothing parameters. We complement the upper bounds with appropriate lower bounds derived from information-theoretic arguments, which show that some smoothness of the regression function is indeed required to achieve the parametric rate. Simulations illustrate the practical feasibility of the proposed methods. |
| 1:30pm - 3:30pm | Discrete time series Location: 0.002 Session Chair: Christian H. Weiß |
|
|
Asymptotic Inference for Rank Correlations 1: Karlsruhe Institute of Technology; 2: Heidelberg Institute for Theoretical Studies; 3: Goethe University Frankfurt; 4: Helmut-Schmidt-University Kendall's tau and Spearman's rho are widely used tools for measuring dependence. Surprisingly, when it comes to asymptotic inference for these rank correlations, some fundamental results and methods have not yet been developed, in particular for discrete random variables and in the time series case, and concerning variance estimation in general. Consequently, asymptotic confidence intervals are not available. We provide a comprehensive treatment of asymptotic inference for classical rank correlations, including Kendall's tau, Spearman's rho, Goodman-Kruskal's gamma, Kendall's tau-b, and grade correlation. We derive asymptotic distributions for both iid and time series data, resorting to asymptotic results for U-statistics, and introduce consistent variance estimators. This enables the construction of confidence intervals and tests, generalizes classical results for continuous random variables and leads to corrected versions of widely used tests of independence. We analyze the finite-sample performance of our variance estimators, confidence intervals, and tests in simulations and illustrate their use in case studies. Inference for INAR Models with Structural Breaks: Classical and Bayesian Approaches 1: Universidade de Aveiro; CIDMA, Portugal; 2: ESTGA, Universidade de Aveiro; CIDMA, Portugal; 3: Universidade de Aveiro, Portugal Integer-valued autoregressive (INAR) models provide a flexible framework for modeling count time series through thinning operators that emulate autoregressive dynamics while respecting the discrete nature of the data. These models naturally accommodate both equidispersion and overdispersion, features commonly observed in count-valued processes. This paper investigates INAR models with structural breaks, with particular emphasis on the detection and estimation of parameter changes over time—an issue of critical importance in dynamic settings such as epidemics, policy interventions, and other regime-shifting phenomena. We consider both classical and Bayesian inferential approaches for identifying change points and estimating model parameters. The classical framework is based on maximum likelihood estimation, where structural changes are detected using a CUSUM-based procedure, followed by a focused grid search within a window centered around the candidate breakpoint. The Bayesian approach employs advanced Markov Chain Monte Carlo (MCMC) techniques, incorporating hidden Markov chains to model latent regimes and infer structural shifts probabilistically. A comprehensive simulation study is conducted under a variety of scenarios, including differing regime lengths and sample size proportions, and distributional characteristics. Finally, the proposed methodologies are illustrated through an application to real-world health indicator data, demonstrating their practical effectiveness in capturing complex dynamics and structural changes in count time series. Model diagnostics and semi-parametric inference for count time series 1: TU Dortmund University, Germany; 2: TU Dortmund University, Germany; 3: Helmut-Schmidt-University Hamburg, Germany; 4: Cyprus Academy of Sciences, Letters, and Arts, Cyprus For modeling the serial dependence in discrete-valued time series, various approaches have been proposed in the literature. In particular, models based on a recursive, autoregressive-type structure such as the integer-valued autoregressive (INAR) models for count time series are very popular in practice. While their estimation typically relies on purely parametric approaches that impose restrictive assumptions on the innovation distribution, we consider semi-parametric estimation techniques that jointly estimate the autoregressive coefficients and the innovation distribution without requiring parametric specification. Building on this, we propose a general semi-parametric bootstrap procedure for INAR models and prove its consistency for general classes of statistics that are functions of the estimated model coefficients and the estimated innovation distribution. This semi-parametric bootstrap approach can be leveraged for various statistical tasks such as goodness-of-fit testing, predictive inference, and dispersion analysis. Additionally, we introduce novel semi-parametric goodness-of-fit tests tailored for the INAR model class. Relying on the INAR-specific shape of the joint probability generating function, our approach allows for model validation of INAR models without specifying the parametric family of the innovation distribution. We derive the limiting null distribution of our proposed test statistics, prove consistency under fixed alternatives and discuss its asymptotic behavior under local alternatives. Moreover, when it comes to predictive inference for discrete-valued time series, this task cannot be implemented through the construction of prediction intervals as they are generally not able to retain a desired coverage level neither in finite samples nor asymptotically. To address this problem, we propose to reverse the construction principle by considering preselected sets of interest and estimating the corresponding predictive probability. The accuracy of this prediction is then evaluated by quantifying the uncertainty associated with the estimation of these predictive probabilities. In this context, we consider parametric and non-parametric approaches and derive asymptotic as well as bootstrap theory, which also covers the practically important case of model misspecification. Nonparametric symmetry tests for integer-valued time series Friedrich-Schiller-Universität Jena, Germany During the last years, there have been many proposals for modelling integer-valued time series. We propose tests of hypotheses related to certain symmetry and antisymmetry properties. For example, we consider the hypotheses that the conditional mean is an odd function or that the conditional variance is an even function. The proposed test statistics are nonparametric and have non-standard limit distributions. We show that the wild bootstrap offers a simple method of generating asymptotically correct critical values. The talk is based on joint work with Paul Doukhan and Christian Weiß. |
| 1:30pm - 3:30pm | Multivariate Statistics and Copulas Location: 0.004 Session Chair: Eckhard Liebscher |
|
|
Characterization of multi-way binary tables with uniform margins and fixed correlations 1: Politecnico di Torino, Italy; 2: Eindhoven University of Technology, the Netherlands; 3: Università di Genova, Italy In many applications involving binary variables, only pairwise dependence measures, such as correlations, are available. However, for multi-way tables involving more than two variables, these quantities do not uniquely determine the joint distribution, but instead define a family of admissible distributions that share the same pairwise dependence while potentially differing in higher-order interactions. In this talk, we introduce a geometric framework to describe the entire feasible set of such joint distributions with uniform margins. We show that this admissible set forms a convex polytope, analyze its symmetry properties, and characterize its extreme rays. These extremal distributions provide fundamental insights into how higher-order dependence structures may vary while preserving the prescribed pairwise information. Unlike traditional methods for table generation, which return a single table, our framework makes it possible to explore and understand the full admissible space of dependence structures, enabling more flexible choices for modeling and simulation. We illustrate the usefulness of our theoretical results through examples and a real case study on rater agreement. Copula robustness in quantitative risk management Saarland University, Germany Characteristics of d-variate risks, such as downside risk measures of aggregate positions or optimal portfolio values, play a central role in financial and actuarial applications. This talk addresses the question of when such characteristics are robust to (small) misspecifications in the copula. DIRECTIONAL FOOTRULE-COEFFICIENTS University of Almería, Spain Measures of association based on ranks, such as Spearman’s footrule[1], play a central role in multivariate statistics due to their robustness and invariance properties. However, classical versions of these coefficients are often unable to capture directional dependence structures that arise in high-dimensional settings. Motivated by this limitation and by the newly defined coefficients described subsequently[2] [3], we introduce a novel family of directional Spearman’s footrule coefficients designed to quantify multivariate dependence along prescribed directions in the unit d-dimensional hypercube. The proposed coefficients are formulated within the framework of copula theory, which allows for a clear separation between marginal behavior and the underlying dependence structure. Our construction extends the classical Spearman’s footrule by incorporating directional information, enabling the detection of dependence patterns that remain undetected by standard measures. We establish a general definition for arbitrary dimensions and directions and investigate its main theoretical properties. In particular, we analyze their behavior under independence and maximal positive dependence, their relation to stochastic orders, as well as their relationship with marginal distributions and lower-dimensional structures. These properties are shown to be consistent with those of the classical footrule coefficient. To facilitate practical implementation, we also introduce nonparametric estimators based on ranks. These estimators are easy to compute and suitable for multivariate data. Their asymptotic behavior is discussed, highlighting consistency and stability properties analogous to those of existing rank-based dependence measures. Several illustrative examples are provided to demonstrate the usefulness of the proposed coefficients. Explicit expressions are derived for well-known families of d-copulas, including the Farlie–Gumbel–Morgenstern and Cuadras–Augé, allowing for a detailed analysis of how directional dependence varies with model parameters. These examples show that the proposed coefficients are able to distinguish different directional dependence patterns even when classical global measures coincide. Overall, this work provides a new tool for directional dependence analysis in multivariate settings, complementing existing rank-based measures and offering a finer understanding of complex dependence structures with applications in finance, reliability, and multivariate risk analysis. [1] Spearman, C. (1906). ‘Footrule’ for measuring correlation. Brithis Journal of Psychology, 2, 89-108. [2] Úbeda-Flores, M. (2004). Multivariate versions of Blomqvist’s beta and Spearman’s footrule. Ann. Inst. Statist. Math., 57(4), 781-788. [3] Decancq, K., Pérez, A., Prieto-Alaiz, M. (2025). Multivariate Dependence Based on Diagonal Sections: Spearman’s Footrule and Related Measures. In: Steland, A., Rafajłowicz, E., Parolya, N. (eds) Stochastic Models, Statistics and Their Applications. SMSA 2024. Springer Proceedings in Mathematics & Statistics, vol 499. Springer, Cham. Estimating Portfolio Risk with Product Copulas: A GARCH-EVT Approach Applied to Financial Data Hochschule Merseburg, Germany This talk introduces a sophisticated GARCH-EVT-Copula framework designed A key innovation presented is the application of product copulas to model the Our empirical analysis demonstrates the superior performance of the product |
| 1:30pm - 3:30pm | Statistics in sports Location: 1.002 Session Chair: Jakob Söhl |
|
|
The Best of Both Worlds: Predicting Coverage Schemes in American Football with Supervised and Unsupervised Learning 1: TU Dortmund; 2: WU Vienna; 3: Bielefeld University Choosing between man and zone coverage is one of the most critical strategic decisions a defensive coordinator must make before each offensive play in American football. In simple terms, in man coverage each defender is assigned to guard a specific offensive player, while zone coverage requires defenders to protect designated areas of the field. This choice fundamentally shapes how the defense reacts to offensive formations and movements. Traditionally, experienced offensive coordinators and quarterbacks rely on visual cues, such as defenders’ alignment or pre-snap motion, to infer these defensive schemes. However, with the increasing availability of high-resolution player tracking data, statistical models can now uncover such tactical patterns quantitatively rather than relying solely on expert intuition. In this project, we first employ an elastic net and an XGBoost classifier to predict whether a defense is in man or zone coverage based on all players’ positions once both teams are set before the snap. The models thus captures spatial configurations that often reveal underlying defensive intentions. In a second step, we incorporate dynamic information from pre-snap player movements. Finally, in a third step, we employ features derived from a hidden Markov model (HMM). Specifically, we use an HMM to represent defenders’ movement trajectories over time. The hidden states correspond to potential offensive players being covered by each defender. From the decoded state sequences, we extract summary statistics, such as the number of state (defender) switches. Including these HMM-based features in the aforementioned models significantly enhances the model’s predictive accuracies. Beyond the pure classification performance, our approach also enables deeper tactical analyses. For instance, it allows us to explore how pre-snap motion helps offenses identify defensive coverages more effectively. Comparing these pre- and post-motion probabilities provides insight into how well offensive movements reveal defensive strategies. Overall, this framework demonstrates how modern machine learning techniques in combination with a statistical model can provide quantitative insights into complex team sports tactics. While developed within an American football context, the methodology may generalize to other sports where spatial positioning and interaction dynamics play similarly crucial roles. Modelling momentum in tennis: A latent-state approach to point outcomes and rally lengths 1: Bielefeld University, Germany; 2: TU Dortmund, Germany Tennis matches are often characterised by momentum shifts – i.e., changes in match dynamics over time – marked by transitions between phases where either player 1 or player 2 dominates. While dominance is clearly reflected in a player’s point wins, rally lengths provide additional valuable information for modelling momentum; short rallies suggest strong momentum, whereas long rallies and point losses indicate pressure. To effectively model momentum shifts, we hence propose considering both the outcomes of the points and the rally lengths. These sequentially observed outcomes reflect the current dynamics of the match (i.e., the level of pressure a player exerts on their opponent), which we regard as an unobserved state process. Thus, we employ a latent-state approach to investigate these momentum shifts. Specifically, we model the outcomes of server wins and rally lengths jointly using Markov-modulated marked Poisson processes (MMMPPs). This flexible framework allows us to relate the events (server wins or loses the point) and the event times (rally length) to an underlying latent state process, modelled as a continuous-time Markov chain. Its states determine the distribution of the outcomes and can be interpreted as proxies for the players’ momentum. For data from all Grand Slam tournaments from 2016 to 2024, we identify momentum shifts within tennis matches using MMMPPs with two latent states, accounting for player- and match-specific effects such as player rankings and court surfaces. The Accuracy–Complexity Trade-Off in the Expected Threat model for Football 1: TU Delft, The Netherlands; 2: AFC Ajax, The Netherlands The Expected Threat model is a possession value model in football (soccer) with a Markov chain structure that allows for interpretation and visualization. To create a Markov chain, the pitch is discretized into different Markov states. However, selecting the right discretization of the pitch is still a challenging design choice. A model with more game states can better distinguish between different scenarios, but has less samples per state when estimating the Markov chain. This creates a trade-off between the model complexity in terms of the number of Markov states and the accuracy of the probability estimates. Theoretical analysis of the model gives error bounds, but interpretation of the results indicates that these might be on the conservative side. Simulations provide a more accurate characterization of the model’s error, which is indeed more optimistic than the theoretical bound. Finally, these insights are converted into a practical rule of thumb to help practitioners choose the right balance between the number of Markov states and accuracy of the probability estimates of the Expected Threat model. |
| 1:30pm - 3:30pm | Computational Biostatistics Location: 1.012 Session Chair: Dennis Dobler |
|
|
Computational and Biostatistical Challenges in Polygenic Score Modelling and Gene–Environment Integration 1: IUF - Leibniz Research Institute for Environmental Medicine; 2: TU Dortmund University Polygenic scores (PGS) quantify genetic predisposition to complex traits and clinical outcomes based on genotype data. This talk addresses recent computational and biostatistical challenges in PGS modelling, including their integration with environmental risk factors. First, training PGS models on high-dimensional and large-scale genotype data with hundreds of thousands of genetic variants and individuals requires scalable yet interpretable statistical learning methods. Second, the transferability of PGS models to diverse populations with different ancestries remains limited, as models are typically trained on cohorts predominantly of European ancestry. Third, the evaluation of predictive performance is complicated by different and sometimes conflicting definitions of the commonly used R-squared measure on test data. To address these challenges, scalable statistical learning approaches for PGS modelling based on individual-level genotype data are presented, including boosting and anchor regression. Finally, open problems and directions for future research are highlighted, with the aim of improving robustness, interpretability and gene–environment integration in personalized medicine. Robust Feature Selection for High-Dimensional Mixtures of Cox Models University of Augsburg, Germany Time-to-event analysis is fundamental for studying patient survival in modern biomedical research, particularly in the presence of high-dimensional covariate information. When survival data are collected over long time horizons, population heterogeneity naturally arises due to evolving clinical practices and patient characteristics. Mixtures of Cox proportional hazards models offer an effective way to account for such heterogeneity by modeling latent subpopulations with distinct risk profiles. In high-dimensional settings, feature selection is crucial for improving model interpretability and predictive performance. This talk presents a robust feature selection approach for mixtures of Cox models based on a combined ℓ1–ℓ2 penalty, which encourages sparsity while stabilizing estimation across mixture components. The resulting optimization problem is non-smooth and challenging to solve within mixture models. We address this challenge by developing an efficient Expectation–Maximization (EM) algorithm that effectively handles the non-smooth penalty structure. Empirical results demonstrate that the proposed method improves patient-specific survival time prediction across heterogeneous populations while achieving stable and interpretable feature selection. A regularized Cox model for selecting interactions and time-varying covariate effects 1: Institute for Medical Biometry, Informatics and Epidemiology, Medical Faculty, University of Bonn; 2: Department of Mathematics, Informatics and Technology, Koblenz University of Applied Sciences, RheinAhrCampus Remagen, The Cox proportional hazards model is a widely used method for analyzing clinical time-to-event data. In its standard form, the Cox model assumes the covariate effects on the hazard function to be constant over time. However, in many clinical settings, covariate effects may vary with time, and covariate interactions may significantly influence survival. Selecting interactions and time-varying effects within the Cox model framework may be challenging and often requires manual pre-screening followed by model selection steps. These selection steps are often carried out through automated stepwise procedures, which, however, can be unstable or even infeasible—particularly if a large number of potential effects is considered. We introduce a linked-shrinkage adaptive elastic net procedure for selecting two-way interactions and time-varying effects in Cox regression models. The proposed approach integrates an adaptive elastic net with penalty weights derived from an initial ridge regression that includes main effects only. Time-varying effects are modeled as piecewise constant functions. Penalty weights for interactions and time-varying terms are specified using a linked-shrinkage strategy based on the pre-estimated main effects, such that these effects are penalized more strongly than the main effects. We assessed the proposed modeling approach through a simulation study based on Weibull-distributed survival times, incorporating various structures of time-varying covariate effects. Using a simulation study, we compared the proposed method with several established approaches, including the classical elastic net extended to the Cox regression model. Model performance was assessed in terms of the mean squared error (MSE) of the estimated survival probabilities and the accuracy of variable selection. The proposed method reliably identified true time-varying and two-way interaction effects. The true positive rates ranged between 80%-90% depending on the scenario. Compared to standard regularized Cox regression models, the proposed method yielded better performance in terms of MSE and the ability to select informative main/interaction/timevarying effects in a more precise way. Furthermore, we illustrate the proposed approach by analyzing real-world data from the National Cancer Institute Surveillance, Epidemiology, and End Results (SEER) program. By addressing the limitations of manual covariate selection and stepwise procedures, the proposed method extends penalized estimation techniques to Cox regression with time-varying coefficients. Further, it facilitates the simultaneous selection of relevant interaction terms and time-varying covariate effects. Inferring Individual-Level Cell Type-Specific Transcriptomic Profiles from Bulk RNA-Seq Using a Bayesian Hierarchical Model University of North Carolina Wilmington, United States of America The high cost of single-cell sequencing often compels large cohort studies to rely on bulk RNA-seq, which presents challenges in resolving tissue heterogeneity and understanding the roles of individual cell types. In bulk RNA-seq analysis, deconvolution is essential for extracting cell-type-specific information. Most tools focus on estimating cell type proportions, but only a few aim to infer cell-type-specific gene expression profiles (ctsGEPs). Among these, very few estimate ctsGEPs at the individual sample level. The technical challenges of this task highlight the need for more advanced approaches capable of generating accurate individual-level ctsGEP estimates. Such estimates are critical for downstream analyses, including cell-type-specific differential expression and expression quantitative trait locus studies. To address this, we developed a novel deconvolution method to estimate individual-level ctsGEPs and cell type proportions simultaneously from bulk RNA-seq data. Using a hierarchical Bayesian framework, our method captures the stochastic variation of ctsGEPs across individuals. Parameters are estimated via Markov Chain Monte Carlo (MCMC), with hyperparameters optimized for robust inference. We benchmarked our method using 48 in silico mixtures generated from single-cell RNA-seq data of human brain donors. The results demonstrated strong performance, with correlations of ~0.9 for ctsGEP estimates and >0.6 for gene expression variation across samples for ~80% of genes. Our method outperformed existing tools, reducing Root-Mean-Square Errors by ~16%. Additionally, we showcased its application in cell-type-specific differential expression analysis. Our method provides a powerful tool to computationally unravel cell-type-specific expression profiles in bulk RNA-seq data, enabling advances in understanding cellular heterogeneity in biological and pathological contexts. |
| 3:30pm - 4:00pm | Coffee break 2 |
| 4:00pm - 5:00pm | Plenary Lecture 2 Location: 0.004 |
|
|
Statistical Optimal Transport in Action: From Theory to Applications University of Göttingen, Germany While optimal transport has been a long standing mathematical, physical and economic concept for more than two centuries, recent developments in statistics, optimization and machine learning suggests its use as a tool for modern data analysis. Extensions, such as Gromov-Wasserstein transport respect the inner metric structure of data sets and have been proven to be useful for image registration and object matching. In this talk we introduce some basic statistical methods related to optimal transport and illustrate these with examples from cell biology and biometric identification. |
| 5:05pm - 6:35pm | Applied Econometrics Location: 0.001 Session Chair: Yannick Hoga |
|
|
The impact of central bank backstops on sovereign risk premia: Evidence from the ECB's Transmission Protection Instrument 1: European Central Bank, Germany; 2: European Central Bank, Germany We study the effects of central bank backstops on sovereign risk premia using the Eurosystem’s Transmission Protection Instrument (TPI) announced in July 2022. We develop a nonlinear non-Gaussian state-space model that decomposes euro area sovereign yields into expected short rates, a common term premium, and country-specific default, redenomination, liquidity, and convenience premia. Structural shocks are identified through heteroscedasticity and fat tails. Using euro area data from 2015 to 2025, we extract latent risk premia and assess the impact of the TPI using event-time and differences-in-differences designs. The results show that the TPI primarily increased the convenience value of sovereign bonds and reduced the volatility of a subset of shocks, while leaving other risk premia largely unchanged. Lower convenience-adjusted yields partially dampened the transmission of policy rate hikes to medium-term sovereign yields. Forecast Combination for Tail Risk: Virtues of the Harmonic Mean University of Freiburg, Germany This paper examines the properties of the loss functions used for forecasting Value-at-Risk (VaR) and Expected Shortfall (ES). We show that the weighted arithmetic average commonly used to construct a forecast combination utilises the convexity property of the loss function only in case of Value-at-Risk. This paper introduces a novel forecasting combination approach for Expected Shortfall, which is constructed using weighted harmonic means. We show that only in this case the insurance against model risk is guaranteed. To construct combination weights consistent with this aggregation result, we propose a novel forecast combination for tail risk measures based on the Bagged Pretested Forecast Combination (BPFC) algorithm. The combination weights assigned to candidate models are determined by their predictive performance using the Model Confidence Set (MCS) test. Unlike many traditional combination methods, BPFC adapts to changing market conditions while simultaneously facilitating model selection and improving forecast stability. We evaluate the performance of forecasting combinations for VaR and ES within the framework of consistent loss functions, highlighting the role of convexity in performance improvements. Our results show that the advantages of combining forecasts are especially evident when there is substantial disagreement among candidate models, a situation that commonly arises during turbulent financial periods. To empirically validate our approach, we apply it to a dataset of 90 stocks spanning various market capitalizations and covering periods of severe financial stress, including the Global Financial Crisis and the COVID-19 pandemic. The results illustrate the ability of BPFC to dynamically select and combine the most effective models from a pool of over 60 candidates, continuously adjusting weights based on model’s forecasting performance and evolving market conditions. Systemic Risk Surveillance 1: Goethe University Frankfurt, Germany; 2: University Duisburg-Essen, Germany Following several episodes of financial market turmoil in recent decades, changes in systemic risk have drawn growing attention. Therefore, we propose surveillance schemes for systemic risk, which allow to detect misspecified systemic risk forecasts in an “on-line” fashion. This enables daily monitoring of the forecasts while controlling for the accumulation of false test rejections. Such online schemes are vital in taking timely countermeasures to avoid financial distress. Our monitoring procedures allow multiple series at once to be monitored, thus increasing the likelihood and the speed at which early signs of trouble may be picked up. The tests hold size by construction, such that the null of correct systemic risk assessments is only rejected during the monitoring period with (at most) a pre-specified probability. Monte Carlo simulations illustrate the good finite-sample properties of our procedures. An empirical application to US banks during multiple crises demonstrates the usefulness of our surveillance schemes for both regulators and financial institutions. |
| 5:05pm - 6:35pm | Statistical Inverse Problems Location: 0.002 Session Chair: Frank Werner |
|
|
Linear methods for non-linear inverse problems 1: Delft University of Technology; 2: Bocconi University, Italy We propose a novel Bayesian linearization approach for non-linear PDE constrained inverse problems. We split the non-linear inverse problem into a linear statistical and a non-linear analytic component. We derive optimal posterior contraction rates, reliable uncertainty quantification, data driven tuning and scalable approximations. The general approaches is applied to specific examples, including Darcy flow and heat equation with absorption term. Learning with Heavy-tailes TU Braunschweig, Germany We examine the performance of ridge regression in reproducing kernel Hilbert spaces in the presence of noise that exhibits a finite number of higher moments. We establish excess risk bounds consisting of subgaussian and polynomial terms based on the well known integral operator framework. The dominant subgaussian component allows to achieve convergence rates that have previously only been derived under subexponential noise - a prevalent assumption in related work from the last two decades. These rates are optimal under standard eigenvalue decay conditions, demonstrating the asymptotic robustness of regularized least squares against heavy-tailed noise. Our derivations are based on a Fuk-Nagaev inequality for Hilbert-space valued random variables. Comparing regularisation paths of (conjugate) gradient estimators in ridge regression 1: Humboldt-Universität zu Berlin, Germany; 2: Aarhus Universitet, Denmark We consider standard gradient descent, gradient flow and conjugate gradients as iterative algorithms for minimising a penalised ridge criterion in linear regression. While it is well known that conjugate gradients exhibit fast numerical convergence, the statistical properties of their iterates are more difficult to assess due to inherent non-linearities and dependencies. On the other hand, standard gradient flow is a linear method with well-known regularising properties when stopped early. By an explicit non-standard error decomposition we are able to bound the prediction error for conjugate gradient iterates by a corresponding prediction error of gradient flow at transformed iteration indices. This way, the risk along the entire regularisation path of conjugate gradient iterations can be compared to that for regularisation paths of standard linear methods like gradient flow and ridge regression. In particular, the oracle conjugate gradient iterate shares the optimality properties of the gradient flow and ridge regression oracles up to a constant factor. Numerical examples show the similarity of the regularisation paths in practice. |
| 5:05pm - 6:35pm | Inference in Wasserstein Spaces and Optimal Transport Location: 0.004 Session Chair: Ansgar Steland |
|
|
Statistical Aspects of Optimal Transport: Regularization, Estimation, and Applications University of Twente, The Netherlands In recent years, statistical methodology based on optimal transport (OT) witnessed a considerable increase in practical and theoretical interest. A central reason for this trend is the ability of optimal transport to efficiently compare data in a geometrically meaningful way. This development was further amplified by computational advances spurred by the introduction of entropy regularized optimal transport (EOT). In applications, the OT or EOT cost are often estimated through an empirical plug-in approach, raising statistical questions about the performance and uncertainty of these estimators. This talk will survey recent theoretical and methodological insights to these topics and discusses future opportunities. This talk is based on joint work with Thomas Staudt, Marcel Klatt, Michel Groppe, Alberto-Gonzáles-Sanz, Gilles Mordant, Christoph Weitkamp, and Axel Munk. On the cut-offs of Optimal Transport based statistical tests University of British Columbia, Canada Tests for equality of distributions based on Optimal Transport functionals are often referred to as being not distribution free: asymptotic laws for tests statistics depend on the underlying true distributions, and this dependence seems unavoidable. Here we show that these tests are ``almost" distribution free, in a sense that there exist cut-offs independent of the true distributions that result in tests with given level of significance. These cut-offs are easy to compute and may serve as a rule-of-thumb-type heuristics, making Optimal Transport based tests more accessible for practical applications. Detecting change-points of univariate time series using the empirical Wasserstein distance 1: RWTH Aachen University, Germany; 2: Delft University of Technology, Netherlands In this talk we are interested in detecting change-points of univariate nonstationary time series in a nonparametric setting. We introduce statistics based on the Wasserstein distance between local empirical distribution functions of the time series which are suitable to detect change-points. The one-dimensional Wasserstein distance is characterized by the sequential quantile process, and we show that this weakly converges to a Gaussian limit. Due to the nonlinearity of the quantile process, difficulties arise from the localization. A new Bahadur representation result is needed to address this, which allows us to consider the asymptotic behavior of the empirical process instead of the quantile process. The proof of this requires further study of the modulus of continuity of the empirical process. As the limit distributions of the test statistics depend on the unknown underlying distributions, a Gaussian multiplier bootstrap scheme is introduced. Lastly, a simulation study shows how well the significance level is retained under the null hypothesis of no change, and an outlook towards the power of the tests will be given. |
| 5:05pm - 6:35pm | Advances in Latent Variable Models Location: 1.002 Session Chair: Daniele Tancini |
|
|
A multilevel discrete latent variable model for joint modeling of response accuracy and times 1: University of Milano-Bicocca, Italy; 2: University of Perugia, Italy In recent years, the widespread adoption of computer-based testing has produced large volumes of data on examinee behavior. Beyond traditional binary indicators of correct responses, these datasets now typically include item-level response times, providing a richer and more informative perspective on the performance. The Bradley–Terry Stochastic Block Model University College Dublin, Ireland The Bradley-Terry model is widely used for the analysis of pairwise comparison data and, in essence, produces a ranking of the items under comparison. We embed the Bradley-Terry model within a stochastic block model, allowing items to cluster. The resulting Bradley-Terry SBM (BT-SBM) ranks clusters so that items within a cluster share the same tied rank. We develop a fully Bayesian specification in which all quantities-the number of blocks, their strengths, and item assignments-are jointly learned via a fast Gibbs sampler derived through a Thurstonian data augmentation. Despite its efficiency, the sampler yields coherent and interpretable posterior summaries for all model components. Our motivating application analyzes men's tennis results from ATP tournaments over the seasons 2000-2022. We find that the top 100 players can be broadly partitioned into three or four tiers in most seasons. Moreover, the size of the strongest tier was small from the mid-2000s to 2018 and has increased since, providing evidence that men's tennis has become more competitive in recent years. A latent space approach for jointly modelling social influence on binary outcomes in networks 1: University of Cambridge, United Kingdom; 2: University College Dublin, Ireland A central task in network analysis is to model social influence, that is, how individual behaviours and outcomes are shaped by their social environment. Classical regression models are not suitable for this purpose, as they frequently rely on independence assumptions that are violated in network data, where individuals' behaviours are inherently interdependent. Although several methods have been proposed to address this problem, existing approaches either treat the network as fixed, rely on multi-step estimation procedures, or are limited to continuous outcome variables. |
| 5:05pm - 6:35pm | Contributions to Computational Biostatistics and Data Science Location: 1.012 Session Chair: Dennis Dobler |
|
|
Bootstrap-based inference in regression using jackknife pseudo-observations 1: RWTH Aachen University, Germany; 2: Aarhus University, Denmark The pseudo-observation regression approach provides a flexible alternative to the omnipresent proportional hazards model when modeling time-to-event outcomes. In this approach, estimands representable as expectations are fitted to regression models using covariates of interest. Exemplary estimands that fit this framework are the restricted mean time lost (in competing risks models) or the survival function at a fixed time-point (in simple survival models). Likelihood-Based Inference for Dirichlet Mixture Models via Unconstrained Parameterization 1: TU Kaiserslautern, Germany; 2: LMU Munich, Germany Dirichlet mixture models (DMMs) provide a flexible and interpretable framework for clustering and modeling compositional data and have found widespread application in genomics, ecology, and the social sciences. Despite their popularity, formal likelihood-based inference for DMM parameters remains underdeveloped, primarily due to the presence of simplex constraints on mixture weights and the complex dependence structure induced by latent component memberships. In this paper, we develop a unified framework for classical likelihood-based inference in Dirichlet mixture models by working on an unconstrained parameterization that combines an additive log-ratio transformation of the mixture weights with the original Dirichlet concentration parameters. Within this framework, we derive closed-form expressions for score functions and observed Fisher information matrices, including full cross-component information terms obtained via the Louis identity. These results enable the construction of Wald, score (Lagrange Multiplier), and likelihood ratio tests for a broad class of regular parametric hypotheses, including fixed-value restrictions and equality constraints across mixture components. We show how the proposed methods apply seamlessly to both soft and hard EM-based estimation schemes and provide a numerically stable implementation that yields consistent standard errors and confidence intervals on the original parameter scale. Through simulation experiments and a real-data application, we demonstrate that the proposed inferential procedures perform well in finite samples and provide meaningful uncertainty quantification for DMM parameters. |
| 6:40pm - 8:30pm | Welcome Reception |

