Publications
List of publications and working papers. In case I am absent-minded and do not update this page, visit the [scholar page](https://scholar.google.com/citations?user=s2uxUGwAAAAJ)
2026
- PreprintUsing protein language models for pangenome constructionNiels Jakob Larsen, Pep Charusanti, Henry Webel, and 3 more authors2026
- PreprintAre We Lost in the Woods? Detecting Silent Semantic Faults for Random Forest Classifiers with Data-informed Static AnalysisWillem Meijer, Louis Ohl, Kristian Sandahl, and 1 more author2026
2025
- EHJ Dig. Hea.Unsupervised Machine Learning Analysis to Enhance Risk Stratification in Patients with Asymptomatic Aortic StenosisMarie-Ange Fleury, Louis Ohl, Lionel Tastet, and 16 more authorsEuropean Heart Journal - Digital Health, Oct 2025
There is a lack of studies investigating the pathophysiologic and phenotypic distinctiveness of aortic stenosis (AS). This heterogeneity has important implications for identifying optimal intervention timing and potential medical management. This study seeks to identify phenogroups of AS using unsupervised machine learning to improve risk stratification.A total of 349 patients with asymptomatic AS from the PROGRESSA study were included in this analysis. Echocardiographic, clinical and blood sample data were used in the unsupervised clustering process. Longitudinal echocardiographic data were used to evaluate AS progression.Five clusters of patients were revealed using 18 variables selected by an unsupervised machine learning algorithm. Amongst them, aortic valvular phenotype, mean gradient, peak jet velocity (Vpeak), and left ventricle stroke volume were selected as discriminatory variables. Following the clustering process, characteristics differed between clusters, including age, body mass index, and sex ratio (all p<0.001). Of note, cluster 1 showed higher AS severity at baseline with significantly higher initial Vpeak (344 [314; 376] cm/s) and calcium score (1257 [806; 1837]UA) (p<0.001). Patients from cluster 1 had a faster AS progression (progression of Vpeak=22 [9; 39] cm/s/year), and calcium score (213 [111; 307] UA/year) (p<0.001). Cluster 1 was also associated with a higher composite risk of mortality and aortic valve replacement when adjusted for age, sex, and baseline AS severity (p<0.001).Artificial intelligence-guided phenotypic classification revealed 5 distinct groups and enhanced risk stratification of patients with AS. This approach may be useful to optimize and individualize medical and interventional management of AS.
@article{fleury_unsupervised_2025, title = {Unsupervised {Machine} {Learning} {Analysis} to {Enhance} {Risk} {Stratification} in {Patients} with {Asymptomatic} {Aortic} {Stenosis}}, issn = {2634-3916}, url = {https://doi.org/10.1093/ehjdh/ztaf115}, doi = {10.1093/ehjdh/ztaf115}, journal = {European Heart Journal - Digital Health}, author = {Fleury, Marie-Ange and Ohl, Louis and Tastet, Lionel and Leclercq, Mickaël and Precioso, Frédéric and Mattei, Pierre-Alexandre and Capoulade, Romain and Abdoun, Kathia and Bédard, Élisabeth and Arsenault, Marie and Beaudoin, Jonathan and Bernier, Mathieu and Salaun, Erwan and Bernard, Jérémy and Shen, Mylène and Hecht, Sébastien and Côté, Nancy and Droit, Arnaud and Pibarot, Philippe}, month = oct, year = {2025}, pages = {ztaf115}, } - ACM Comp. Sur.A Tutorial on Discriminative Clustering and Mutual InformationLouis Ohl, Pierre-Alexandre Mattei, and Frederic PreciosoACM Comput. Surv., Oct 2025Place: New York, NY, USA
To cluster data is to separate samples into distinctive groups that should ideally have some cohesive properties. Today, numerous clustering algorithms exist, and their differences lie essentially in what can be perceived as “cohesive properties”. Therefore, hypotheses on the nature of clusters must be set: they can be either generative or discriminative. As the last decade witnessed the impressive growth of deep clustering methods that involve neural networks to handle high-dimensional data often in a discriminative manner; we concentrate mainly on the discriminative hypotheses. In this article, our aim is to provide an accessible historical perspective on the evolution of discriminative clustering methods and notably how the nature of assumptions of the discriminative models changed over time: from decision boundaries to invariance critics. We notably highlight how mutual information has been a historical cornerstone of the progress of (deep) discriminative clustering methods. We also show some known limitations of mutual information and how discriminative clustering methods tried to circumvent those. We then discuss the challenges that discriminative clustering faces with respect to the selection of the number of clusters. Finally, we showcase these techniques using the dedicated Python package, GemClus , that we have developed for discriminative clustering.
@article{ohl_tutorial_2025, title = {A {Tutorial} on {Discriminative} {Clustering} and {Mutual} {Information}}, volume = {58}, issn = {0360-0300}, url = {https://doi.org/10.1145/3748255}, doi = {10.1145/3748255}, number = {4}, journal = {ACM Comput. Surv.}, publisher = {Association for Computing Machinery}, author = {Ohl, Louis and Mattei, Pierre-Alexandre and Precioso, Frederic}, month = oct, year = {2025}, note = {Place: New York, NY, USA}, keywords = {Clustering, discriminative models, mutual information, neural networks, unsupervised learning}, pages = {1--36}, } - Discriminative Ordering Through Ensemble ConsensusLouis Ohl and Fredrik LindstenIn Proceedings of the Forty-first Conference on Uncertainty in Artificial Intelligence, Jul 2025
Evaluating the performance of clustering models is a challenging task where the outcome depends on the definition of what constitutes a cluster. Due to this design, current existing metrics rarely handle multiple clustering models with diverse cluster definitions, nor do they comply with the integration of constraints when available. In this work, we take inspiration from consensus clustering and assume that a set of clustering models is able to uncover hidden structures in the data. We propose to construct a discriminative ordering through ensemble clustering based on the distance between the connectivity of a clustering model and the consensus matrix. We first validate the proposed method with synthetic scenarios, highlighting that the proposed score ranks the models that best match the consensus first. We then show that this simple ranking score significantly outperforms other scoring methods when comparing sets of different clustering algorithms that are not restricted to a fixed number of clusters and is compatible with clustering constraints.
@inproceedings{ohl_discriminative_2025, series = {Proceedings of {Machine} {Learning} {Research}}, title = {Discriminative {Ordering} {Through} {Ensemble} {Consensus}}, volume = {286}, url = {https://proceedings.mlr.press/v286/ohl25a.html}, booktitle = {Proceedings of the {Forty}-first {Conference} on {Uncertainty} in {Artificial} {Intelligence}}, publisher = {PMLR}, author = {Ohl, Louis and Lindsten, Fredrik}, editor = {Chiappa, Silvia and Magliacane, Sara}, month = jul, year = {2025}, pages = {3252--3271}, }
2024
- ThesisStatistical learning applied to cardiology: discriminative clustering and aortic stenosis phenogroupsLouis OhlUniversité Côte d’Azur; Université Laval, Aug 2024
- Stat. and Comp.Sparse and geometry-aware generalisation of the mutual information for joint discriminative clustering and feature selectionStatistics and Computing, 2024
@article{ohl_sparse_2024, title = {Sparse and geometry-aware generalisation of the mutual information for joint discriminative clustering and feature selection}, volume = {34}, number = {5}, journal = {Statistics and Computing}, publisher = {Springer}, author = {Ohl, Louis and Mattei, Pierre-Alexandre and Bouveyron, Charles and Leclercq, Mickaël and Droit, Arnaud and Precioso, Frédéric}, absract = {Feature selection in clustering is a hard task which involves simultaneously the discovery of relevant clusters as well as relevant variables with respect to these clusters. While feature selection algorithms are often model-based through optimised model selection or strong assumptions on the data distribution, we introduce a discriminative clustering model trying to maximise a geometry-aware generalisation of the mutual information called GEMINI with a simple L1 penalty: the Sparse GEMINI. This algorithm avoids the burden of combinatorial feature subset exploration and is easily scalable to high-dimensional data and large amounts of samples while only designing a discriminative clustering model. We demonstrate the performances of Sparse GEMINI on synthetic datasets and large-scale datasets. Our results show that Sparse GEMINI is a competitive algorithm and has the ability to select relevant subsets of variables with respect to the clustering without using relevance criteria or prior hypotheses.}, year = {2024}, pages = {155}, } - PreprintKernel KMeans clustering splits for end-to-end unsupervised decision trees2024
Trees are convenient models for obtaining explainable predictions on relatively small datasets. Although there are many proposals for the end-to-end construction of such trees in supervised learning, learning a tree end-to-end for clustering without labels remains an open challenge. As most works focus on interpreting with trees the result of another clustering algorithm, we present here a novel end-to-end trained unsupervised binary tree for clustering: Kauri. This method performs a greedy maximisation of the kernel KMeans objective without requiring the definition of centroids. We compare this model on multiple datasets with recent unsupervised trees and show that Kauri performs identically when using a linear kernel. For other kernels, Kauri often outperforms the concatenation of kernel KMeans and a CART decision tree.
- JACC Adv.AI-Enhanced Prediction of Aortic Stenosis ProgressionMelissa Sanabria, Lionel Tastet, Simon Pelletier, and 8 more authorsJACC: Advances, Oct 2024
*Background*. Aortic valve stenosis (AS) is a progressive chronic disease with progression rates that vary in patients and therefore difficult to predict. *Objectives*. The aim of this study was to predict the progression of AS using comprehensive and longitudinal patient data. *Methods*. Machine and deep learning algorithms were trained on a data set of 303 patients enrolled in the PROGRESSA (Metabolic Determinants of the Progression of Aortic Stenosis) study who underwent clinical and echocardiographic follow-up on an annual basis. Performance of the models was measured to predict disease progression over long (next 5 years) and short (next 2 years) terms and was compared to a standard clinical model with usually used features in clinical settings based on logistic regression. *Results*. For each annual follow-up visit including baseline, we trained various supervised learning algorithms in predicting disease progression at 2- and 5-year terms. At both terms, LightGBM consistently outperformed other models with the highest average area under curves across patient visits (0.85 at 2 years, 0.83 at 5 years). Recurrent neural network-based models (Gated Recurrent Unit and Long Short-Term Memory) and XGBoost also demonstrated strong predictive capabilities, while the clinical model showed the lowest performance. *Conclusions*. This study demonstrates how an artificial intelligence-guided approach in clinical routine could help enhance risk stratification of AS. It presents models based on multisource comprehensive data to predict disease progression and clinical outcomes in patients with mild-to-moderate AS at baseline.
@article{sanabria_ai-enhanced_2024, title = {{AI}-{Enhanced} {Prediction} of {Aortic} {Stenosis} {Progression}}, volume = {3}, url = {https://www.jacc.org/doi/10.1016/j.jacadv.2024.101234}, doi = {10.1016/j.jacadv.2024.101234}, number = {10}, urldate = {2024-12-05}, journal = {JACC: Advances}, publisher = {American College of Cardiology Foundation}, author = {Sanabria, Melissa and Tastet, Lionel and Pelletier, Simon and Leclercq, Mickael and Ohl, Louis and Hermann, Lara and Mattei, Pierre-Alexandre and Precioso, Frederic and Coté, Nancy and Pibarot, Philippe and Droit, Arnaud}, month = oct, year = {2024}, keywords = {aortic stenosis, machine learning, deep learning, risk prediction}, pages = {101234}, }
2022
- Generalised Mutual Information for Discriminative ClusteringIn Advances in Neural Information Processing Systems, 2022
In the last decade, recent successes in deep clustering majorly involved the mutual information (MI) as an unsupervised objective for training neural networks with increasing regularisations. While the quality of the regularisations have been largely discussed for improvements, little attention has been dedicated to the relevance of MI as a clustering objective. In this paper, we first highlight how the maximisation of MI does not lead to satisfying clusters. We identified the Kullback-Leibler divergence as the main reason of this behaviour. Hence, we generalise the mutual information by changing its core distance, introducing the generalised mutual information (GEMINI): a set of metrics for unsupervised neural network training. Unlike MI, some GEMINIs do not require regularisations when training. Some of these metrics are geometry-aware thanks to distances or kernels in the data space. Finally, we highlight that GEMINIs can automatically select a relevant number of clusters, a property that has been little studied in deep clustering context where the number of clusters is a priori unknown.
@inproceedings{ohl_generalised_2022, title = {Generalised {Mutual} {Information} for {Discriminative} {Clustering}}, volume = {35}, booktitle = {Advances in {Neural} {Information} {Processing} {Systems}}, publisher = {Curran Associates, Inc.}, author = {Ohl, Louis and Mattei, Pierre-Alexandre and Bouveyron, Charles and Harchaoui, Warith and Leclercq, Mickaël and Droit, Arnaud and Precioso, Frederic}, editor = {Koyejo, S. and Mohamed, S. and Agarwal, A. and Belgrave, D. and Cho, K. and Oh, A.}, year = {2022}, pages = {3377--3390}, }