Skip to main navigation Skip to search Skip to main content

False selection rate control in mixture models

  • Ariane Marandon
  • , Tabea Rebafka
  • , Etienne Roquain
  • , Nataliya Sokolovska
  • The Alan Turing Institute
  • Sorbonne Université

Research output: Contribution to journalArticlepeer-review

Abstract

The clustering task consists in partitioning elements of a sample into homogeneous groups. Most datasets contain individuals that are ambiguous and intrinsically difficult to attribute to one or another cluster. However, in practical applications, misclassifying individuals is potentially disastrous and should be avoided. To keep the misclassification rate small, one can decide to classify only a part of the sample. In the supervised setting, this approach is well known and referred to as classification with an abstention option. In this paper, the approach is revisited in an unsupervised mixture-model framework. The purpose is to develop a method that guarantees the false selection rate (FSR) does not exceed a predefined level (Formula presented.). We propose a plug-in procedure and provide a theoretical analysis, quantifying the deviation of the FSR from the target (Formula presented.) with explicit remainder terms. Bootstrap versions of the procedure are shown to improve the performance in numerical experiments.

Original languageEnglish
Pages (from-to)2014-2060
Number of pages47
JournalScandinavian Journal of Statistics
Volume52
Issue number4
DOIs
Publication statusPublished - 1 Dec 2025

Keywords

  • abstention option
  • bootstrap
  • clustering
  • false discovery rate
  • mixture models

Fingerprint

Dive into the research topics of 'False selection rate control in mixture models'. Together they form a unique fingerprint.

Cite this