Skip to main navigation Skip to search Skip to main content

Classification of sparse high-dimensional vectors

  • St. Petersburg State Electrotechnical University
  • Aix Marseille Université
  • ENSAE

Research output: Contribution to journalArticlepeer-review

28 Citations (Scopus)

Abstract

We study the problem of classification of d-dimensional vectors into two classes (one of which is 'pure noise') based on a training sample of size m. The main specific feature is that the dimension d can be very large. We suppose that the difference between the distribution of the population and that of the noise is only in a shift, which is a sparse vector. For Gaussian noise, fixed sample size m, and dimension d that tends to infinity, we obtain the sharp classification boundary, i.e. the necessary and sufficient conditions for the possibility of successful classification. We propose classifiers attaining this boundary. We also give extensions of the result to the case where the sample size m depends on d and satisfies the condition (log m)/log d - y ,0 ≤ y< 1, and to the case of non-Gaussian noise satisfying the Cramer condition. This journal is

Original languageEnglish
Pages (from-to)4427-4448
Number of pages22
JournalPhilosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences
Volume367
Issue number1906
DOIs
Publication statusPublished - 13 Nov 2009

Keywords

  • Bayes risk
  • Classification boundary
  • High-dimensional data
  • Optimal classifier
  • Sparse vectors

Fingerprint

Dive into the research topics of 'Classification of sparse high-dimensional vectors'. Together they form a unique fingerprint.

Cite this