Passer à la navigation principale Passer à la recherche Passer au contenu principal

A musically motivated mid-level representation for pitch estimation and musical audio source separation

  • ENAC-IIC-GEL
  • CNRS LTCI

Résultats de recherche: Contribution à un journalArticleRevue par des pairs

122 Citations (Scopus)

Résumé

When designing an audio processing system, the target tasks often influence the choice of a data representation or transformation. Low-level time-frequency representations such as the short-time Fourier transform (STFT) are popular, because they offer a meaningful insight on sound properties for a low computational cost. Conversely, when higher level semantics, such as pitch, timbre or phoneme, are sought after, representations usually tend to enhance their discriminative characteristics, at the expense of their invertibility. They become so-called mid-level representations. In this paper, a source/filter signal model which provides a mid-level representation is proposed. This representation makes the pitch content of the signal as well as some timbre information available, hence keeping as much information from the raw data as possible. This model is successfully used within a main melody extraction system and a lead instrument/accompaniment separation system. Both frameworks obtained top results at several international evaluation campaigns.

langue originaleAnglais
Numéro d'article5784290
Pages (de - à)1180-1191
Nombre de pages12
journalIEEE Journal on Selected Topics in Signal Processing
Volume5
Numéro de publication6
Les DOIs
étatPublié - 1 oct. 2011
Modification externeOui

Empreinte digitale

Examiner les sujets de recherche de « A musically motivated mid-level representation for pitch estimation and musical audio source separation ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation