Résumé
Extracting the main melody from a polyphonic music recording seems natural even to untrained human listeners. To a certain extent it is related to the concept of source separation, with the human ability of focusing on a specific source in order to extract relevant information. In this paper, we propose a new approach for the estimation and extraction of the main melody (and in particular the leading vocal part) from polyphonic audio signals. To that aim, we propose a new signal model where the leading vocal part is explicitly represented by a specific source/filter model. The proposed representation is investigated in the framework of two statistical models: a Gaussian Scaled Mixture Model (GSMM) and an extended Instantaneous Mixture Model (IMM). For both models, the estimation of the different parameters is done within a maximum-likelihood framework adapted from single-channel source separation techniques. The desired sequence of fundamental frequencies is then inferred from the estimated parameters. The results obtained in a recent evaluation campaign (MIREX08) show that the proposed approaches are very promising and reach state-of-the-art performances on all test sets.
| langue originale | Anglais |
|---|---|
| Numéro d'article | 5410055 |
| Pages (de - à) | 564-575 |
| Nombre de pages | 12 |
| journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 18 |
| Numéro de publication | 3 |
| Les DOIs | |
| état | Publié - 1 mars 2010 |
| Modification externe | Oui |
Empreinte digitale
Examiner les sujets de recherche de « Source/filter model for unsupervised main melody extraction from polyphonic audio signals ». Ensemble, ils forment une empreinte digitale unique.Contient cette citation
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver