Abstract
Extracting the main melody from a polyphonic music recording seems natural even to untrained human listeners. To a certain extent it is related to the concept of source separation, with the human ability of focusing on a specific source in order to extract relevant information. In this paper, we propose a new approach for the estimation and extraction of the main melody (and in particular the leading vocal part) from polyphonic audio signals. To that aim, we propose a new signal model where the leading vocal part is explicitly represented by a specific source/filter model. The proposed representation is investigated in the framework of two statistical models: a Gaussian Scaled Mixture Model (GSMM) and an extended Instantaneous Mixture Model (IMM). For both models, the estimation of the different parameters is done within a maximum-likelihood framework adapted from single-channel source separation techniques. The desired sequence of fundamental frequencies is then inferred from the estimated parameters. The results obtained in a recent evaluation campaign (MIREX08) show that the proposed approaches are very promising and reach state-of-the-art performances on all test sets.
| Original language | English |
|---|---|
| Article number | 5410055 |
| Pages (from-to) | 564-575 |
| Number of pages | 12 |
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 18 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - 1 Mar 2010 |
| Externally published | Yes |
Keywords
- Blind audio source separation
- Expectation-;Maximization (EM) algorithm
- Gaussian scaled mixture model (GSMM)
- Main melody extraction
- Maximum likelihood
- Music
- Non-negative matrix factorization (NMF)
- Source/filter model
- Spectral analysis
Fingerprint
Dive into the research topics of 'Source/filter model for unsupervised main melody extraction from polyphonic audio signals'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver