Abstract
Real-world sounds often exhibit time-varying spectral shapes, as observed in the spectrogram of a harpsichord tone or that of a transition between two pronounced vowels. Whereas the standard non-negative matrix factorization (NMF) assumes fixed spectral atoms, an extension is proposed where the temporal activations (coefficients of the decomposition on the spectral atom basis) become frequency dependent and follow a time-varying autoregressive moving average (ARMA) modeling. This extension can thus be interpreted with the help of a source/filter paradigm and is referred to as source/filter factorization. This factorization leads to an efficient single-atom decomposition for a single audio event with strong spectral variation (but with constant pitch). The new algorithm is tested on real audio data and shows promising results.
| Original language | English |
|---|---|
| Article number | 5535132 |
| Pages (from-to) | 744-753 |
| Number of pages | 10 |
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 19 |
| Issue number | 4 |
| DOIs | |
| Publication status | Published - 21 Feb 2011 |
| Externally published | Yes |
Keywords
- Music information retrieval (MIR)
- non-negative matrix factorization (NMF)
- unsupervised machine learning
Fingerprint
Dive into the research topics of 'NMF with time-frequency activations to model nonstationary audio events'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver