TY - GEN
T1 - Wavetransfer
T2 - 34th IEEE International Workshop on Machine Learning for Signal Processing, MLSP 2024
AU - Baoueb, Teysir
AU - Bie, Xiaoyu
AU - Janati, Hicham
AU - Richard, Gaël
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024/1/1
Y1 - 2024/1/1
N2 - As diffusion-based deep generative models gain prevalence, researchers are actively investigating their potential applications across various domains, including music synthesis and style alteration. Within this work, we are interested in timbre transfer, a process that involves seamlessly altering the instrumental characteristics of musical pieces while preserving essential musical elements. This paper introduces WaveTransfer, an end-to-end diffusion model designed for timbre transfer. We specifically employ the bilateral denoising diffusion model (BDDM) for noise scheduling search. Our model is capable of conducting timbre transfer between audio mixtures as well as individual instruments. Notably, it exhibits versatility in that it accommodates multiple types of timbre transfer between unique instrument pairs in a single model, eliminating the need for separate model training for each pairing. Furthermore, unlike recent works limited to 16 kHz, WaveTransfer can be trained at various sampling rates, including the industry-standard 44.1 kHz, a feature of particular interest to the music community.
AB - As diffusion-based deep generative models gain prevalence, researchers are actively investigating their potential applications across various domains, including music synthesis and style alteration. Within this work, we are interested in timbre transfer, a process that involves seamlessly altering the instrumental characteristics of musical pieces while preserving essential musical elements. This paper introduces WaveTransfer, an end-to-end diffusion model designed for timbre transfer. We specifically employ the bilateral denoising diffusion model (BDDM) for noise scheduling search. Our model is capable of conducting timbre transfer between audio mixtures as well as individual instruments. Notably, it exhibits versatility in that it accommodates multiple types of timbre transfer between unique instrument pairs in a single model, eliminating the need for separate model training for each pairing. Furthermore, unlike recent works limited to 16 kHz, WaveTransfer can be trained at various sampling rates, including the industry-standard 44.1 kHz, a feature of particular interest to the music community.
KW - Multi-instrumental timbre transfer
KW - diffusion models
KW - generative AI
KW - music transformation
U2 - 10.1109/MLSP58920.2024.10734786
DO - 10.1109/MLSP58920.2024.10734786
M3 - Conference contribution
AN - SCOPUS:85210576883
T3 - IEEE International Workshop on Machine Learning for Signal Processing, MLSP
BT - 34th IEEE International Workshop on Machine Learning for Signal Processing, MLSP 2024 - Proceedings
PB - IEEE Computer Society
Y2 - 22 September 2024 through 25 September 2024
ER -