1. FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
- Author
-
Ferreira, Alef Iury Siqueira, Gris, Lucas Rafael, da Rosa, Augusto Seben, de Oliveira, Frederico Santos, Casanova, Edresson, Sousa, Rafael Teixeira, Junior, Arnaldo Candido, Soares, Anderson da Silva, and Filho, Arlindo Galvão
- Subjects
Computer Science - Sound ,Electrical Engineering and Systems Science - Audio and Speech Processing - Abstract
This work presents FreeSVC, a promising multilingual singing voice conversion approach that leverages an enhanced VITS model with Speaker-invariant Clustering (SPIN) for better content representation and the State-of-the-Art (SOTA) speaker encoder ECAPA2. FreeSVC incorporates trainable language embeddings to handle multiple languages and employs an advanced speaker encoder to disentangle speaker characteristics from linguistic content. Designed for zero-shot learning, FreeSVC enables cross-lingual singing voice conversion without extensive language-specific training. We demonstrate that a multilingual content extractor is crucial for optimal cross-language conversion. Our source code and models are publicly available.
- Published
- 2025