Any-to-Many and Any-to-Any Voice Conversion based on BNE-Seq2seqMoL+

نویسندگان

1 Faculty of Electrical Engineering, Shahrood University of Technology, Shahrood, Iran

2 Faculty of Electrical Engineering, Shahrood University of Technology, Shahrood, Iran

doi
10.5829/ije.2026.39.12c.04
چکیده

Voice conversion (VC) aims to transform a source speaker’s voice to that of a target speaker while preserving linguistic content. Although recent sequence-to-sequence and transformer-based approaches have achieved promising results, robust any-to-many voice conversion remains challenging due to pitch variability, speaker imbalance, and limited disentanglement of phonetic and prosodic features. In this paper, we propose BNE-Seq2SeqMoL+, an architecturally enhanced sequence-to-sequence framework that introduces explicit refinements to improve temporal modeling, pitch representation, and spectral quality. The proposed method incorporates a transformer-based bottleneck feature modeling strategy, dedicated pitch encoding, and an efficient spectral refinement module, enabling more stable and natural voice conversion across diverse speaker pairs. Extensive experiments on multiple public datasets demonstrate that the proposed approach consistently outperforms the baseline in both any-to-many and any-to-any scenarios, including conversions involving unseen speakers. Objective evaluations show improvements in spectral accuracy, pitch stability, and linguistic content preservation, while subjective listening tests reveal an approximately 10% relative improvement in Mean Opinion Score (MOS) compared to the baseline. These results confirm that the proposed architectural enhancements provide an effective and computationally efficient solution for high-quality voice conversion.