書誌事項
- 公開日
- 2019-05
- 権利情報
-
- https://ieeexplore.ieee.org/Xplorehelp/downloads/license-information/IEEE.html
- https://doi.org/10.15223/policy-029
- https://doi.org/10.15223/policy-037
- DOI
-
- 10.1109/icassp.2019.8682897
- 10.48550/arxiv.1904.04631
- 公開者
- IEEE
説明
Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training conditions. Recently, CycleGAN-VC has provided a breakthrough and performed comparably to a parallel VC method without relying on any extra data, modules, or time alignment procedures. However, there is still a large gap between the real target and converted speech, and bridging this gap remains a challenge. To reduce this gap, we propose CycleGAN-VC2, which is an improved version of CycleGAN-VC incorporating three new techniques: an improved objective (two-step adversarial losses), improved generator (2-1-2D CNN), and improved discriminator (PatchGAN). We evaluated our method on a non-parallel VC task and analyzed the effect of each technique in detail. An objective evaluation showed that these techniques help bring the converted feature sequence closer to the target in terms of both global and local structures, which we assess by using Mel-cepstral distortion and modulation spectra distance, respectively. A subjective evaluation showed that CycleGAN-VC2 outperforms CycleGAN-VC in terms of naturalness and similarity for every speaker pair, including intra-gender and inter-gender pairs.
Accepted to ICASSP 2019. Project page: http://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/cyclegan-vc2/index.html
収録刊行物
-
- ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
-
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 6820-6824, 2019-05
IEEE
- Tweet
キーワード
- FOS: Computer and information sciences
- Computer Science - Machine Learning
- Sound (cs.SD)
- Machine Learning (stat.ML)
- Computer Science - Sound
- Machine Learning (cs.LG)
- Statistics - Machine Learning
- Audio and Speech Processing (eess.AS)
- FOS: Electrical engineering, electronic engineering, information engineering
- Electrical Engineering and Systems Science - Audio and Speech Processing
詳細情報 詳細情報について
-
- CRID
- 1362262943594490624
-
- データソース種別
-
- Crossref
- OpenAIRE

