Item request has been placed!

Item request cannot be made.

Processing Request

Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT

Item request has been placed!

Item request cannot be made.

Processing Request

اقرأ أكثر حفظ في قائمتي

المؤلفون: Yamauchi, Kazuki; Saito, Yuki; Saruwatari, Hiroshi
الموضوع:
Computer Science - Sound; Computer Science - Computation and Language; Electrical Engineering and Systems Science - Audio and Speech Processing
نوع التسجيلة:
Working Paper
الدخول الالكتروني :
http://arxiv.org/abs/2409.07265

معلومة اضافية
- الموضوع:
  2024
- Collection:
  Computer Science
- نبذة مختصرة :
  We explore cross-dialect text-to-speech (CD-TTS), a task to synthesize learned speakers' voices in non-native dialects, especially in pitch-accent languages. CD-TTS is important for developing voice agents that naturally communicate with people across regions. We present a novel TTS model comprising three sub-modules to perform competitively at this task. We first train a backbone TTS model to synthesize dialect speech from a text conditioned on phoneme-level accent latent variables (ALVs) extracted from speech by a reference encoder. Then, we train an ALV predictor to predict ALVs tailored to a target dialect from input text leveraging our novel multi-dialect phoneme-level BERT. We conduct multi-dialect TTS experiments and evaluate the effectiveness of our model by comparing it with a baseline derived from conventional dialect TTS methods. The results show that our model improves the dialectal naturalness of synthetic speech in CD-TTS.
  Comment: Accepted by IEEE SLT 2024
- الرقم المعرف:
  edsarx.2409.07265

تعليقات

No Comments.