試験
通常
試験時間: 00:00問題時間: 00:00
問題一覧
問題 19 / 65
Which speech synthesis architecture is recommended when you want to generate natural prosody and high-quality speech but keep the model's latency low?
①Classical concatenative TTS (cut-and-paste from a corpus)
②Completely rule-based parametric speech synthesis
③A simple model that just generates mel spectrograms from text.
④Combination of a sequence-to-sequence acoustic model and a lightweight neural vocoder (e.g., Tacotron2 + WaveRNN)
⑤How to apply diffusion models directly to speech waveform generation
