HomeLearnCoursesHackathonsAccount
Speech Recognition & Synthesis
Text-to-Speech: From Concatenative to Neural · 1/2

The old approach: stitching together recordings

Text-to-speech (TTS) is the reverse problem of ASR: turning written text into a spoken audio waveform. Older concatenative TTS systems worked by recording a voice actor saying a huge number of short speech fragments in advance, then stitching together the appropriate fragments at run time to spell out whatever new sentence was needed. This produced genuinely intelligible speech, but it tended to sound robotic and choppy at the seams, the points where separately-recorded fragments were joined rarely blended smoothly, because pitch, pacing, and emphasis didn't naturally connect across pieces recorded at different times and in different contexts.

Concatenative systems were also inflexible. Because the output was built entirely from a fixed library of pre-recorded clips, changing the emotional tone, speaking rate, or emphasis of the output meant either not having the right fragment recorded at all, or accepting an unnatural result.