week 73
r/MachineLearning - [D] Machine Learning - WAYR (What Are You Reading) - Week 73
It is inspired by the Location-Sensitive attention applied in Tacotron. However, that kind of mechanism is not well suited to generalize over long utterances. It means, we can synthesize texts up to the longest sentence in the data set which poses a major limitation. For example if we synthesize very long texts, the synthesized audio after some point consists only of repetitions and gibberish voice. With the newly invented Dynamic Convolution Attention (DCA) this is no longer a case. We can train with a data sets with relatively short utterances and still synthesize very long texts, without any significant loss.