A Comparison of Discrete Latent Variable Models for Speech Representation Learning

Zhou, Henry, Baevski, Alexei, Auli, Michael

Oct-23-2020–arXiv.org Artificial Intelligence

Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly based on predicting future time-steps or auto-encoding the input signal. Our study compares the representations learned by vq-vae and vq-wav2vec in terms of sub-word unit discovery and phoneme recognition performance. Results show that future time-step prediction with vq-wav2vec achieves better performance. The best system achieves an error rate of 13.22 on the ZeroSpeech 2019 ABX phoneme discrimination challenge.

artificial intelligence, representation, speech recognition, (14 more...)

arXiv.org Artificial Intelligence

Oct-23-2020

arXiv.org PDF

Add feedback

Country:
- North America > Canada > Ontario > Toronto (0.14)

Genre:
- Research Report > New Finding (0.49)

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning
    - Learning Graphical Models (0.85)
    - Statistical Learning (0.99)
  - Speech > Speech Recognition (0.69)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found