Crowdsourcing the Pronunciation of Out-of-Vocabulary Words
Shirali-Shahreza, Sajad (University of Toronto) | Luitjens, Pieter (University of Toronto) | Morcos, Natalie (University of Toronto) | Xiao, Wen (University of Toronto) | Qian, Zhenghong (University of Toronto) | Penn, Gerald (University of Toronto)
This is an Out-of-vocabulary (OOV) words still account for a significant extremely conservative use of crowdsourcing, particularly number of the mistakes by both speech recognizers as their crowdsource workers really do speak the words in and text-to-speech synthesizers. These are not words that their experiments, rather than selecting the correct pronunciation are merely very rare, but words that were unknown to the in a multiple choice question format. Our approach lexicon used by the automatic speech recognizer (ASR) or uses nothing more than a larger number of speakers (101) text-to-speech synthesizer (TTS). In the case of ASR, even and an acoustic model in order to find the pronunciation if the pronunciation is accurately modelled, there can be a almost ab nihilo, by constructing phone lattices and submitting question as to how to spell it correctly. In the case of TTS candidate pronunciation paths to a simple weighted systems, the pronunciation of the word may be unknown, as voting algorithm that combines results across crowdsource the component euphemistically known as "letter-to-sound" workers. Our only assumption is that the basic phonetic inventory or "grapheme-to-phoneme" rules may in fact not be able to is known to the acoustic model (e.g., the pronunciation infer the pronunciation from the spelling of the word, particularly of Rodriguez selected using an English acoustic model if its provenance is unknown, or the writing system is would never trill the r's). Furthermore, whereas Rutherford more logographically constructed.
Feb-4-2017