Spoken Word2Vec: A Perspective And Some Techniques
Sayeed, Mohammad Amaan, Aldarmaki, Hanan
–arXiv.org Artificial Intelligence
Text word embeddings that encode distributional semantic features work by modeling contextual similarities of frequently occurring words. Acoustic word embeddings, on the other hand, typically encode low-level phonetic similarities. Semantic embeddings for spoken words have been previously explored using similar algorithms to Word2Vec, but the resulting vectors still mainly encoded phonetic rather than semantic features. In this paper, we examine the assumptions and architectures used in previous works and show experimentally how Word2Vec algorithms fail to encode distributional semantics when the input units are acoustically correlated. In addition, previous works relied on the simplifying assumptions of perfect word segmentation and clustering by word type. Given these conditions, a trivial solution identical to text-based embeddings has been overlooked. We follow this simpler path using automatic word type clustering and examine the effects on the resulting embeddings, highlighting the true challenges in this task.
arXiv.org Artificial Intelligence
Nov-15-2023
- Country:
- North America > United States
- Louisiana > Orleans Parish > New Orleans (0.04)
- Europe
- Italy > Trentino-Alto Adige/Südtirol
- Trentino Province > Trento (0.04)
- Czechia > South Moravian Region
- Brno (0.04)
- Italy > Trentino-Alto Adige/Südtirol
- Asia > Middle East
- UAE > Abu Dhabi Emirate > Abu Dhabi (0.14)
- North America > United States
- Genre:
- Research Report (1.00)
- Technology: