A model of early word acquisition based on realistic-scale audiovisual naming events
Khorrami, Khazar, Räsänen, Okko
–arXiv.org Artificial Intelligence
As they grow, infants gradually acquire understanding of their native language without direct supervision. By the age of 6 months, infants' perception has already attuned to native language phonetic contrasts [1] and they show first signs of word comprehension [2-4] and familiar word identification [5-7]. By 12 months, they already recognize dozens of words [8]. During this learning process, the infants must learn to parse the speech stream into words and to associate the words with their referential meanings in the external world. From a cognitive perspective, the discovery of words and word-meaning mappings is a task far from trivial: acoustic speech is a continuous and complex signal without transparent linguistic structure (see, e.g., [9, 10]), and there is substantial ambiguity in how individual words embedded in larger utterances are related to specific objects and events in the visual scene, also known as referential ambiguity [11]. Figure 1 illustrates referential ambiguity in everyday life situations between speech and its visual references through some examples.
arXiv.org Artificial Intelligence
Jun-7-2024
- Country:
- North America > United States
- Europe
- United Kingdom > England
- Oxfordshire > Oxford (0.04)
- Netherlands > Gelderland
- Nijmegen (0.04)
- Finland > Pirkanmaa
- Tampere (0.04)
- United Kingdom > England
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Education (1.00)
- Leisure & Entertainment > Sports (0.93)
- Transportation (0.68)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Cognitive Science (1.00)
- Representation & Reasoning (0.94)
- Natural Language > Text Processing (0.68)
- Machine Learning > Neural Networks
- Deep Learning (0.67)
- Information Technology > Artificial Intelligence