A model of early word acquisition based on realistic-scale audiovisual naming events

Khorrami, Khazar, Räsänen, Okko

arXiv.org Artificial Intelligence 

As they grow, infants gradually acquire understanding of their native language without direct supervision. By the age of 6 months, infants' perception has already attuned to native language phonetic contrasts [1] and they show first signs of word comprehension [2-4] and familiar word identification [5-7]. By 12 months, they already recognize dozens of words [8]. During this learning process, the infants must learn to parse the speech stream into words and to associate the words with their referential meanings in the external world. From a cognitive perspective, the discovery of words and word-meaning mappings is a task far from trivial: acoustic speech is a continuous and complex signal without transparent linguistic structure (see, e.g., [9, 10]), and there is substantial ambiguity in how individual words embedded in larger utterances are related to specific objects and events in the visual scene, also known as referential ambiguity [11]. Figure 1 illustrates referential ambiguity in everyday life situations between speech and its visual references through some examples.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found