IPA-CLIP: Integrating Phonetic Priors into Vision and Language Pretraining
Matsuhira, Chihaya, Kastner, Marc A., Komamizu, Takahiro, Hirayama, Takatsugu, Doman, Keisuke, Kawanishi, Yasutomo, Ide, Ichiro
–arXiv.org Artificial Intelligence
Recently, large-scale Vision and Language (V\&L) pretraining has become the standard backbone of many multimedia systems. While it has shown remarkable performance even in unseen situations, it often performs in ways not intuitive to humans. Particularly, they usually do not consider the pronunciation of the input, which humans would utilize to understand language, especially when it comes to unknown words. Thus, this paper inserts phonetic prior into Contrastive Language-Image Pretraining (CLIP), one of the V\&L pretrained models, to make it consider the pronunciation similarity among its pronunciation inputs. To achieve this, we first propose a phoneme embedding that utilizes the phoneme relationships provided by the International Phonetic Alphabet (IPA) chart as a phonetic prior. Next, by distilling the frozen CLIP text encoder, we train a pronunciation encoder employing the IPA-based embedding. The proposed model named IPA-CLIP comprises this pronunciation encoder and the original CLIP encoders (image and text). Quantitative evaluation reveals that the phoneme distribution on the embedding space represents phonetic relationships more accurately when using the proposed phoneme embedding. Furthermore, in some multimodal retrieval tasks, we confirm that the proposed pronunciation encoder enhances the performance of the text encoder and that the pronunciation encoder handles nonsense words in a more phonetic manner than the text encoder. Finally, qualitative evaluation verifies the correlation between the pronunciation encoder and human perception regarding pronunciation similarity.
arXiv.org Artificial Intelligence
Mar-6-2023
- Country:
- North America > United States
- New York > New York County
- New York City (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- Florida > Miami-Dade County
- Miami (0.04)
- California > San Diego County
- San Diego (0.04)
- New York > New York County
- Europe
- Spain > Catalonia (0.04)
- Germany > Berlin (0.04)
- Czechia > Prague (0.04)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.14)
- Italy > Tuscany
- Florence (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.04)
- Asia
- North America > United States
- Genre:
- Research Report (0.50)
- Technology: