Revisiting Cosine Similarity via Normalized ICA-transformed Embeddings
Yamagiwa, Hiroaki, Oyama, Momose, Shimodaira, Hidetoshi
–arXiv.org Artificial Intelligence
Cosine similarity is widely used to measure the similarity between two embeddings, while interpretations based on angle and correlation coefficient are common. In this study, we focus on the interpretable axes of embeddings transformed by Independent Component Analysis (ICA), and propose a novel interpretation of cosine similarity as the sum of semantic similarities over axes. To investigate this, we first show experimentally that unnormalized embeddings contain norm-derived artifacts. We then demonstrate that normalized ICA-transformed embeddings exhibit sparsity, with a few large values in each axis and across embeddings, thereby enhancing interpretability by delineating clear semantic contributions. Finally, to validate our interpretation, we perform retrieval experiments using ideal embeddings with and without specific semantic components.
arXiv.org Artificial Intelligence
Jun-16-2024
- Country:
- Africa > Ethiopia
- Addis Ababa > Addis Ababa (0.04)
- Asia
- China
- Anhui Province > Hefei (0.04)
- Beijing > Beijing (0.04)
- India > Maharashtra
- Mumbai (0.04)
- Japan > Honshū
- Kansai > Kyoto Prefecture > Kyoto (0.04)
- Middle East > Qatar
- Singapore (0.04)
- China
- Europe
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Czechia > Prague (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- France (0.04)
- Portugal > Lisbon
- Lisbon (0.04)
- Slovakia > Bratislava
- Bratislava (0.04)
- Slovenia > Central Slovenia
- Municipality of Ljubljana > Ljubljana (0.04)
- United Kingdom > Wales (0.04)
- Belgium > Brussels-Capital Region
- North America
- Canada
- British Columbia > Metro Vancouver Regional District
- Vancouver (0.04)
- Manitoba (0.04)
- Ontario (0.04)
- Quebec (0.04)
- British Columbia > Metro Vancouver Regional District
- United States
- Arizona (0.04)
- California
- Los Angeles County > Long Beach (0.04)
- Santa Clara County > Palo Alto (0.04)
- Colorado (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Nevada (0.04)
- New York > New York County
- New York City (0.04)
- Canada
- Oceania > Australia
- New South Wales > Sydney (0.04)
- Queensland (0.04)
- Africa > Ethiopia
- Genre:
- Research Report > New Finding (0.48)
- Industry:
- Health & Medicine > Therapeutic Area (0.46)
- Materials > Chemicals (0.46)
- Technology: