Clustering, Coding, and the Concept of Similarity
–arXiv.org Artificial Intelligence
This paper develops a theory of clustering and coding which combines a geometric model with a probabilistic model in a principled way. The geometric model is a Riemannian manifold with a Riemannian metric, ${g}_{ij}({\bf x})$, which we interpret as a measure of dissimilarity. The probabilistic model consists of a stochastic process with an invariant probability measure which matches the density of the sample input data. The link between the two models is a potential function, $U({\bf x})$, and its gradient, $\nabla U({\bf x})$. We use the gradient to define the dissimilarity metric, which guarantees that our measure of dissimilarity will depend on the probability measure. Finally, we use the dissimilarity metric to define a coordinate system on the embedded Riemannian manifold, which gives us a low-dimensional encoding of our original data.
arXiv.org Artificial Intelligence
May-16-2018
- Country:
- North America
- United States
- New York (0.04)
- Massachusetts > Suffolk County
- Boston (0.04)
- Canada > Ontario
- Toronto (0.14)
- United States
- Europe
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Russia > Northwestern Federal District
- Leningrad Oblast > Saint Petersburg (0.04)
- Netherlands > North Holland
- Amsterdam (0.04)
- United Kingdom > England
- Asia
- Russia (0.04)
- Middle East > Jordan (0.04)
- India > West Bengal
- Kolkata (0.04)
- North America
- Genre:
- Research Report (0.81)
- Technology: