msu
Representatividad Muestral en la Incertidumbre Sim\'etrica Multivariada para la Selecci\'on de Atributos
Author: Gustavo Daniel Sosa Cabrera Advisors: Miguel García Torres Santiago Gómez Christian E. Schaerer Serra SUMMARY In this work, we analyze the behavior of the multivariate symmetric uncertainty (MSU) measure through the use of statistical simulation techniques under various mixes of informative and non-informative randomly generated features. Experiments show how the number of attributes, their cardinalities, and the sample size affect the MSU. In this thesis, through observation of results, it is proposed an heuristic condition that preserves good quality in the MSU under different combinations of these three factors, providing a new useful criterion to help drive the process of dimension reduction. Definición 5. La incertidumbre simétrica de dos variables aleatorias X, Y se define como Hierarchical clustering based on mutual information.
Convergent Bounds on the Euclidean Distance
Given a set V of n vectors in d-dimensional space, we provide an efficient method for computing quality upper and lower bounds of the Euclidean distances between a pair of vectors in V. For this purpose, we define a distance measure, called the MS-distance, by using the mean and the standard deviation values of vectors in V. Once we compute the mean and the standard deviation values of vectors in V in O(dn) time, the MS-distance provides upper and lower bounds of Euclidean distance between any pair of vectors in V in constant time. Furthermore, these bounds can be refined further in such a way to converge monotonically to the exact Euclidean distance within d refinement steps. An analysis on a random sequence of refinement steps shows that the MS-distance provides very tight bounds in only a few refinement steps. The MS-distance can be used to various applications where the Euclidean distance is used to measure the proximity or similarity between objects. We provide experimental results on the nearest and the farthest neighbor searches.
Understanding a Version of Multivariate Symmetric Uncertainty to assist in Feature Selection
Sosa-Cabrera, Gustavo, García-Torres, Miguel, Gómez, Santiago, Schaerer, Christian, Divina, Federico
In these spaces of high dimensionality, feature selection is a way to exclude those irrelevant and redundant features, whose presence might complicate the task of knowledge discovery. In classification tasks, a feature is considered irrelevant if it contains no information about the class and therefore it is not necessary at all for the predictive task. Besides, it is widely accepted that two features are redundant if their values are correlated. There are several well known measures that compare features and determine their importance, such as the symmetrical uncertainty (SU)[2]. SU is a measure based on information that uses entropy and conditional entropy values to determine the correlation between pairs of features.
General Information, PRIP Lab at MSU
The Pattern Recognition and Image Processing (PRIP) Lab faculty and students investigate the use of machines to recognize a variety of patterns or objects. Methods are developed to sense objects, to discover which of their features distinguish them from others, and to design algorithms which can be used by a machine to do classification or clustering. Many practical applications use a sensed image to initially represent the object, and so much of the PRIP Lab research deals with images. A significant portion of our research focuses on the development of algorithms to do feature extraction and matching and on the organization of data to support efficient matching. Important applications include face recognition, fingerprint identification, document image analysis, 3D object recognition, robot navigation, and visualization/exploration of 3D volumetric data.
Convergent Bounds on the Euclidean Distance
Given a set V of n vectors in d-dimensional space, we provide an efficient method for computing quality upper and lower bounds of the Euclidean distances between a pair of the vectors in V . For this purpose, we define a distance measure, called the MS-distance, by using the mean and the standard deviation values of vectors in V . Once we compute the mean and the standard deviation values of vectors in V in O(dn) time, the MS-distance between them provides upper and lower bounds of Euclidean distance between a pair of vectors in V in constant time. Furthermore, these bounds can be refined further such that they converge monotonically to the exact Euclidean distance within d refinement steps. We also provide an analysis on a random sequence of refinement steps which can justify why MS-distance should be refined to provide very tight bounds in a few steps of a typical sequence. The MS-distance can be used to various problems where the Euclidean distance is used to measure the proximity or similarity between objects. We provide experimental results on the nearest and the farthest neighbor searches.