Country
Scale Normalization
Lo, Henry Z., Amaral, Kevin, Ding, Wei
One of the difficulties of training deep neural networks is caused by improper scaling between layers. Scaling issues introduce exploding / gradient problems, and have typically been addressed by careful scale-preserving initialization. We investigate the value of preserving scale, or isometry, beyond the initial weights. We propose two methods of maintaing isometry, one exact and one stochastic. Preliminary experiments show that for both determinant and scale-normalization effectively speeds up learning. Results suggest that isometry is important in the beginning of learning, and maintaining it leads to faster learning.
Evaluating the effect of topic consideration in identifying communities of rating-based social networks
Reihanian, Ali, Minaei-Bidgoli, Behrouz, Yousefnezhad, Muhammad
-- Finding meaningful communities in social network has attracted the attentions of many researchers. The community structure of complex networks reveals both their organization and hidd en relations among their constituents. Most of the researches in the field of community detection mainly focus on the topological structure of the network without performing any content analysis. Nowadays, real world social networks are containing a vast r ange of information including shared objects, comments, following information, etc. In recent years, a number of researches have proposed approaches which consider both the contents that are interchanged in the networks and the topological structures of th e networks in order to find more meaningful communities. In this research, the effect of topic analysis in finding more meaningful communities in social networking sites in which the users express their feelings toward different object s (like movies) by the means of rating is demonstrated by performing extensive experiments. With the emergence of social networks, people have been attracted to them, and have been sharing valuable information by means of communicating with each other. For example, folksonomies are social tagging sites which their users collaboratively express th eir feelings and sentiments toward a special resource like a movie or music by means of descriptive keywords (tags) [1] or ratings. One of the most important issues considered when analyzing these kinds of network s is community detection.
A New Approach in Persian Handwritten Letters Recognition Using Error Correcting Output Coding
Kazemi, Maziar, Yousefnezhad, Muhammad, Nourian, Saber
Classification Ensemble, which uses the weighed polling of outputs, is the art of combining a set of basic classifiers for generating high-performance, robust and more stable results. This study aims to improve the results of identifying the Persian handwritten letters using Error Correcting Output Coding (ECOC) ensemble method. Furthermore, the feature selection is used to reduce the costs of errors in our proposed method. ECOC is a method for decomposing a multi-way classification problem into many binary classification tasks; and then combining the results of the subtasks into a hypothesized solution to the original problem. Firstly, the image features are extracted by Principal Components Analysis (PCA). After that, ECOC is used for identification the Persian handwritten letters which it uses Support Vector Machine (SVM) as the base classifier. The empirical results of applying this ensemble method using 10 real-world data sets of Persian handwritten letters indicate that this method has better results in identifying the Persian handwritten letters than other ensemble methods and also single classifications. Moreover, by testing a number of different features, this paper found that we can reduce the additional cost in feature selection stage by using this method.
The Mean Partition Theorem of Consensus Clustering
Clustering is a standard technique for exploratory data analysis that finds applications across different disciplines such as computer science, biology, marketing, and social science. The goal of clustering is to group a set of unlabeled data points into several clusters based on some notion of dissimilarity. Inspired by the success of classifier ensembles, consensus clustering has emerged as a research topic [8, 23]. Consensus clustering first generates several partitions of the same dataset. Then it combines the sample partitions to a single consensus partition. The assumption is that a consensus partition better fits to the hidden structure in the data than individual partitions. One standard approach of consensus clustering combines the sample partitions to a mean partition [3, 4, 5, 6, 9, 17, 20, 21, 22].
Neural network-based clustering using pairwise constraints
This paper presents a neural network-based end-to-end clustering framework. We design a novel strategy to utilize the contrastive criteria for pushing data-forming clusters directly from raw data, in addition to learning a feature embedding suitable for such clustering. The network is trained with weak labels, specifically partial pairwise relationships between data instances. The cluster assignments and their probabilities are then obtained at the output layer by feed-forwarding the data. The framework has the interesting characteristic that no cluster centers need to be explicitly specified, thus the resulting cluster distribution is purely data-driven and no distance metrics need to be predefined. The experiments show that the proposed approach beats the conventional two-stage method (feature embedding with k-means) by a significant margin. It also compares favorably to the performance of the standard cross entropy loss for classification. Robustness analysis also shows that the method is largely insensitive to the number of clusters. Specifically, we show that the number of dominant clusters is close to the true number of clusters even when a large k is used for clustering.
Mixtures of Sparse Autoregressive Networks
We consider high-dimensional distribution estimation through autoregressive networks. By combining the concepts of sparsity, mixtures and parameter sharing we obtain a simple model which is fast to train and which achieves state-of-the-art or better results on several standard benchmark datasets. Specifically, we use an L1-penalty to regularize the conditional distributions and introduce a procedure for automatic parameter sharing between mixture components. Moreover, we propose a simple distributed representation which permits exact likelihood evaluations since the latent variables are interleaved with the observable variables and can be easily integrated out. Our model achieves excellent generalization performance and scales well to extremely high dimensions.
Simple, Robust and Optimal Ranking from Pairwise Comparisons
Shah, Nihar B., Wainwright, Martin J.
We consider data in the form of pairwise comparisons of n items, with the goal of precisely identifying the top k items for some value of k < n, or alternatively, recovering a ranking of all the items. We analyze the Copeland counting algorithm that ranks the items in order of the number of pairwise comparisons won, and show it has three attractive features: (a) its computational efficiency leads to speed-ups of several orders of magnitude in computation time as compared to prior work; (b) it is robust in that theoretical guarantees impose no conditions on the underlying matrix of pairwise-comparison probabilities, in contrast to some prior work that applies only to the BTL parametric model; and (c) it is an optimal method up to constant factors, meaning that it achieves the information-theoretic limits for recovering the top k-subset. We extend our results to obtain sharp guarantees for approximate recovery under the Hamming distortion metric, and more generally, to any arbitrary error requirement that satisfies a simple and natural monotonicity condition.
Train and Test Tightness of LP Relaxations in Structured Prediction
Meshi, Ofer, Mahdavi, Mehrdad, Weller, Adrian, Sontag, David
Structured prediction is used in areas such as computer vision and natural language processing to predict structured outputs such as segmentations or parse trees. In these settings, prediction is performed by MAP inference or, equivalently, by solving an integer linear program. Because of the complex scoring functions required to obtain accurate predictions, both learning and inference typically require the use of approximate solvers. We propose a theoretical explanation to the striking observation that approximations based on linear programming (LP) relaxations are often tight on real-world instances. In particular, we show that learning with LP relaxed inference encourages integrality of training instances, and that tightness generalizes from train to test data.
Meet China's cool robo-monk
Buddhist monks in China have harnessed technology to create a robot monk. The 2-foot tall robot is called Xian'er and can chant Buddhist mantras, move via voice command and even hold a simple conversation, according to Reuters. Xian'er resembles a novice monk and "lives" at Longquan temple on the outskirts of Beijing. The cartoon-style robot holds a touchscreen to its chest and can answer about 20 questions on Buddhism and daily life. Master Xianfan, Xian'er's creator, described the robot as the perfect vessel for spreading the wisdom of Buddhism in China, via the fusion of science and Buddhism.
To Build the Best Robotic Exoskeleton, Make It on the Cheap
In a small startup space just down the street from UC Berkeley's campus, robotics pioneer Homayoon Kazerooni is bragging about how no-frills his invention is. "We're trying to make the Honda," he says, "not the sportscar." Kazerooni is showing me his latest robotic exoskeleton that gives paraplegics and people with mobility problems the ability to stand up from their wheelchairs and walk again. He's been building such bionic systems for more than a decade, and back in 2005 he cofounded Ekso Bionics, the current market leader for exoskeletons. So it's no surprise that Kazerooni says the new device from his new company, SuitX, is the most advanced yet.