Technology
Optimal Fuzzy Model Construction with Statistical Information using Genetic Algorithm
Hossain, Md. Amjad, Shill, Pintu Chandra, Sarker, Bishnu, Murase, Kazuyuki
Fuzzy rule based models have a capability to approximate any continuous function to any degree of accuracy on a compact domain. The majority of FLC design process relies on heuristic knowledge of experience operators. In order to make the design process automatic we present a genetic approach to learn fuzzy rules as well as membership function parameters. Moreover, several statistical information criteria such as the Akaike information criterion (AIC), the Bhansali-Downham information criterion (BDIC), and the Schwarz-Rissanen information criterion (SRIC) are used to construct optimal fuzzy models by reducing fuzzy rules. A genetic scheme is used to design Takagi-Sugeno-Kang (TSK) model for identification of the antecedent rule parameters and the identification of the consequent parameters. Computer simulations are presented confirming the performance of the constructed fuzzy logic controller.
A comparison of two suffix tree-based document clustering algorithms
Rafi, Muhammad, Maujood, M., Fazal, M. M., Ali, S. M.
Document clustering as an unsupervised approach extensively used to navigate, filter, summarize and manage large collection of document repositories like the World Wide Web (WWW). Recently, focuses in this domain shifted from traditional vector based document similarity for clustering to suffix tree based document similarity, as it offers more semantic representation of the text present in the document. In this paper, we compare and contrast two recently introduced approaches to document clustering based on suffix tree data model. The first is an Efficient Phrase based document clustering, which extracts phrases from documents to form compact document representation and uses a similarity measure based on common suffix tree to cluster the documents. The second approach is a frequent word/word meaning sequence based document clustering, it similarly extracts the common word sequence from the document and uses the common sequence/ common word meaning sequence to perform the compact representation, and finally, it uses document clustering approach to cluster the compact documents. These algorithms are using agglomerative hierarchical document clustering to perform the actual clustering step, the difference in these approaches are mainly based on extraction of phrases, model representation as a compact document, and the similarity measures used for clustering. This paper investigates the computational aspect of the two algorithms, and the quality of results they produced.
A Split-Merge MCMC Algorithm for the Hierarchical Dirichlet Process
The hierarchical Dirichlet process (HDP) has become an important Bayesian nonparametric model for grouped data, such as document collections. The HDP is used to construct a flexible mixed-membership model where the number of components is determined by the data. As for most Bayesian nonparametric models, exact posterior inference is intractable---practitioners use Markov chain Monte Carlo (MCMC) or variational inference. Inspired by the split-merge MCMC algorithm for the Dirichlet process (DP) mixture model, we describe a novel split-merge MCMC sampling algorithm for posterior inference in the HDP. We study its properties on both synthetic data and text corpora. We find that split-merge MCMC for the HDP can provide significant improvements over traditional Gibbs sampling, and we give some understanding of the data properties that give rise to larger improvements.
The Interaction of Entropy-Based Discretization and Sample Size: An Empirical Study
An empirical investigation of the interaction of sample size and discretization - in this case the entropy-based method CAIM (Class-Attribute Interdependence Maximization) - was undertaken to evaluate the impact and potential bias introduced into data mining performance metrics due to variation in sample size as it impacts the discretization process. Of particular interest was the effect of discretizing within cross-validation folds averse to outside discretization folds. Previous publications have suggested that discretizing externally can bias performance results; however, a thorough review of the literature found no empirical evidence to support such an assertion. This investigation involved construction of over 117,000 models on seven distinct datasets from the UCI (University of California-Irvine) Machine Learning Library and multiple modeling methods across a variety of configurations of sample size and discretization, with each unique "setup" being independently replicated ten times. The analysis revealed a significant optimistic bias as sample sizes decreased and discretization was employed. The study also revealed that there may be a relationship between the interaction that produces such bias and the numbers and types of predictor attributes, extending the "curse of dimensionality" concept from feature selection into the discretization realm. Directions for further exploration are laid out, as well some general guidelines about the proper application of discretization in light of these results.
Fusion de classifieurs pour la classification d'images sonar
In this paper, we present some high level information fusion approaches for numeric and symbolic data. We study the interest of such method particularly for classifier fusion. A comparative study is made in a context of sea bed characterization from sonar images. The classi- fication of kind of sediment is a difficult problem because of the data complexity. We compare high level information fusion and give the obtained performance.
Classification under Data Contamination with Application to Remote Sensing Image Mis-registration
Yan, Donghui, Gong, Peng, Chen, Aiyou, Zhong, Liheng
This work is motivated by the problem of image mis-registration in remote sensing and we are interested in determining the resulting loss in the accuracy of pattern classification. A statistical formulation is given where we propose to use data contamination to model and understand the phenomenon of image mis-registration. This model is widely applicable to many other types of errors as well, for example, measurement errors and gross errors etc. The impact of data contamination on classification is studied under a statistical learning theoretical framework. A closed-form asymptotic bound is established for the resulting loss in classification accuracy, which is less than $\epsilon/(1-\epsilon)$ for data contamination of an amount of $\epsilon$. Our bound is sharper than similar bounds in the domain adaptation literature and, unlike such bounds, it applies to classifiers with an infinite Vapnik-Chervonekis (VC) dimension. Extensive simulations have been conducted on both synthetic and real datasets under various types of data contamination, including label flipping, feature swapping and the replacement of feature values with data generated from a random source such as a Gaussian or Cauchy distribution. Our simulation results show that the bound we derive is fairly tight.
How People Talk with Robots: Designing Dialog to Reduce User Uncertainty
Fischer, Kerstin (University of Southern Denmark)
If human-robot interaction is mainly shaped by users' strategies to deal with their unfamiliar artificial com munication partner, as it is suggested here, robot dialog design should orient at reducing users' uncertainty about the affordances of the robot and the joint task. Two experiments are presented that investigate the impact of verbal robot utterances on users' behavior; results show that users react sensitively to subtle linguistic cues that may guide them into appropriate understandings of the robot. Furthermore, the role of user expectations and robot appearance are discussed in the light of the model presented.
Crowdsourcing Real World Human-Robot Dialog and Teamwork through Online Multiplayer Games
Chernova, Sonia (Worcester Polytechnic Institute) | DePalma, Nick (Massachusetts Institute of Technology) | Breazeal, Cynthia (Massachusetts Institute of Technology)
We present an innovative approach for large-scale data collection in human-robot interaction research through the use of online multi-player games. By casting a robotic task as a collaborative game, we gather thousands of examples of human-human interactions online, and then leverage this corpus of action and dialog data to create contextually relevant, social and task-oriented behaviors for human-robot interaction in the real world. We demonstrate our work in a collaborative search and retrieval task requiring dialog, action synchronization and action sequencing between the human and robot partners. A user study performed at the Boston Museum of Science shows that the autonomous robot exhibits many of the same patterns of behavior that were observed in the online dataset and survey results rate the robot similarly to human partners in several critical measures.
Introduction to the Special Issue on Dialog with Robots
Bohus, Dan (Microsoft Research) | Horvitz, Eric (Microsoft Research) | Kanda, Takayuki (ATR Intelligent Robotics and Communication Laboratories) | Mutlu, Bilge (University of Wisconsin - Madison) | Raux, Antoine (Honda Research Institute USA)
This special issue of AI Magazine on dialog with robots brings together a collection of articles on situated dialog. The contributing authors have been working in interrelated fields of human-robot interaction, dialog systems, virtual agents, and other related areas and address core concepts in spoken dialog with embodied robots or agents. Several of the contributors participated in the AAAI Fall Symposium on Dialog with Robots, held in November 2010, and several articles in this issue are extensions of work presented there. The articles in this collection address diverse aspects of dialog with robots, but are unified in addressing opportunities with spoken language interaction, physical embodiment, and enriched representations of context.