Asia
Machine olfaction using time scattering of sensor multiresolution graphs
Gugel, Leonid, Shkolnisky, Yoel, Dekel, Shai
In this paper we construct a learning architecture for high dimensional time series sampled by sensor arrangements. Using a redundant wavelet decomposition on a graph constructed over the sensor locations, our algorithm is able to construct discriminative features that exploit the mutual information between the sensors. The algorithm then applies scattering networks to the time series graphs to create the feature space. We demonstrate our method on a machine olfaction problem, where one needs to classify the gas type and the location where it originates from data sampled by an array of sensors. Our experimental results clearly demonstrate that our method outperforms classical machine learning techniques used in previous studies.
Infusing Human Factors into Algorithmic Crowdsourcing
Yu, Han (Nanyang Technological University) | Miao, Chunyan (Nanyang Technological University) | Shen, Zhiqi (Nanyang Technological University) | Lin, Jun (Nanyang Technological University) | Leung, Cyril (University of British Columbia) | Yang, Qiang (Hong Kong University of Science and Technology)
The emergence of crowdsourcing systems have provided a viable mechanism for incorporating humans into the computational loop at large scale and in real-time. This offers an unprecedent opportunity to study how artificial intelligence (AI) techniques and humans can collaborate to solve problems. An important challenge in crowdsourcing is how to make optimal use of human resources as people have different skills and their availability may be limited. In this paper, we provide the research community with a new dataset derived from an online game-based platform to address this challenge. Six crowdsourcing task allocation scenarios with different overall workload levels and worker population characteristics were presented to over 400 players to solve. With close to 3,000 game sessions and over 300,000 task allocation decisions from human and AI players, the dataset provides an efficient focal point for the research community to design solutions that can sustainably tap into the pool of human resources through crowdsourcing.
Automated Capture and Execution of Manufacturability Rules Using Inductive Logic Programming
Moitra, Abha (GE Global Research) | Palla, Ravi (GE Global Research) | Rangarajan, Arvind (GE Global Research)
Capturing domain knowledge can be a time-consuming process that typically requires the collaboration of a Subject Matter Expert and a modeling expert to encode the knowledge. In a number of domains and applications, this situation is further exacerbated by the fact that the Subject Matter Expert may find it difficult to articulate the domain knowledge as a procedure or rules, but instead may find it easier to classify instance data. To facilitate this type of knowledge elicitation from Subject Matter Experts, we have developed a system that automatically generates formal and executable rules from provided labeled instance data. We do this by leveraging the techniques of Inductive Logic Programming (ILP) to generate Horn clause based rules to separate out positive and negative instance data. We illustrate our approach on a Design For Manufacturability (DFM) platform where the goal is to design products that are easy to manufacture by providing early manufacturability feedback. Specifically we show how our approach can be used to generate feature recognition rules from positive and negative instance data supplied by Subject Matter Experts. Our platform is interactive, provides visual feedback and is iterative. The feature identification rules generated can be inspected, manually refined and vetted.
Wikipedia in the Tourism Industry: Forecasting Demand and Modeling Usage Behavior
Khadivi, Pejman (Virginia Polytechnic Institute and State University) | Ramakrishnan, Naren (Virginia Polytechnic Institute and State University)
Due to the economic and social impacts of tourism, both private and public sectors are interested in precisely forecasting the tourism demand volume in a timely manner. With recent advances in social networks, more people use online resources to plan their future trips. In this paper we explore the application of Wikipedia usage trends (WUTs) in tourism analysis. We propose a framework that deploys WUTs for forecasting the tourism demand of Hawaii. We also propose a data-driven approach, using WUTs, to estimate the behavior of tourists when they plan their trips.
Data-Augmented Software Diagnosis
Elmishali, Amir (Ben Gurion University of the Negev) | Stern, Roni (Ben Gurion University of the Negev) | Kalech, Meir (Ben Gurion University of the Negev)
Software fault prediction algorithms predict which software components is likely to contain faults using machine learning techniques. Software diagnosis algorithm identify the faulty software components that caused a failure using model-based or spectrum based approaches. We show how software fault prediction algorithms can be used to improve software diagnosis. The resulting data-augmented diagnosis algorithm overcomes key problems in software diagnosis algorithms: ranking diagnoses and distinguishing between diagnoses with high probability and low probability. We demonstrate the efficiency of the proposed approach empirically on three open sources domains, showing significant increase in accuracy of diagnosis and efficiency of troubleshooting. These encouraging results suggests broader use of data-driven methods to complement and improve existing model-based methods.
Document Type Classification in Online Digital Libraries
Caragea, Cornelia (University of North Texas) | Wu, Jian (Pennsylvania State University) | Gollapalli, Sujatha Das (Institute for Infocomm Research, A*STAR) | Giles, C. Lee (Pennsylvania State University)
Online digital libraries make it easier for researchers to search for scientific information. They have been proven as powerful resources in many data mining, machine learning and information retrieval applications that require high-quality data. The quality of the data highly depends on the accuracy of classifiers that identify the types of documents that are crawled from the Web, e.g., as research papers, slides, books, etc., for appropriate indexing. These classifiers in turn depend on the choice of the feature representation. We propose novel features that result in high-accuracy classifiers for document type classification. Experimental results on several datasets show that our classifiers outperform models that are employed in current systems.
Ontology Re-Engineering: A Case Study from the Automotive Industry
Rychtyckyj, Nestor (Ford Motor Company) | Raman, Venkatesh (Ford Motor Company) | Sankaranarayanan, Baskaran (Indian Institute of Technology Madras) | Kumar, P. Sreenivasa (Indian Institute of Technology Madras) | Khemani, Deepak (Indian Institute of Technology Madras)
For over twenty five years Ford has been utilizing an AI-based system to manage process planning for vehicle assembly at our assembly plants around the world. The scope of the AI system, known originally as the Direct Labor Management System and now as the Global Study Process Allocation System (GSPAS),has increased over the years to include additional functionality on Ergonomics and Powertrain Assembly (Engine and Transmission plants). The knowledge about Ford’s manufacturing processes is contained in an ontology originally developed using the KL-ONE representation language and methodology. To preserve the viability of the GSPAS ontology and to make it easily usable for other applications within Ford, we needed to re-engineer and convert the KL-ONE ontology into a semantic web OWL/RDF format. In this paper, we will discuss the process by which we re-engineered the existing GSPAS KL-ONE ontology and deployed semantic web technology in our application.
Deploying PAWS: Field Optimization of the Protection Assistant for Wildlife Security
Fang, Fei (University of Southern California) | Nguyen, Thanh H. (University of Southern California) | Pickles, Rob (Panthera) | Lam, Wai Y. (Panthera, Rimba) | Clements, Gopalasamy R. (Panthera, Rimba, Kenyir Research Institute, and Universiti Malaysia Terengganu) | An, Bo (Nanyang Technological University) | Singh, Amandeep (Columbia University) | Tambe, Milind (University of Southern California) | Lemieux, Andrew (The Netherlands Institute for the Study of Crime and Law Enforcement (NSCR))
Poaching is a serious threat to the conservation of key species and whole ecosystems. While conducting foot patrols is the most commonly used approach in many countries to prevent poaching, such patrols often do not make the best use of limited patrolling resources. To remedy this situation, prior work introduced a novel emerging application called PAWS (Protection Assistant for Wildlife Security); PAWS was proposed as a game-theoretic (``security games'') decision aid to optimize the use of patrolling resources. This paper reports on PAWS's significant evolution from a proposed decision aid to a regularly deployed application, reporting on the lessons from the first tests in Africa in Spring 2014, through its continued evolution since then, to current regular use in Southeast Asia and plans for future worldwide deployment. In this process, we have worked closely with two NGOs (Panthera and Rimba) and incorporated extensive feedback from professional patrolling teams. We outline key technical advances that lead to PAWS's regular deployment: (i) incorporating complex topographic features, e.g., ridgelines, in generating patrol routes; (ii) handling uncertainties in species distribution (game theoretic payoffs); (iii) ensuring scalability for patrolling large-scale conservation areas with fine-grained guidance; and (iv) handling complex patrol scheduling constraints.
Nonparametric Canonical Correlation Analysis
Michaeli, Tomer, Wang, Weiran, Livescu, Karen
Canonical correlation analysis (CCA) is a classical representation learning technique for finding correlated variables in multi-view data. Several nonlinear extensions of the original linear CCA have been proposed, including kernel and deep neural network methods. These approaches seek maximally correlated projections among families of functions, which the user specifies (by choosing a kernel or neural network structure), and are computationally demanding. Interestingly, the theory of nonlinear CCA, without functional restrictions, had been studied in the population setting by Lancaster already in the 1950s, but these results have not inspired practical algorithms. We revisit Lancaster's theory to devise a practical algorithm for nonparametric CCA (NCCA). Specifically, we show that the solution can be expressed in terms of the singular value decomposition of a certain operator associated with the joint density of the views. Thus, by estimating the population density from data, NCCA reduces to solving an eigenvalue system, superficially like kernel CCA but, importantly, without requiring the inversion of any kernel matrix. We also derive a partially linear CCA (PLCCA) variant in which one of the views undergoes a linear projection while the other is nonparametric. Using a kernel density estimate based on a small number of nearest neighbors, our NCCA and PLCCA algorithms are memory-efficient, often run much faster, and perform better than kernel CCA and comparable to deep CCA.
A Tractable Fully Bayesian Method for the Stochastic Block Model
Hayashi, Kohei, Konishi, Takuya, Kawamoto, Tatsuro
The stochastic block model (SBM) is a generative model revealing macroscopic structures in graphs. Bayesian methods are used for (i) cluster assignment inference and (ii) model selection for the number of clusters. In this paper, we study the behavior of Bayesian inference in the SBM in the large sample limit. Combining variational approximation and Laplace's method, a consistent criterion of the fully marginalized log-likelihood is established. Based on that, we derive a tractable algorithm that solves tasks (i) and (ii) concurrently, obviating the need for an outer loop to check all model candidates. Our empirical and theoretical results demonstrate that our method is scalable in computation, accurate in approximation, and concise in model selection.