Overview
On Causality Inference in Time Series
Bahadori, Mohammad Taha (University of Southern Califoria) | Liu, Yan (University of Southern California)
Causality discovery has been one of the core tasks in scientific research since the beginning of human scientific history. In the age of data tsunami, the causality discovery task involves identification of causality among millions of variables which cannot be done manually by humans. However, the identification of causality relationships using artificial intelligence and statistical techniques in non-experimental settings faces several challenges. In this work, we address three of the challenges regarding Granger causality, one of the most popular causality inference techniques. First, we analyze the consistency of two most popular Granger causality techniques and show that the significance test is not consistent in high dimensions. Second, we review our nonparametric generalization of the Lasso-Granger technique called Generalized Lasso Granger (GLG) to uncover Granger causality relationships among irregularly sampled time series. Finally, we describe two techniques to uncover the casual dependence in non-linear datasets. Extensive experiments are provided to show the significant advantages of the proposed algorithms over their state-of-the-art counterparts.
Discovery Informatics: AI Opportunities in Scientific Discovery
Gil, Yolanda (University of Southern California) | Hirsh, Haym (Rutgers University)
Artificial Intelligence researchers have long sought to understand and replicate processes of scientific discovery. This article discusses Discovery Informatics as an emerging area of research that builds on that tradition and applies principles of intelligent computing and information systems to understand, automate, improve, and innovate processes of scientific discovery.
Transforming Graph Data for Statistical Relational Learning
Rossi, R. A., McDowell, L. K., Aha, D. W., Neville, J.
Relational data representations have become an increasingly important topic due to the recent proliferation of network datasets (e.g., social, biological, information networks) and a corresponding increase in the application of Statistical Relational Learning (SRL) algorithms to these domains. In this article, we examine and categorize techniques for transforming graph-based relational data to improve SRL algorithms. In particular, appropriate transformations of the nodes, links, and/or features of the data can dramatically affect the capabilities and results of SRL algorithms. We introduce an intuitive taxonomy for data representation transformations in relational domains that incorporates link transformation and node transformation as symmetric representation tasks. More specifically, the transformation tasks for both nodes and links include (i) predicting their existence, (ii) predicting their label or type, (iii) estimating their weight or importance, and (iv) systematically constructing their relevant features. We motivate our taxonomy through detailed examples and use it to survey competing approaches for each of these tasks. We also discuss general conditions for transforming links, nodes, and features. Finally, we highlight challenges that remain to be addressed.
A Tutorial on Dual Decomposition and Lagrangian Relaxation for Inference in Natural Language Processing
Dual decomposition, and more generally Lagrangian relaxation, is a classical method for combinatorial optimization; it has recently been applied to several inference problems in natural language processing (NLP). This tutorial gives an overview of the technique. We describe example algorithms, describe formal guarantees for the method, and describe practical issues in implementing the algorithms. While our examples are predominantly drawn from the NLP literature, the material should be of general relevance to inference problems in machine learning. A central theme of this tutorial is that Lagrangian relaxation is naturally applied in conjunction with a broad class of combinatorial algorithms, allowing inference in models that go significantly beyond previous work on Lagrangian relaxation for inference in graphical models.
Latent Composite Likelihood Learning for the Structured Canonical Correlation Model
Latent variable models are used to estimate variables of interest quantities which are observable only up to some measurement error. In many studies, such variables are known but not precisely quantifiable (such as "job satisfaction" in social sciences and marketing, "analytical ability" in educational testing, or "inflation" in economics). This leads to the development of measurement instruments to record noisy indirect evidence for such unobserved variables such as surveys, tests and price indexes. In such problems, there are postulated latent variables and a given measurement model. At the same time, other unantecipated latent variables can add further unmeasured confounding to the observed variables. The problem is how to deal with unantecipated latents variables. In this paper, we provide a method loosely inspired by canonical correlation that makes use of background information concerning the "known" latent variables. Given a partially specified structure, it provides a structure learning approach to detect "unknown unknowns," the confounding effect of potentially infinitely many other latent variables. This is done without explicitly modeling such extra latent factors. Because of the special structure of the problem, we are able to exploit a new variation of composite likelihood fitting to efficiently learn this structure. Validation is provided with experiments in synthetic data and the analysis of a large survey done with a sample of over 100,000 staff members of the National Health Service of the United Kingdom.
Distributed Problem Solving
Yeoh, William (Singapore Management University) | Yokoo, Makoto (Kyushu University)
Distributed problem solving is a subfield within multiagent systems, where agents are assumed to be part of a team and collaborate with each other to reach a common goal. In this article, we illustrate the motivations for distributed problem solving and provide an overview of two distributed problem solving models, namely distributed constraint satisfaction problems (DCSPs) and distributed constraint optimization problems (DCOPs), and some of their algorithms.
An Overview of Recent Application Trends at the AAMAS Conference: Security, Sustainability and Safety
Jain, Manish (University of Southern California) | An, Bo (University of Southern California) | Tambe, Milind (University of Southern California)
A key feature of the AAMAS conference is its emphasis on ties to real-world applications. The focus of this article is to provide a broad overview of application-focused papers published at the AAMAS 2010 and 2011 conferences. More specifically, recent applications at AAMAS could be broadly categorized as belonging to research areas of security, sustainability and safety. We outline the domains of applications, key research thrusts underlying each such application area, and emerging trends.
Agent-Based Modeling and Simulation
Klügl, Franziska (Orebro University) | Bazzan, Ana L. C. (Universidade Federal do Rio Grande do Sul)
This article gives an introduction to agent-based modeling and simulation (ABMS). After a general discussion about modeling and simulation, we address the basic concept of ABMS, focusing on its generative and bottom-up nature, its advantages as well as its pitfalls. The subsequent part of the article deals with application-oriented aspects, including selected tools and well-known applications. In order to illustrate the benefits of using ABMS, we focus on several aspects of a well-known area related to simulation of complex systems, namely traffic. At the end, a brief look into future challenges is given.