Expert Systems
Finding New Information Via Robust Entity Detection
Iacobelli, Francisco (Northwestern University) | Nichols, Nathan (Northwestern University) | Birnbaum, Larry (Northwestern University) | Hammond, Kristian (Northwestern University)
Journalists and editors work under pressure to collect relevant details and background information about specific events. They spend a significant amount of time sifting through documents and finding new information such as facts, opinions or stakeholders (i.e. people, places and organizations that have a stake in the news). Spotting them is a tedious and cognitively intense process. One task, essential to this process, is to find and keep track of stakeholders. This task is taxing cognitively and in terms of memory. Tell Me More offers an automatic aid to this task. Tell Me More is a system that, given a seed story, mines the web for similar stories reported by different sources and selects only those stories which offer new information with respect to that original seed story. Much like a journalist, the task of detecting named entities is central to its success. In this paper we briefly describe Tell Me More and, in particular, we focus on Tell Me More's entity detection component. We describe an approach that combines off-the-shelf named entity recognizers (NERs) with WPED, an in-house publicly available NER that uses Wikipedia as its knowledge base. We show significant increase in precision scores with respect to traditional NERs. Lastly, we present an overall evaluation of Tell Me More using this approach.
Goal-Oriented Knowledge Collection
Kuo, Yen-Ling (National Taiwan University) | Hsu, Jane Yung-jen (National Taiwan University)
Games with A Purpose (GWAP) has been demonstrated to be efficient in collecting large amount of knowledge from online users, e.g. Verbosity and Virtual Pet game. However, its effectiveness in knowledge base (KB) construction has not been explored in previous research. This paper examines the knowledge collected in the Vir- tual Pet game and presents an approach to collect more knowledge driven by the existing relations in KB. In this paper, goal-oriented knowledge collection successfully draws 10572 answers for the "foodโ domain. The answers are verified by online voting to show that 92.07% of them are good sentences and 95.89% of them are new sentences. This result is a significant improvement over the original Virtual Pet game, with 80.58% good sentences and 67.56% weekly new information.
Significance of Classification Techniques in Prediction of Learning Disabilities
Balakrishnan, Julie M. David And Kannan
The aim of this study is to show the importance of two classification techniques, viz. decision tree and clustering, in prediction of learning disabilities (LD) of school-age children. LDs affect about 10 percent of all children enrolled in schools. The problems of children with specific learning disabilities have been a cause of concern to parents and teachers for some time. Decision trees and clustering are powerful and popular tools used for classification and prediction in Data mining. Different rules extracted from the decision tree are used for prediction of learning disabilities. Clustering is the assignment of a set of observations into subsets, called clusters, which are useful in finding the different signs and symptoms (attributes) present in the LD affected child. In this paper, J48 algorithm is used for constructing the decision tree and K-means algorithm is used for creating the clusters. By applying these classification techniques, LD in any child can be identified.
Learning under Concept Drift: an Overview
Concept drift refers to a non stationary learning problem over time. The training and the application data often mismatch in real life problems. In this report we present a context of concept drift problem 1. We focus on the issues relevant to adaptive training set formation. We present the framework and terminology, and formulate a global picture of concept drift learners design. We start with formalizing the framework for the concept drifting data in Section 1. In Section 2 we discuss the adaptivity mechanisms of the concept drift learners. In Section 3 we overview the principle mechanisms of concept drift learners. In this chapter we give a general picture of the available algorithms and categorize them based on their properties. Section 5 discusses the related research fields and Section 5 groups and presents major concept drift applications. This report is intended to give a bird's view of concept drift research field, provide a context of the research and position it within broad spectrum of research fields and applications.
Project Halo Update--Progress Toward Digital Aristotle
Gunning, David (Vulcan, Inc.) | Chaudhri, Vinay K. (SRI International) | Clark, Peter E. (Boeing Research and Technology) | Barker, Ken (University of Texas at Austin) | Chaw, Shaw-Yi (University of Texas at Austin) | Greaves, Mark (Vulcan, Inc.) | Grosof, Benjamin (Vulcan, Inc.) | Leung, Alice (Raytheon BBN Technologies Corporation) | McDonald, David D. (Raytheon BBN Technologies Corporation) | Mishra, Sunil (SRI International) | Pacheco, John (SRI International) | Porter, Bruce (University of Texas at Austin) | Spaulding, Aaron (SRI International) | Tecuci, Dan (University of Texas at Austin) | Tien, Jing (SRI International)
In the winter, 2004 issue of AI Magazine, we reported Vulcan Inc.'s first step toward creating a question-answering system called "Digital Aristotle." The goal of that first step was to assess the state of the art in applied Knowledge Representation and Reasoning (KRR) by asking AI experts to represent 70 pages from the advanced placement (AP) chemistry syllabus and to deliver knowledge-based systems capable of answering questions from that syllabus. This paper reports the next step toward realizing a Digital Aristotle: we present the design and evaluation results for a system called AURA, which enables domain experts in physics, chemistry, and biology to author a knowledge base and that then allows a different set of users to ask novel questions against that knowledge base. These results represent a substantial advance over what we reported in 2004, both in the breadth of covered subjects and in the provision of sophisticated technologies in knowledge representation and reasoning, natural language processing, and question answering to domain experts and novice users.
AI Theory and Practice: A Discussion on Hard Challenges and Opportunities Ahead
Horvitz, Eric (Microsoft Research) | Getoor, Lise (University of Maryland) | Guestrin, Carlos (Carnegie Mellon University) | Hendler, James (Rensselaer Polytechnic Institute) | Konstan, Joseph (University of Minnesota) | Subramanian, Devika (Rice University) | Wellman, Michael (University of Michigan) | Kautz, Henry (University of Rochester)
So, we have a variety of people here with different interests and backgrounds that I asked to talk about not just the key challenges ahead but potential opportunities and promising pathways, trajectories to solving those problems, and their predictions about how R&D might proceed in terms of the timing of various kinds of development over time. I asked the panelists briefly to frame their comments sharing a little bit about fundamental questions, such as, "What is the research goal?" Not everybody stays up late at night hunched over a computer or a simulation or a robotic system, pondering the foundations of intelligence and human-level AI. We have here today Lise Getoor from the University ipate the liability and insurance industry; and the of Maryland; Devika Subramanian, who other one, that it was a human interface problem, comes to us from Rice University; we have Carlos that people don't necessarily want to go and type Guestrin from Carnegie Mellon University (CMU); a bunch of yes/no questions into a computer to get James Hendler from Rensselaer Polytechnic Institute an answer, even with a rule-based explanation, (RPI); Mike Wellman at the University of that if you'd taken that just a step further and Michigan; Henry Kautz at tjhe University of solved the human problem, it might have worked. Rochester; and Joe Konstan, who comes to us from Related to that, I was remembering a bunch of the Midwest, as our Minneapolis person here on these smart house projects. And I have to admit I the panel. I think everyone Joe Konstan: I was actually surprised when you hates smart spaces. I think of myself at the core there's nobody there, do you warn people and give in human-computer interaction. So I went back them a chance to answer? There's no good answer and started looking at what I knew of artificial to this question. I can tell you if that person is in intelligence to try to see where the path forward bed asleep, the answer is no, don't wake them up was, and I was inspired by the past.
Project Halo UpdateโProgress Toward Digital Aristotle
Gunning, David (Vulcan, Inc.) | Chaudhri, Vinay K. (SRI International) | Clark, Peter E. (Boeing Research and Technology) | Barker, Ken (University of Texas at Austin) | Chaw, Shaw-Yi (University of Texas at Austin) | Greaves, Mark (Vulcan, Inc.) | Grosof, Benjamin (Vulcan, Inc.) | Leung, Alice (Raytheon BBN Technologies Corporation) | McDonald, David D. (Raytheon BBN Technologies Corporation) | Mishra, Sunil (SRI International) | Pacheco, John (SRI International) | Porter, Bruce (University of Texas at Austin) | Spaulding, Aaron (SRI International) | Tecuci, Dan (University of Texas at Austin) | Tien, Jing (SRI International)
In the winter, 2004 issue of AI Magazine, we reported Vulcan Inc.'s first step toward creating a question-answering system called "Digital Aristotle." The goal of that first step was to assess the state of the art in applied Knowledge Representation and Reasoning (KRR) by asking AI experts to represent 70 pages from the advanced placement (AP) chemistry syllabus and to deliver knowledge-based systems capable of answering questions from that syllabus. This paper reports the next step toward realizing a Digital Aristotle: we present the design and evaluation results for a system called AURA, which enables domain experts in physics, chemistry, and biology to author a knowledge base and that then allows a different set of users to ask novel questions against that knowledge base. These results represent a substantial advance over what we reported in 2004, both in the breadth of covered subjects and in the provision of sophisticated technologies in knowledge representation and reasoning, natural language processing, and question answering to domain experts and novice users.
True Knowledge: Open-Domain Question Answering Using Structured Knowledge and Inference
Tunstall-Pedoe, William (True Knowledge Ltd)
The motivation for the project was to tackle what might be regarded as the "holy grail" of Internet search, replacing larger and larger numbers of keyword-based lists of links with perfect, direct answers to naturally phrased queries on any subject. The platform was also designed to scale, with the primary mechanism for answering more and more questions being the addition of knowledge to the platform rather than writing more program code. Additional knowledge areas are typically included by adding "knowledge about knowledge." The system is live and answers millions of questions per month, asked by real Internet users. Questions can be tried at (and API access obtained from) www.trueknowledge.com. All the intellectual External computer systems can connect to the property was subsequently transferred in 2006 platform at two points through an API.
Efficient Knowledge Base Management in DCSP
DCSP (Distributed Constraint Satisfaction Problem) has been a very important research area in AI (Artificial Intelligence). There are many application problems in distributed AI that can be formalized as DSCPs. With the increasing complexity and problem size of the application problems in AI, the required storage place in searching and the average searching time are increasing too. Thus, to use a limited storage place efficiently in solving DCSP becomes a very important problem, and it can help to reduce searching time as well. This paper provides an efficient knowledge base management approach based on general usage of hyper-resolution-rule in consistence algorithm. The approach minimizes the increasing of the knowledge base by eliminate sufficient constraint and false nogood. These eliminations do not change the completeness of the original knowledge base increased. The proofs are given as well. The example shows that this approach decrease both the new nogoods generated and the knowledge base greatly. Thus it decreases the required storage place and simplify the searching process.
A Comprehensive Survey of Data Mining-based Fraud Detection Research
Phua, Clifton, Lee, Vincent, Smith, Kate, Gayler, Ross
This survey paper categorises, compares, and summarises from almost all published technical and review articles in automated fraud detection within the last 10 years. It defines the professional fraudster, formalises the main types and subtypes of known fraud, and presents the nature of data evidence collected within affected industries. Within the business context of mining the data to achieve higher cost savings, this research presents methods and techniques together with their problems. Compared to all related reviews on fraud detection, this survey covers much more technical articles and is the only one, to the best of our knowledge, which proposes alternative data and solutions from related domains.