Expert Systems
Inferring Missing Entity Type Instances for Knowledge Base Completion: New Dataset and Methods
Neelakantan, Arvind, Chang, Ming-Wei
Most of previous work in knowledge base (KB) completion has focused on the problem of relation extraction. In this work, we focus on the task of inferring missing entity type instances in a KB, a fundamental task for KB competition yet receives little attention. Due to the novelty of this task, we construct a large-scale dataset and design an automatic evaluation methodology. Our knowledge base completion method uses information within the existing KB and external information from Wikipedia. We show that individual methods trained with a global objective that considers unobserved cells from both the entity and the type side gives consistently higher quality predictions compared to baseline methods. We also perform manual evaluation on a small subset of the data to verify the effectiveness of our knowledge base completion methods and the correctness of our proposed automatic evaluation method.
RoboBrain: Large-Scale Knowledge Engine for Robots
Saxena, Ashutosh, Jain, Ashesh, Sener, Ozan, Jami, Aditya, Misra, Dipendra K., Koppula, Hema S.
In this paper we introduce a knowledge engine, which learns and shares knowledge representations, for robots to carry out a variety of tasks. Building such an engine brings with it the challenge of dealing with multiple data modalities including symbols, natural language, haptic senses, robot trajectories, visual features and many others. The \textit{knowledge} stored in the engine comes from multiple sources including physical interactions that robots have while performing tasks (perception, planning and control), knowledge bases from the Internet and learned representations from several robotics research groups. We discuss various technical aspects and associated challenges such as modeling the correctness of knowledge, inferring latent information and formulating different robotic tasks as queries to the knowledge engine. We describe the system architecture and how it supports different mechanisms for users and robots to interact with the engine. Finally, we demonstrate its use in three important research areas: grounding natural language, perception, and planning, which are the key building blocks for many robotic tasks. This knowledge engine is a collaborative effort and we call it RoboBrain.
Knowledge reduction of dynamic covering decision information systems with varying attribute values
Knowledge reduction of dynamic covering information systems involves with the time in practical situations. In this paper, we provide incremental approaches to computing the type-1 and type-2 characteristic matrices of dynamic coverings because of varying attribute values. Then we present incremental algorithms of constructing the second and sixth approximations of sets by using characteristic matrices. We employ experimental results to illustrate that the incremental approaches are effective to calculate approximations of sets in dynamic covering information systems. Finally, we perform knowledge reduction of dynamic covering information systems with the incremental approaches.
Using Semantics and Statistics to Turn Data into Knowledge
Pujara, Jay (University of Maryland, College Park) | Miao, Hui (University of California, Santa Cruz) | Getoor, Lise (Carnegie Mellon University) | Cohen, William W.
Many information extraction and knowledge base construction systems are addressing the challenge of deriving knowledge from text. In this article, we represent the desired knowledge base as a knowledge graph and introduce the problem of knowledge graph identification, collectively resolving the entities, labels, and relations present in the knowledge graph. Knowledge graph identification requires reasoning jointly over millions of extractions simultaneously, posing a scalability challenge to many approaches. We use probabilistic soft logic (PSL), a recently-introduced statistical relational learning framework, to implement an efficient solution to knowledge graph identification and present state-of-the-art results for knowledge graph construction while performing an order of magnitude faster than competing methods.
Using Semantics and Statistics to Turn Data into Knowledge
Pujara, Jay (University of Maryland, College Park) | Miao, Hui (University of California, Santa Cruz) | Getoor, Lise (Carnegie Mellon University) | Cohen, William W.
Many information extraction and knowledge base construction systems are addressing the challenge of deriving knowledge from text. A key problem in constructing these knowledge bases from sources like the web is overcoming the erroneous and incomplete information found in millions of candidate extractions. To solve this problem, we turn to semantics — using ontological constraints between candidate facts to eliminate errors. In this article, we represent the desired knowledge base as a knowledge graph and introduce the problem of knowledge graph identification, collectively resolving the entities, labels, and relations present in the knowledge graph. Knowledge graph identification requires reasoning jointly over millions of extractions simultaneously, posing a scalability challenge to many approaches. We use probabilistic soft logic (PSL), a recently-introduced statistical relational learning framework, to implement an efficient solution to knowledge graph identification and present state-of-the-art results for knowledge graph construction while performing an order of magnitude faster than competing methods.
Towards Learning a Knowledge Base of Actions from Experiential Microblogs
Kiciman, Emre (Microsoft Research)
While today's structured knowledge bases (e.g., Freebase) contain a sizable collection of information about entities, from celebrities and locations to concepts and common objects, there is a class of knowledge that has minimal coverage: actions. A large-scale knowledge base of actions would provide an opportunity for computinng devices to aid and support people's reasoning about their own actions and outcomes, leading to improved decision-making and goal achievement. In this short paper, we describe our first efforts towards building a distributional representation of actions and their outcomes, as learned from the timelines of individuals posting experiential microblogs.
Compositional Vector Space Models for Knowledge Base Inference
Neelakantan, Arvind (University of Massachusetts, Amherst) | Roth, Benjamin (University of Massachusetts, Amherst) | McCallum, Andrew (University of Massachusetts, Amherst)
Traditional approaches to knowledge base completion have been based on symbolic representations. Low-dimensional vector embedding models proposed recently for this task are attractive since they generalize to possibly unlimited sets of relations. A significant draw- back of previous embedding models for KB completion is that they merely support reasoning on individual relations (e.g., bornIn ( X, Y ) ⇒ nationality ( X, Y ) ). In this work, we develop models for KB completion that support chains of reasoning on paths of any length using compositional vector space models. We construct compositional vector representations for the paths in the KB graph from the semantic vector representations of the binary relations in that path and perform inference directly in the vector space. Unlike previous methods, our approach can generalize to paths that are unseen in training and, in a zero-shot setting, predict target relations without supervised training data for that relation.
Process Diagnosis System (PDS) – A 30 Year History
Thompson, Edward D. (Siemens Energy, Inc.) | Frolich, Ethan (Siemens Energy, Inc.) | Bellows, James C. (Siemens Energy, Inc.) | Bassford, Benjamin E. (Siemens Energy, Inc.) | Skiko, Edward J. (Siemens Energy, Inc.) | Fox, Mark S. (University of Toronto)
PDS (Process Diagnosis System) is an expert system shell developed in the early 1980's. It could handle thousands of sensor inputs and produce thousands of diagnostic messages with confidence factors based on complex logic designed to mimic the thinking of human experts. PDS went into commercial operation in 1985 to monitor seven power plant generators from a centralized diagnostic center at Westinghouse Power Generation headquarters. In the 1990’s the popularity of advanced technology gas turbines provided a renaissance in PDS utilization. The software has undergone rewrites and improvements since its inception, and the current PCPDS now supports the Siemens Power Diagnostics® Center with centralized rule based monitoring of over 1200 gas turbines, steam turbines, and generators.
Exploring Social Context for Topic Identification in Short and Noisy Texts
Wang, Xin (Jilin University;Key Laboratory of Symbolic Computation and Knowledge Engineering, Ministry of Education) | Wang, Ying (Changchun Institute of Tech) | Zuo, Wanli (Jilin University) | Cai, Guoyong (Jilin University)
With the pervasion of social media, topic identification in short texts attracts increasing attention in recent years. However, in nature the texts of social media are short and noisy, and the structures are sparse and dynamic, resulting in difficulty to identify topic categories exactly from online social media. Inspired by social science findings that preference consistency and social contagion are observed in social media, we investigate topic identification in short and noisy texts by exploring social context from the perspective of social sciences. In particular, we present a mathematical optimization formulation that incorporates the preference consistency and social contagion theories into a supervised learning method, and conduct feature selection to tackle short and noisy texts in social media, which result in a Sociological framework for Topic Identification (STI). Experimental results on real-world datasets from Twitter and Citation Network demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the importance of social context in topic identification.
Incremental Update of Datalog Materialisation: the Backward/Forward Algorithm
Motik, Boris (University of Oxford) | Nenov, Yavor (University of Oxford) | Piro, Robert Edgar Felix (University of Oxford) | Horrocks, Ian (University of Oxford)
Datalog-based systems often materialise all consequences of a datalog program and the data, allowing users' queries to be evaluated directly in the materialisation. This process, however, can be computationally intensive, so most systems update the materialisation incrementally when input data changes. We argue that existing solutions, such as the well-known Delete/Rederive (DRed) algorithm, can be inefficient in cases when facts have many alternate derivations. As a possible remedy, we propose a novel Backward/Forward (B/F) algorithm that tries to reduce the amount of work by a combination of backward and forward chaining. In our evaluation, the B/F algorithm was several orders of magnitude more efficient than the DRed algorithm on some inputs, and it was never significantly less efficient.