Ontologies
Shared Model of Sense-making for Human-Machine Collaboration
Tecuci, Gheorghe, Marcu, Dorin, Kaiser, Louis, Boicu, Mihai
We present a model of sense-making that greatly facilitates the collaboration between an intelligent analyst and a knowledge-based agent. It is a general model grounded in the science of evidence and the scientific method of hypothesis generation and testing, where sense-making hypotheses that explain an observation are generated, relevant evidence is then discovered, and the hypotheses are tested based on the discovered evidence. We illustrate how the model enables an analyst to directly instruct the agent to understand situations involving the possible production of weapons (e.g., chemical warfare agents) and how the agent becomes increasingly more competent in understanding other situations from that domain (e.g., possible production of centrifuge-enriched uranium or of stealth fighter aircraft).
Extraction of common conceptual components from multiple ontologies
Asprino, Luigi, Carriero, Valentina Anita, Presutti, Valentina
Understanding large ontologies - by humans or machines - is both a struggle and crucially important for performing ontology engineering tasks such as ontology reuse, ontology matching, ontology evaluation, and (federated) querying [2]. According to [6], existing visualisation tools fail in providing overviews of large ontologies, which is crucial for ontology understanding, while none of them allows to compare multiple ontologies. Besides the layout and interaction features, the problem lays in the lack of effective methods for producing summaries of large ontologies. Many summarisation approaches focus on analysing the data level, e.g. to reduce the size of a knowledge graph and allow simplified queries for testing its coverage [16, 3]. Available summarisation methods addressing the conceptual level are based on extractive approaches that select and return a subset of nodes from the original ontology, i.e. the key concepts, as a summary [16]. However, an overall understanding of all the facts an ontology can represent, and a comparison between multiple ontologies, are not supported. For example, we may identify that in a cultural heritage ontology the concepts Cultural Property and Collection are key ones, however this is insufficient to understand if one ontology allows to answer whether a cultural property has been in a collection. Two ontologies having the same key concept would appear they address the same modelling problem, which may not be the case.
Marriage is a Peach and a Chalice: Modelling Cultural Symbolism on the SemanticWeb
Sartini, Bruno, van Erp, Marieke, Gangemi, Aldo
In this work, we fill the gap in the Semantic Web in the context of Cultural Symbolism. Building upon earlier work in, we introduce the Simulation Ontology, an ontology that models the background knowledge of symbolic meanings, developed by combining the concepts taken from the authoritative theory of Simulacra and Simulations of Jean Baudrillard with symbolic structures and content taken from "Symbolism: a Comprehensive Dictionary" by Steven Olderr. We re-engineered the symbolic knowledge already present in heterogeneous resources by converting it into our ontology schema to create HyperReal, the first knowledge graph completely dedicated to cultural symbolism. A first experiment run on the knowledge graph is presented to show the potential of quantitative research on symbolism.
Creating Knowledge Graphs Subsets using Shape Expressions
The initial adoption of knowledge graphs by Google and later by big companies has increased their adoption and popularity. In this paper we present a formal model for three different types of knowledge graphs which we call RDF-based graphs, property graphs and wikibase graphs. In order to increase the quality of Knowledge Graphs, several approaches have appeared to describe and validate their contents. Shape Expressions (ShEx) has been proposed as concise language for RDF validation. We give a brief introduction to ShEx and present two extensions that can also be used to describe and validate property graphs (PShEx) and wikibase graphs (WShEx). One problem of knowledge graphs is the large amount of data they contain, which jeopardizes their practical application. In order to palliate this problem, one approach is to create subsets of those knowledge graphs for some domains. We propose the following approaches to generate those subsets: Entity-matching, simple matching, ShEx matching, ShEx plus Slurp and ShEx plus Pregel which are based on declaratively defining the subsets by either matching some content or by Shape Expressions. The last approach is based on a novel validation algorithm for ShEx based on the Pregel algorithm that can handle big data graphs and has been implemented on Apache Spark GraphX.
Fuzzy Conceptual Graphs: a comparative discussion
Faci, Adam, Lesot, Marie-Jeanne, Laudy, Claire
Conceptual Graphs (CG) are a graph-based knowledge representation and reasoning formalism; fuzzy Conceptual Graphs (fCG) constitute an extension that enriches their expressiveness, exploiting the fuzzy set theory so as to relax their constraints at various levels. This paper proposes a comparative study of existing approaches over their respective advantages and possible limitations. The discussion revolves around three axes: (a) Critical view of each approach and comparison with previous propositions from the state of the art; (b) Presentation of the many possible interpretations of each definition to illustrate its potential and its limits; (c) Clarification of the part of CG impacted by the definition as well as the relaxed constraint.
Development of Semantic Web-based Imaging Database for Biological Morphome
Kume, Satoshi, Masuya, Hiroshi, Maeda, Mitsuyo, Suga, Mitsuo, Kataoka, Yosky, Kobayashi, Norio
We introduce the RIKEN Microstructural Imaging Metadatabase, a semantic web-based imaging database in which image metadata are described using the Resource Description Framework (RDF) and detailed biological properties observed in the images can be represented as Linked Open Data. The metadata are used to develop a large-scale imaging viewer that provides a straightforward graphical user interface to visualise a large microstructural tiling image at the gigabyte level. We applied the database to accumulate comprehensive microstructural imaging data produced by automated scanning electron microscopy. As a result, we have successfully managed vast numbers of images and their metadata, including the interpretation of morphological phenotypes occurring in sub-cellular components and biosamples captured in the images. We also discuss advanced utilisation of morphological imaging data that can be promoted by this database.
Principled Representation Learning for Entity Alignment
Guo, Lingbing, Sun, Zequn, Chen, Mingyang, Hu, Wei, Zhang, Qiang, Chen, Huajun
Embedding-based entity alignment (EEA) has recently received great attention. Despite significant performance improvement, few efforts have been paid to facilitate understanding of EEA methods. Most existing studies rest on the assumption that a small number of pre-aligned entities can serve as anchors connecting the embedding spaces of two KGs. Nevertheless, no one investigates the rationality of such an assumption. To fill the research gap, we define a typical paradigm abstracted from existing EEA methods and analyze how the embedding discrepancy between two potentially aligned entities is implicitly bounded by a predefined margin in the scoring function. Further, we find that such a bound cannot guarantee to be tight enough for alignment learning. We mitigate this problem by proposing a new approach, named NeoEA, to explicitly learn KG-invariant and principled entity embeddings. In this sense, an EEA model not only pursues the closeness of aligned entities based on geometric distance, but also aligns the neural ontologies of two KGs by eliminating the discrepancy in embedding distribution and underlying ontology knowledge. Our experiments demonstrate consistent and significant improvement in performance against the best-performing EEA methods.
Why Settle for Just One? Extending EL++ Ontology Embeddings with Many-to-Many Relationships
Mohapatra, Biswesh, Bhatia, Sumit, Mutharaju, Raghava, Srinivasaraghavan, G.
Knowledge Graph (KG) embeddings provide a low-dimensional representation of entities and relations of a Knowledge Graph and are used successfully for various applications such as question answering and search, reasoning, inference, and missing link prediction. However, most of the existing KG embeddings only consider the network structure of the graph and ignore the semantics and the characteristics of the underlying ontology that provides crucial information about relationships between entities in the KG. Recent efforts in this direction involve learning embeddings for a Description Logic (logical underpinning for ontologies) named EL++. However, such methods consider all the relations defined in the ontology to be one-to-one which severely limits their performance and applications. We provide a simple and effective solution to overcome this shortcoming that allows such methods to consider many-to-many relationships while learning embedding representations. Experiments conducted using three different EL++ ontologies show substantial performance improvement over five baselines. Our proposed solution also paves the way for learning embedding representations for even more expressive description logics such as SROIQ.
NLP Methods for Extraction of Symptoms from Unstructured Data for Use in Prognostic COVID-19 Analytic Models
Silverman, Greg M. | Sahoo, Himanshu S. (NLP/IE Program, Department of Electrical and Computer Engineering, University of Minnesota) | Ingraham, Nicholas E. (Division of Pulmonary, Allergy, Critical Care, and Sleep Medicine, University of Minnesota) | Lupei, Monica (Division of Critical Care, Department of Anesthesiology, University of Minnesota) | Puskarich, Michael A. (Department of Emergency Medicine, University of Minnesota) | Usher, Michael (Department of Medicine, University of Minnesota) | Dries, James (University of Minnesota) | Finzel, Raymond L. (NLP/IE Program, College of Pharmacy, University of Minnesota) | Murray, Eric (Information Technology, M Health Fairview) | Sartori, John (Department of Electrical and Computer Engineering, University of Minnesota) | Simon, Gyorgy (Institute for Health Informatics, University of Minnesota ) | Zhang, Rui | Melton, Genevieve B. (NLP/IE Program, Department of Surgery, and Institute for Health Informatics, University of Minnesota, Fairview Health Services, Information Technology) | Tignanelli, Christopher J. (NLP/IE Program, Department of Surgery, University of Minnesota ) | Pakhomov, Serguei VS (NLP/IE Program, College of Pharmacy, University of Minnesota )
Statistical modeling of outcomes based on a patient's presenting symptoms (symptomatology) can help deliver high quality care and allocate essential resources, which is especially important during the COVID-19 pandemic. Patient symptoms are typically found in unstructured notes, and thus not readily available for clinical decision making. In an attempt to fill this gap, this study compared two methods for symptom extraction from Emergency Department (ED) admission notes. Both methods utilized a lexicon derived by expanding The Center for Disease Control and Prevention's (CDC) Symptoms of Coronavirus list. The first method utilized a word2vec model to expand the lexicon using a dictionary mapping to the Uni ed Medical Language System (UMLS). The second method utilized the expanded lexicon as a rule-based gazetteer and the UMLS. These methods were evaluated against a manually annotated reference (f1-score of 0.87 for UMLS-based ensemble; and 0.85 for rule-based gazetteer with UMLS). Through analyses of associations of extracted symptoms used as features against various outcomes, salient risks among the population of COVID-19 patients, including increased risk of in-hospital mortality (OR 1.85, p-value < 0.001), were identified for patients presenting with dyspnea. Disparities between English and non-English speaking patients were also identified, the most salient being a concerning finding of opposing risk signals between fatigue and in-hospital mortality (non-English: OR 1.95, p-value = 0.02; English: OR 0.63, p-value = 0.01). While use of symptomatology for modeling of outcomes is not unique, unlike previous studies this study showed that models built using symptoms with the outcome of in-hospital mortality were not significantly different from models using data collected during an in-patient encounter (AUC of 0.9 with 95% CI of [0.88, 0.91] using only vital signs; AUC of 0.87 with 95% CI of [0.85, 0.88] using only symptoms). These findings indicate that prognostic models based on symptomatology could aid in extending COVID-19 patient care through telemedicine, replacing the need for in-person options. The methods presented in this study have potential for use in development of symptomatology-based models for other diseases, including for the study of Post-Acute Sequelae of COVID-19 (PASC).
Semi-automated checking for regulatory compliance in e-Health
Amantea, Ilaria Angela, Robaldo, Livio, Sulis, Emilio, Boella, Guido, Governatori, Guido
One of the main issues of every business process is to be compliant with legal rules. This work presents a methodology to check in a semi-automated way the regulatory compliance of a business process. We analyse an e-Health hospital service in particular: the Hospital at Home (HaH) service. The paper shows, at first, the analysis of the hospital business using the Business Process Management and Notation (BPMN) standard language, then, the formalization in Defeasible Deontic Logic (DDL) of some rules of the European General Data Protection Regulation (GDPR). The aim is to show how to combine a set of tasks of a business with a set of rules to be compliant with, using a tool.