Goto

Collaborating Authors

 Ontologies


Deep Multiple Instance Learning for Taxonomic Classification of Metagenomic read sets

arXiv.org Machine Learning

Metagenomic studies have increasingly utilized sequencing technologies in order to analyze DNA fragments found in environmental samples. It can provide useful insights for studying the interactions between hosts and microbes, infectious disease proliferation, and novel species discovery. One important step in this analysis is the taxonomic classification of those DNA fragments. Of particular interest is the determination of the distribution of the taxa of microbes in metagenomic samples. Recent attempts using deep learning focus on architectures that classify single DNA reads independently from each other. In this work, we attempt to solve the task of directly predicting the distribution over the taxa of whole metagenomic read sets. We formulate this task as a Multiple Instance Learning (MIL) problem. We extend architectures used in single-read taxonomic classification with two different types of permutation-invariant MIL pooling layers: a) deepsets and b) attention-based pooling. We illustrate that our architecture can exploit the co-occurrence of species in metagenomic read sets and outperforms the single-read architectures in predicting the distribution over the taxa at higher taxonomic ranks.


Artificial Intelligence BlockCloud (AIBC) Technical Whitepaper

arXiv.org Machine Learning

The AIBC is an Artificial Intelligence and blockchain technology based large-scale decentralized ecosystem that allows system-wide low-cost sharing of computing and storage resources. The AIBC consists of four layers: a fundamental layer, a resource layer, an application layer, and an ecosystem layer. The AIBC implements a two-consensus scheme to enforce upper-layer economic policies and achieve fundamental layer performance and robustness: the DPoEV incentive consensus on the application and resource layers, and the DABFT distributed consensus on the fundamental layer. The DABFT uses deep learning techniques to predict and select the most suitable BFT algorithm in order to achieve the best balance of performance, robustness, and security. The DPoEV uses the knowledge map algorithm to accurately assess the economic value of digital assets.


Data Interpretation Support in Rescue Operations: Application for French Firefighters

arXiv.org Artificial Intelligence

--This work aims at developing a system that supports French firefighters in data interpretation during rescue operations. An application ontology is proposed based on existing crisis management ones and operational expertise collection. After that, a knowledge-based system will be developed and integrated in firefighters' environment. Our first studies are shown in this paper. Rescue of people consists in saving their life in case of distress situations by applying responsive operations. In France, it is defined as specific tasks to be accomplished by public services in order to ensure the safety of patients and victims by making them able to escape from dangers, securing intervention sites, providing medical help, and finally, ensuring the evacuation to an appropriate place of reception [1].


Parsa Mirhaji Montefiore Health System - PMWC Precision Medicine World Conference

#artificialintelligence

Dr. Mirhaji was the former director of the Center for Biosecurity and Public Health Informatics Research at the University of Texas at Houston where he developed clinical text understanding, semantic information integration, and EMR interoperability solutions, for public health and disaster preparedness. He is an inventor with several patents covering information integration, biomedical vocabularies and taxonomy services, clinical text understanding and natural language processing, electronic data capture, and knowledge-based information retrieval. Dr. Mirhaji and his fellow researchers were awarded "The Best Practice in Public Health. He is a member of W3C working groups for application of Semantic Technologies in Healthcare and Life Sciences, and organizer and committee member for several national and international conferences on Bio-Ontologies and Semantic Technologies.


Webinar summary - Semantic annotation of images in the FAIR data era CGIAR Platform for Big Data in Agriculture

#artificialintelligence

Digital agriculture increasingly relies on the generation of large quantity of images. These images are processed with machine learning techniques to speed up the identification of objects, their classification, visualization, and interpretation. However, images must comply with the FAIR principles to facilitate their access, reuse, and interoperability. As stated in recent paper authored by the Planteome team (Trigkakis et al, 2018), "Plant researchers could benefit greatly from a trained classification model that predicts image annotations with a high degree of accuracy." In this third Ontologies Community of Practice webinar, Justin Preece, Senior Faculty Research Assistant Oregon State University, presents the module developed by the Planteome project using the Bio-Image Semantic Query User Environment (BISQUE), an online image analysis and storage platform of Cyverse.


The future of Pharma: harnessing AI to decentralise data

#artificialintelligence

As Chief Data Officer for the OSTHUS Group, Eric Little co-founded LeapAnalysis, a new approach to AI, data integration and analytics. LeapAnalysis is the first fully federated and virtualised search and analytics engine that runs on semantic metadata. It allows users to combine semantic models (ontologies) with machine learning algorithms to provide customers with unparalleled flexibility in utilizing their data. Nearly all technologies surrounding AI and analytics are purely statistical in nature, using algorithmic approaches that are not incredibly novel, such as decision trees, neural networks, etc. The logical framework that contextualises these things is often missing.


Enabling Semantic Data Access for Toxicological Risk Assessment

arXiv.org Artificial Intelligence

Experimental effort and animal welfare are concerns when exploring the effects a compound has on an organism. Appropriate methods for extrapolating chemical effects can further mitigate these challenges. In this paper we present the efforts to (i) (pre)process and gather data from public and private sources, varying from tabular files to SPARQL endpoints, (ii) integrate the data and represent them as a knowledge graph with richer semantics. This knowledge graph is further applied to facilitate the retrieval of the relevant data for a ecological risk assessment task, extrapolation of effect data, where two prediction techniques are developed.


General Fragment Model for Information Artifacts

arXiv.org Artificial Intelligence

The use of semantic descriptions in data intensive domains require a systematic model for linking semantic descriptions with their manifestations in fragments of heterogeneous information and data objects. Such information heterogeneity requires a fragment model that is general enough to support the specification of anchors from conceptual models to multiple types of information artifacts. While diverse proposals of anchoring models exist in the literature, they are usually focused in audiovisual information. We propose a generalized fragment model that can be instantiated to different kinds of information artifacts. Our objective is to systematize the way in which fragments and anchors can be described in conceptual models, without committing to a specific vocabulary.


Commonsense Reasoning Using WordNet and SUMO: a Detailed Analysis

arXiv.org Artificial Intelligence

We describe a detailed analysis of a sample of large benchmark of commonsense reasoning problems that has been automatically obtained from WordNet, SUMO and their mapping. The objective is to provide a better assessment of the quality of both the benchmark and the involved knowledge resources for advanced commonsense reasoning tasks. By means of this analysis, we are able to detect some knowledge misalignments, mapping errors and lack of knowledge and resources. Our final objective is the extraction of some guidelines towards a better exploitation of this commonsense knowledge framework by the improvement of the included resources.


SQuAP-Ont: an Ontology of Software Quality Relational Factors from Financial Systems

arXiv.org Artificial Intelligence

Quality, architecture, and process are considered the keystones of software engineering. ISO defines them in three separate standards. However, their interaction has been scarcely studied, so far. The SQuAP model (Software Quality, Architecture, Process) describes twenty-eight main factors that impact on software quality in banking systems, and each factor is described as a relation among some characteristics from the three ISO standards. Hence, SQuAP makes such relations emerge rigorously, although informally. In this paper, we present SQuAP-Ont, an OWL ontology designed by following a well-established methodology based on the reuse of Ontology Design Patterns (i.e. SQuAP-Ont formalises the relations emerging from SQuAP to represent and reason via Linked Data about software engineering in a three-dimensional model consisting of quality, architecture, and process ISO characteristics. Industrial standards are widely used in the software engineering practice: they are built on preexisting literature and provide a common ground to scholars and practitioners to analyze, develop, and assess software systems. As far as software quality is concerned, the reference standard is the ISO/IEC 25010:2011 (ISO quality from now on), which defines the quality of software products and their usage (i.e., in-use quality). The ISO quality standard introduces eight characteristics that qualify a software product, and five characteristics that assess its quality in use. A characteristic is a parameter for measuring the quality of a software system-related aspect, e.g., reliability, usability, performance efficiency. The quantitative value associated with a characteristic is measured employing metrics that are dependent on the context of a specific software project and defined following the established literature.