Technology
Assuming Facts Are Expressed More Than Once
Betteridge, Justin (Carnegie Mellon University) | Ritter, Alan (Carnegie Mellon University) | Mitchell, Tom (Carnegie Mellon University)
Distant Supervision (DS) is a method for training sentence-level information extraction models using only an unlabeled corpus and a knowledge base (KB). Fundamental to many DS approaches is the assumption that KB facts are expressed at least once (EALO) in the text corpus. Often, however, KB facts are actually expressed in the corpus many times, in which cases EALO-based systems underuse the available training data. To address this problem, we introduce the "expressed at least alpha percent" (EALA) assumption, which asserts that expressions of KB facts account for up to alpha percent of the corresponding mentions. We show that for the same level of precision as the EALO approach, the EALA approach achieves up to 66 % higher recall on category recognition and 53 % higher recall on relation recognition.
Using Strong Lexical Association Extraction in an Understanding of Managersโ Decision Process
Simard, Frรฉdรฉric (Universitรฉ du Quรฉbec ร Trois-Riviรจres) | Biskri, Ismaรฏl (Universitรฉ du Quรฉbec ร Trois-Riviรจres) | St-Pierre, Josรฉe (Universitรฉ du Quรฉbec ร Trois-Riviรจres) | Bensaber, Boucif Amar (Universitรฉ du Quรฉbec ร Trois-Riviรจres)
Searching information or specific knowledge to understand decisions in a huge amount of data can be a difficult task. To support this task, classification is one of several used strategies. Algorithms used to support the process of automated classification leads to large and often noisy classes that is difficult to interpret. In this paper we present a method that exploits the notion of association rules and maximal association rules, in order to seek strong lexical associations in classes of similarities. We will show in experimentation section how these lexical associations can assist in understanding owner-managers decisions.
An Ensemble Approach to Adaptation-Guided Retrieval
Jalali, Vahid (Indiana University) | Leake, David (Indiana University)
Instance-based learning methods predict the solution of a case from the solutions of similar cases.However, solutions can be generated from less similar cases as well, provided appropriate case adaptation rules are available to adjust the prior solutions to account for dissimilarities. In fact, case-based reasoning research on adaptation-guided retrieval (AGR) shows that it may be beneficial to base retrieval decisions primarily on the availability of suitable adaptation knowledge, rather than on similarity. This paper proposes a new method for adaptation-guided retrieval for numerical prediction (regression) tasks. The method, EAGR (ensemble of adaptations-guided retrieval) works by retrieving an ensemble of cases, with a case favored for retrieval if there exists an ensemble of adaptation rules suitable for adapting its solution to the current problem. The solution for the input problem is then calculated by applying each retrieved case's ensemble of adaptations to that case, and combining the generated values. The approach is evaluated on four sample domains compared to three baseline methods: k-NN, an adaptation-guided retrieval approach, and a previous approach using ensembles of adaptations without adaptation-guided retrieval. EAGR improves accuracy in the tested domains compared to the other methods.
Scan Matching for Graph SLAM in Indoor Dynamic Scenarios
Yin, Jingchun (Politecnico di Torino, Italy) | Carlone, Luca (Politecnico di Torino, Italy) | Rosa, Stefano (Politecnico di Torino, Italy) | Anjum, Muhammad Latif (Politecnico di Torino, Italy) | Bona, Basilio (Politecnico di Torino, Italy)
SLAM (Simultaneous Localization And Mapping) plays an essential and important role for mobile robotic autonomous navigation. SLAM in dynamic environ- ments with moving objects is a challenging problem. We focus on scan matching for Graph-SLAM in indoor dynamic scenarios. Scan matching algorithm is pro- posed and implemented, which consists of the follow- ing phases: first, conditioned Hough Transform based segmentation is performed to extract and group line features; second, occupancy-analysis based moving ob- jects detection is done to detect and discard the seg- ments corresponding to the moving objects; third, lin- ear regression based line feature matching is executed to estimate the roto-translation parameters. Simulations to estimate roto-translation and the entire trajectory of the robot effectively verified the robustness of this al- gorithm in a dynamic scenario. The proposed algorithm is based on the line features of the indoor environment. It is robust to disturbances from moving objects in the dynamic scenario, and is especially suitable for the case when large rotational displacement is present.
Sentiment Analysis Using Dependency Trees and Named-Entities
Yasavur, Ugan (Florida International University) | Travieso, Jorge (Florida International University) | Lisetti, Christine (Florida International University) | Rishe, Naphtali David (Florida International University)
There is an increasing interest for valence and emotion sensing using a variety of signals. Text, as a communication channel, gathers a substantial amount of interest for recognizing its underlying sentiment (valence or polarity), affect or emotion (e.g. happy, sadness). We consider recognizing the valence of a sentence as a prior task to emotion sensing. In this article, we discuss our approach to classify sentences in terms of emotional valence. Our supervised system performs syntactic and semantic analysis for feature extraction. It processes the interactions between words in sentences by using dependency parse trees, and it can decide the current polarity of named-entities based on on-the-fly topic modeling. We compared 3 rule-based approaches and two supervised approaches (i.e. Naive Bayes and Maximum Entropy). We trained and tested our system using the SemEval-2007 affective text dataset, which contains news headlines extracted from news websites. Our results show that our systems outperform the systems demonstrated in SemEval-2007.
Semantic Enrichments in Text Supervised Classification: Application to Medical Domain
Albitar, Shereen (Aix-Marseille Universitรฉ, LSIS) | Espinasse, Bernard (Aix-Marseille Universitรฉ, LSIS) | Fournier, Sรฉbastien (Aix-Marseille Universitรฉ, LSIS)
The use of semantics in supervised text classification can improve its effectiveness especially in specific domains. Most state of the art works use concepts as an alternative to words in order to transform the classical bag of words (BOW) into a Bag of concepts (BOC). This transformation is done through conceptualization task. Furthermore, the resulting BOC can be enriched using other related concepts from semantic resources. This enrichment may enhance classification effectiveness as well. This paper focuses on two strategies for semantic enrichment of conceptualized text representation. The first one is based on semantic kernel method while the second one is based on enriching vectors method. These two semantic enrichment strategies are evaluated through experiments using Rocchio as the supervised classification method in the medical domain, using UMLS ontology and Ohsumed corpus.
Representing and Reasoning about Cultural Contexts in Intelligent Learning Environments
Mohammed, Phaedra (The University of the West Indies) | Mohan, Permanand (The University of the West Indies)
There is a growing interest within educational research to produce culturally-aware intelligent learning environments (ILEs) that capitalize on the affective benefits of positive cultural resonance and avoid the counter-productive effects of culturally ignorant designs. Several challenges arise when attempting to produce culturally-appropriate content for ILEs. These stem from the need for semantic representations of cultural conceptualisations that go beyond folk approaches, have sufficient details for intracultural reasoning, and which can be matched with the cultural backgrounds of students who use these ILEs. This paper tackles these challenges firstly through the formalism of a lower-level ontology for describing the cultural semantics commonly used in educational content and secondly with a software component for reasoning about this ontological knowledge in relation to student cultural backgrounds. An application was developed to test the practicality of the approach and assess its utility in locating culturally-appropriate educational resources for students. The evaluation results revealed that the majority of content selections made by the system were rated as highly appropriate by 90% of the participants on average and confirmed the viability of the approach.
Multi-Document Summarization Using Graph-Based Iterative Ranking Algorithms and Information Theoretical Distortion Measures
Samei, Borhan (University of Memphis) | Estiagh, Marzieh (Shiraz University, Shiraz, Iran) | Eshtiagh, Marzieh (Southeast Missouri State University) | Keshtkar, Fazel (Shiraz University) | Hashemi, Sattar (Shiraz University, Shiraz, Iran)
Text summarization is an important field in the area of natural language processing and text mining. This paper proposes an extraction-based model which uses graph-based and information theoretic concepts for multi-document summarization. Our method constructs a directed weighted graph from the original text by adding a vertex for each sentence, and compute a weighted edge between sentences which is based on distortion measures. In this paper we proposed a combination of these two models by representing the input as a graph, using distortion measures as the weight function and a ranking algorithm. Finally, a ranking algorithm is applied to identify the most important sentences to be included in the summary. By defining a proper distortion measure and ranking algorithm, this model gains promising results on the DUC2002 which is a well known real world data set. The results and ROUGE-1 scores of our model is fairly close to other successful models.