Technology
A Logic Prover Approach to Predicting Textual Similarity
Blanco, Eduardo (Lymba Corporation) | Moldovan, Dan (Lymba Corporation)
This paper presents a logic prover approach to predicting textual similarity. Sentences are represented using three logic forms capturing different levels of knowledge, from only content words to semantic representations extracted with an existing semantic parser. A logic prover is used to find proofs and derive semantic features that are combined in a machine learning framework. Experimental results show that incorporating the semantic structure of sentences yields better results than simpler pairwise word similarity measures.
Ontology-Based Named Entity Recognizer for Behavioral Health
Yasavur, Ugan (Florida International University) | Amini, Reza (Florida International University) | Lisetti, Christine (Florida International University) | Rishe, Naphtali (Florida International University )
Named-Entity Recognizers (NERs) are an important part of information extraction systems in annotation tasks. Although substantial progress has been made in recognizing domain-independent named entities (e.g. location, organization and person), there is a need to recognize named entities for domain-specific applications in order to extract relevant concepts. Due to the growing need for smart health applications in order to address some of the latest worldwide epidemics of behavioral issues (e.g. over eating, lack of exercise, alcohol and drug consumption), we focused on the domain of behavior change, especially {\em lifestyle change}. To the best of our knowledge, there is no named-entity recognizer designed for the lifestyle change domain to enable applications to recognize relevant concepts. We describe the design of an ontology for behavioral health based on which we developed a NER augmented with lexical resources. Our NER automatically tags words and phrases in sentences with relevant (lifestyle) domain-specific tags (e.g. [un/]healthy food, potentially-risky/healthy activity, drug, tobacco and alcoholic beverage). We discuss the evaluation that we conducted with with manually collected test data. In addition, we discuss how our ontology enables systems to make further information acquisition for the recognized named entities by using semantic reasoners.
A Gramulator Analysis of Gendered Language in Cable News Reportage
Wen, Xin (University of Memphis) | McCarthy, Philip Michael (Decooda International) | Strain, Amber Chauncey (University of Memphis)
News reportage is intended to serve the public in terms of nurturing a better understanding of political and societal concerns. But such a goal may be stymied if reporters lack a sufficient understanding of the effect gendered language may have on the conveyance and interpretation of news. To address this issue, we use the Gramulator to conduct an applied natural language processing study of the linguistic and topical features of gendered language in news reportage. Our goal is to offer some insights as to how the choice of language and topics might affect the efficacy of news reportage. Results suggest that current news reportage largely conforms to an established gender divide: Specifically, we find evidence that male reportage is more quantitative and likely to focus on topics such as politics, crime, and the military. By contrast, female reportage is more qualitative, and likely to focus on issues such as home and education. The study is of interest to all current affairs writers (e.g., journalists) because it offers a systematic approach to identifying and assessing the linguistic and topical differences that contribute to gendered language in new reportage.
Extending Word Highlighting in Multiparticipant Chat
Uthus, David C. (NRC/NRL Postdoctoral Fellow) | Aha, David W. (Naval Research Laboratory)
We describe initial work on extensions to word highlighting for multiparticipant chat to aid users in finding messages of interest, especially during times of high traffic in chat rooms. We have annotated a corpus of chat messages from a technical chat domain (Ubuntu’s technical support), indicating whether they are related to Ubuntu’s new desktop environment Unity. We also created an unsupervised learning algorithm, in which relations are represented with a graph, and applied this to find words related to Unity so they can be highlighted in new, unseen chat messages. On the task of finding relevant messages, our approach outperformed two baseline approaches that are similar to current state-of-the-art word highlighting methods in chat clients.
A Study of Probabilistic and Algebraic Methods for Semantic Similarity
Rus, Vasile (The University of Memphis) | Niraula, Nobal Bikram (The University of Memphis) | Banjade, Rajendra (The University of Memphis)
We study and propose in this article several novel solutions to the task of semantic similarity between two short texts. The proposed solutions are based on the probabilistic method of Latent Dirichlet Allocation (LDA) and on the algebraic method of Latent Semantic Analysis (LSA). Both methods, LDA and LSA, are completely automated methods used to discover latent topics or concepts from large collection of documents. We propose a novel word-to-word similarity measure based on LDA as well as several text-to-text similarity measures. We compare these measures with similar, known measures based on LSA. Experiments and results are presented on two data sets: the Microsoft Research Paraphrase corpus and the User Language Paraphrase corpus. We found that the novel word-to-word similarity measure based on LDA is extremely promising.
Towards Finding Relevant Information Graphics: Identifying the Independent and Dependent Axis from User-Written Queries
Li, Zhuo (University of Delaware) | Stagitis, Matthew (University of Delaware) | McCoy, Kathleen (University of Delaware) | Carberry, Sandra (University of Delaware)
Information graphics (non-pictorial graphics such as bar charts and line graphs) contain a great deal of knowledge. Information retrieval research has focused on retrieving textual documents and on extracting images based on words appearing in the accompanying article or based on low-level features such as color or texture. Our goal is to build a system for retrieving information graphics that reasons about the content of the graphic itself in deciding its relevance to the user query. As a first step, we aim to identify, from a full sentence user query, what should be depicted on the independent and dependent axes of potentially relevant graphs. Natural language processing techniques are used to extract features from the query and machine learning is employed to build a model for hypothesizing the content of the axes. Results have shown that our models can achieve accuracy higher than 80% on a corpus of collected user queries.
Automated Non-Content Word List Generation Using hLDA
Krug, Wayne (Language Computer Corporation) | Tomlinson, Marc T. (Language Computer Corporation)
In this paper, we present a language-independent method for the automatic, unsupervised extraction of non-content words from a corpus of documents. This method permits the creation of word lists that may be used in place of traditional function word lists in various natural language processing tasks. As an example we generated lists of words from a corpus of English, Chinese, and Russian posts extracted from Wikipedia articles and Wikipedia Wikitalk discussion pages. We applied these lists to the task of authorship attribution on this corpus to compare the effectiveness of lists of words extracted with this method to expert-created function word lists and frequent word lists (a common alternative to function word lists). hLDA lists perform comparably to frequent word lists. The trials also show that corpus-derived lists tend to perform better than more generic lists, and both sets of generated lists significantly outperformed the expert lists. Additionally, we evaluated the performance of an English expert list on machine translations of our Chinese and Russian documents, showing that our method also outperforms this alternative.
Using Automatic Scoring Models to Detect Changes in Student Writing in an Intelligent Tutoring System
Crossley, Scott (Georgia State University) | Roscoe, Rod (Arizona State University) | McNamara, Danielle (Arizona State University)
This study compares automated scoring increases and linguistic changes for student writers in two groups: a group that used an intelligent tutoring system embedded with an automated writing evaluation component (Writing Pal) and a group that used only the automated writing evaluation component. The primary goal is to examine automated scoring differences in both groups from pretest to posttest essays to investigate score gains and linguistic development. The study finds that both groups show significant increases in automated writing scores and significant development in lexical, syntactic, cohesion, and rhetorical features. However, the Writing-Pal group shows greater raw frequency gains (i.e., negative v. positive gains).
Novel Curve Signatures and a Combination Method for Thai On-Line Handwriting Character Recognition
Chaowicharat, Ekawat (Mahidol University) | Cercone, Nick (York University) | Naruedomkul, Kanlaya (Mahidol University)
There is no commercial character recognition software that supports Thai handwriting. Thai handwritten character recognition is needed to convert handwritten text written on mobile and tablet devices into computer encoded text. We propose a novel method that joins three curve signatures. The first signature is the normalized tangent angle function (TAF), which provides rough classification. The other two novel curve signatures are the relative position matrix (RPM), which is used to compare global curve features, and the straightened tangent angle function (STAF), which is used to compare the tangent angle along the cumulative unsigned curvature domain. In the recognition process, an input curve is extracted for these three signatures and the similarity against each character in the handwriting templates is measured. Then, the similarity scores are weighted and summed for ranking. Our experiment is done on 48 handwriting sample sets (44 Thai consonants appear in each set, and there are 4 sets per handwriting). Our methods yield an accuracy of 94.08% for personal handwriting, and 92.23% for general handwriting.
Automatic Detection of Nominal Entities in Speech for Enriched Content Search
Calix, Ricardo A. (Purdue University Calumet) | Javadpout, Leili (Louisiana State University) | Khazaeli, Mehdi (Louisiana State University) | Knapp, Gerald M. (Louisiana State University)
In this work, a methodology is developed to detect sentient actors in spoken stories. Meta-tags are then saved to XML files associated with the audio files. A recursive approach is used to find actor candidates and features which are then classified using machine learning approaches. Results of the study indicate that the methodology performed well on a narrative based corpus of children’s stories. Using Support Vector Machines for classification, an F-measure accuracy score of 86% was achieved for both named and unnamed entities. Additionally, feature analysis indicated that speech features were very useful when detecting unnamed actors.