Technology
An Eigenvalue-Based Measure for Word-Sense Disambiguation
Hulpus, Ioana (National University of Ireland) | Hayes, Conor (National University of Ireland) | Karnstedt, Marcel (National University of Ireland, Galway) | Greene, Derek (University College Dublin)
Current approaches for word-sense disambiguation (WSD) try to relate the senses of the target words by optimizing a score for each sense in the context of all other words' senses. However, by scoring each sense separately, they often fail to optimize the relations between the resulting senses. We address this problem by proposing a HITS-inspired method that attempts to optimize the score for the entire sense combination rather than one-word-at-a-time. We also exploit word-sense disambiguation via topic-models, when retrieving senses from heterogeneous sense inventories. Although this entails the relaxation of several assumptions behind current WSD algorithms, we show that our proposed method E-WSD achieves better results than current state-of-the-art approaches, without the need for additional background knowledge.
Proper Noun Semantic Clustering Using Bag-of-Vectors
Ebadat, Ali Reza (INRIA-INSA) | Claveau, Vincent (IRISA-CNRS) | Sébillot, Pascale (IRISA-INS)
In this paper, we propose a model for semantic clustering of entities extracted from a text, and we apply it to a Proper Noun classification task.This model is based on a new method to compute the similarity between the entities.Indeed, the classical way of calculating similarity is to build a feature vector or Bag-of-Features for each entity and then use classical similarity functions like Cosine.In practice, the features are contextual, such as words around the different occurrences of each entity. Here, we propose to use an alternative representation for entities, called Bag-of-Vectors, or Bag-of-Bags-of-Features.In this new model, each entity is not defined as a unique vector but as a set of vectors, in which each vector is built based on the contextual features of one occurrence of the entity.In order to use Bag-of-Vectors for clustering, we introduce new versions of classical similarity functions such as Cosine and Scalar Products. Experimentally, we show that the Bag-of-Vectors representation always improve the clustering results compared to classical Bag-of-Features representations.
Syntagmatic, Paradigmatic, and Automatic N-Gram Approaches to Assessing Essay Quality
Crossley, Scott (Georgia State University) | Cai, Zhiqiang (University of Memphis) | McNamara, Danielle S. (Arizona State University)
Computational indices related to n-gram production were developed in order to assess the potential for n-gram indices to predict human scores of essay quality. A regression analyses was conducted on a corpus of 313 argumentative essays. The analyses demonstrated that a variety of n-gram indices were highly correlated to essay quality, but were also highly correlated to the number of words in the text (although many of the n-gram indices were stronger predictors of writing quality than the number of words in a text). A second regression analysis was conducted on a corpus of 88 argumentative essays that were controlled for text length differences. This analysis demonstrated that n-gram indices were still strong predictors of essay quality when text length was not a factor.
Story-Level Inference and Gap Filling to Improve Machine Reading
Chalupsky, Hans (University of Southern California / Information Sciences Institute)
Machine reading aims at extracting formal knowledge representations from text to enable programs to execute some performance task, for example, diagnosis or answering complex queries stated in a formal representation language. Information extraction techniques are a natural starting point for machine reading, however, since they focus on explicit surface features at the phrase and sentence level, they generally miss information only stated implicitly. Moreover, the combination of multiple extraction results leads to error compounding which dramatically affects extraction quality for composite structures. To address these shortcomings, we present a new approach which aggregates locally extracted information into a larger story context and uses abductive constraint reasoning to generate the best story-level interpretation. We demonstrate that this approach significantly improves formal question answering performance on complex questions.
SenticNet 2: A Semantic and Affective Resource for Opinion Mining and Sentiment Analysis
Cambria, Erik (National University of Singapore) | Havasi, Catherine (MIT Media Lab) | Hussain, Amir (University of Stirling)
Web 2.0 has changed the ways people communicate, collaborate, and express their opinions and sentiments. But despite social data on the Web being perfectly suitable for human consumption, they remain hardly accessible to machines. To bridge the cognitive and affective gap between word-level natural language data and the concept-level sentiments conveyed by them, we developed SenticNet 2, a publicly available semantic and affective resource for opinion mining and sentiment analysis. SenticNet 2 is built by means of sentic computing, a new paradigm that exploits both AI and Semantic Web techniques to better recognize, interpret, and process natural language opinions. By providing the semantics and sentics (that is, the cognitive and affective information) associated with over 14,000 concepts, SenticNet 2 represents one of the most comprehensive semantic resources for the development of affect-sensitive applications in fields such as social data mining, multimodal affective HCI, and social media marketing.
Finding Associations between People
Blanco, Eduardo (Lymba Corporation) | Moldovan, Dan (Lymba Corporation)
Associations between people and other concepts are common in text and range from distant to close connections. This paper discusses and justifies the need to consider subtypes of the generic relation ASSOCIATION. Semantic primitives are used as a concise and formal way of specifying the key semantic differences between subtypes. A taxonomy of association relations is proposed, and a method based on composing previously extracted relations is used to extract subtypes. Experimental results show high precision and moderate recall.
The Devil Is in the Details: New Directions in Deception Analysis
McCarthy, Philip Michael (The University of Memphis ) | Duran, Nicholas D. (University of California Merced) | Booker, Lucille M. (The University of Memphis)
In this study, we use the computational textual analysis tool, the Gramulator, to identify and examine the distinctive linguistic features of deceptive and truthful discourse. The theme of the study is abortion rights and the deceptive texts are derived from a Devil’s Advocate approach, conducted to suppress personal beliefs and values. Our study takes the form of a contrastive corpus analysis, and produces systematic differences between truthful and deceptive personal accounts. Results suggest that deceivers employ a distancing strategy that is often associated with deceptive linguistic behavior. Ultimately, these deceivers struggle to adopt a truth perspective. Perhaps of most importance, our results indicate issues of concern with current deception detection theory and methodology. From a theoretical standpoint, our results question whether deceivers are deceiving at all or whether they are merely poorly expressing a rhetorical position, caused by being forced to speculate on a perceived proto-typical position. From a methodological standpoint, our results cause us to question the validity of deception corpora. Consequently, we propose new rigorous standards so as to better understand the subject matter of the deception field. Finally, we question the prevailing approach of abstract data measurement and call for future assessment to consider contextual lexical features. We conclude by suggesting a prudent approach to future research for fear that our eagerness to analyze and theorize may cause us to misidentify deception. After-all, successful deception, which is the kind we seek to detect, is likely to be an elusive and fickle prey.
Special Track on Applied Natural Language Processing
Boonthum-Denecke, Chutima (Hampton University)
Novel human-computer interfaces, for instance talking heads, can benefit from language understanding and generation techniques with big impact on user satisfaction. Dialoguebased intelligent tutoring systems require advanced dialogue processing, language understanding and generation components in order to assess students' natural language inputs and provide appropriate feedback. Moreover, language can facilitate human-computer interaction for the handicapped (no typing needed) and elderly leading to an ever increasing user base for computer systems. Some of the many areas emphasized by the ANLP track to include for contributions include multilingual processing, learning environments, multimodal communication, bioNLP, spam filtering, language acquisition (first and second), textual assessment, language varieties, materials development, generic classification, educational applications, information retrieval, speech processing, machine learning, knowledge representations, English for specific purposes, textual assessment indices, coreference resolution, word sense disambiguation, dialogue management and systems, language generation, language models, ontologies, and reasoning. For 2012, there were 15 submissions, out of which 10 were accepted as long papers and 3 as poster presentations.
Investigating the Robustness of Teager Energy Cepstrum Coefficients for Emotion Recognition in Noisy Conditions
Sun, Rui (Georgia Institute of Technology) | Moore, Elliot II (Georgia Institute of Technology)
This paper investigated the robustness of Teager Energy Cepstrum Coefficient (TECC) in differentiating emotion categories for speech at different White Gaussian noise levels by comparing the performance with MFCC. Experiments involved the normalized squared error measurement, the multi-classes (four classes) emotion classification and the pair-wise emotion classification. This study included four emotion categories (neutral, happy, sad, and happy) from three databases (two English, one German). The result showed that TECC performed equally or outperformed MFCC in both multi-emotion and pair-wise emotion classifications at all noise levels for all three databases. Using TECC features only, up to 89\% for the four-emotion classification and 99\% for the pair-wise emotion classification accuracy rate could be achieved.
Building an On-Demand Avatar-Based Health Intervention for Behavior Change
Lisetti, Christine (Florida International University) | Yasavur, Ugan (Florida International University) | Leon, Claudia de (Florida International University) | Amini, Reza (Florida International University) | Visser, Ubbo (University of Miami) | Rishe, Naphtali (Florida International University)
We discuss the design and implementation of the pro- totype of an avatar-based health system aimed at pro- viding people access to an effective behavior change intervention which can help them to find and cultivate motivation to change unhealthy lifestyles. An empathic Embodied Conversational Agent (ECA) delivers the in- tervention. The health dialog is directed by a compu- tational model of Motivational Interviewing, a novel effective face-to-face patient-centered counseling style which respects an individual’s pace toward behavior change. Although conducted on a small sample size, re- sults of a preliminary user study to asses users’ accep- tance of the avatar counselor indicate that the current early version of the system prototype is well accepted by 75% of users.