Goto

Collaborating Authors

 Technology


Enhancing Publication Description with Resources Metadata

AAAI Conferences

In this paper, we suggest to increase the quality and the precision of a document description using publication’s context description. Today, a lot of linguistic resources are both available on line and described by specific metadata. We first integrate them into an ontology which describes how linguists consider their primary data and tools. Then, we add to this ontology an inference system based on the information flow theory in order to establish causal relations between heterogeneous data. The result of the inference is characterized by a small set of properties which are embedded into three sequences of metadata enhancing the usual metadata describing publications.


Special Track on Artificial Intelligence, Cognitive Semantics, Computational Linguistics, and Logic

AAAI Conferences

Propositional attitudes in noncompositional logic are analysed from the view point of integration of their epistemic and deontic components. A new logical calculus for propositional attitudes inspired by possibility theory, a noncompositional version of fuzzy logic is proposed.


Applying Kernel Methods to Argumentation Mining

AAAI Conferences

The area of argumentation theory is an increasingly important area of artificial intelligence and mechanisms that are able to automatically detect the argument structure provide a novel area of research. This paper considers the use of kernel methods for argumentation detection and classification. It shows that a classification accuracy of 65%, can be attained using Natural Language Processing based kernel approaches, which do not require any heuristic choice of features.


A Linguistic Analysis of Expert-Generated Paraphrases

AAAI Conferences

The authors used the computational tool Coh-Metrix to examine expert writers’ paraphrases and in particular, how experts paraphrase text passages using condensing strategies. The overarching goal of this study was to develop machine learning algorithms to aid in the automatic detection of paraphrases and paraphrase types. To this end, three experts were instructed to paraphrase by condensing a set of target passages. The linguistic differences between the original passages and the condensed paraphrases were then analyzed using Coh-Metrix. The condensed paraphrases were accurately distinguished from the original target passages based on the number of words, word frequency, and syntactic complexity.


Addressing Semantic Ambiguities in Natural Language Constraints

AAAI Conferences

In NL2OCL project, we aim to translate English specification of constraints to formal constraints such as OCL (Object Constraint Language). In English to OCL translation, our contribution is a semantic analyzer that uses the output of the Stanford parser for shallow and deep semantic parsing. Our analysis of the output of shallow semantic parsing showed that semantic roles were mis-identified for a few English constraints due to semantic ambiguity. Similarly, in deep semantic parsing, it is difficult to resolve scope of quantifier operators due to scope ambiguity that is another sub-type of semantic ambiguity. In this paper, we highlight the identified cases of semantic ambiguities in English constraints. We also present a novel approach to automatically resolve the identified cases of the semantic ambiguities. The presented approach is also evaluated to show that by addressing the identified cases of semantic ambiguities, we can generate more accurate and complete formal (OCL) specifications.


Arabic Cross-Document NLP for the Hadith and Biography Literature

AAAI Conferences

Recently cross-document integration and reconciliation of extracted information became of interest to researchers in Arabic natural language processing. Given a set of documents $A$, we use Arabic morphological analysis, finite state machines, and graph transformations to extract named entities N a and relations R a expressed as edges in a graph G = ( N a, R a ). We use the same techniques to extract entities N b and relations R b from a separate set of documents B. We use G to disambiguate N b and R and we integrate the resulting entities back into G by annotating the nodes and edges in G with elements from N b . We apply our approach in an iterative manner. Our results show a significant increase in accuracy from 41% to 93% after applying this cross-document NLP methodology to hadith and biography documents.


Automatic Coherence Profile in Public Speeches of Three Latin American Heads-of-State

AAAI Conferences

Different studies provide evidence that the computational psycholinguistic algorithm called Latent Semantic Analysis (LSA) allows measuring local and global coherence in texts similarly to human evaluation (Foltz, Kintsch, Landauer 1998; McNamara, Cai & Louwerse 2007; McCarthy, Briner, Rus, & McNamara, 2007; McNamara, Louwerse & Jeuniaux 2009; Louwerse, McCarthy & Graesser 2010). The texts used in all these studies are written in English and correspond to scientific and literary texts. In Spanish, there are some studies using LSA that measure the semantic similarity between texts in automatic summary assessment (Pérez, Alfonseca, Rodríguez, Gliozzo, Strapparava & Magnini 2005; León, Olmos, Escudero, Cañas & Salmerón 2006; Venegas 2007, 2009, 2011); however, automatic measurement of coherence in Spanish has not yet been sufficiently investigated. The present study aimed at identifying a global and local coherence profile in a corpus of speeches in Spanish of three Latin American Heads-of-States (Perón, Castro and Pinochet), using Latent Semantic Analysis. Local coherence is calculated through the measurement of implicit semantic similarity between adjacent sentences and global coherence through the measurement of the similarity among the semantic content of the paragraphs. The corpus under analysis corresponds to a sample of 107 speeches. The semantic space was built using a multi-register corpus and it is available through the “Interface for the measurement of lexical-semantic similarity” in the El Grial interface (www.elgrial.cl). Results showed a systematic difference between the speeches of the Heads-of-State in terms of both local and global coherence. The Bonferroni analysis established an effect that distinguishes Perón’s speeches from Pinochet’s and Castro’s speeches. This results show that Perón’s speeches are more topically related than the other leaders’, probably due to a discourse strategy to persuade voters. The identification of a profile of coherence might be relevant to predict cues of government discourse styles.


Measuring Semantic Similarity in Short Texts through Greedy Pairing and Word Semantics

AAAI Conferences

We propose in this paper a greedy method to the problem of measuring semantic similarity between short texts. Our method is based on the principle of compositionality which states that the overall meaning of a sentence can be captured by summing up the meaning of its parts, i.e. the meanings of words in our case. Based on this principle, we extend word-to-word semantic similarity metrics to quantify the semantic similarity at sentence level. We report results using several word-to-word semantic similarity metrics, based on word knowledge or vectorial representations of meaning. Our approach performs better than similar approaches on the tasks of paraphrase identification and recognizing textual entailment, which are two illustrative semantic similarity tasks. We also report the role of word weighting and of function words on the performance of the proposed method.


A Comparative Study on English and Chinese Word Uses with LIWC

AAAI Conferences

This paper compared the linguistic and psychological word uses in English and Chinese languages with LIWC (Linguistic Inquiry and Word Count) programs. A Principal Component Analysis uncovered six linguistic and psychological components, among which five components were significantly correlated. The correlated components were ranked as Negative Valence (r=.92), Embodiment (r=.88), Narrative (r=.68), Achievement (r=.65), and Social Relation (r=.64). However, the results showed the order of the representative features differs in two languages and certain word categories co-occurred with different components in English and Chinese. The differences were interpreted from the perspective of distinctive eastern and western cultures.


Identifying Personality Types Using Document Classification Methods

AAAI Conferences

Are the words that people use indicative of their personality type preferences? In this paper, it is hypothesized that word-usage is not independent of personality type, as measured by the Myers-Briggs Type Indicator (MBTI) personality assessment tool. In-class writing samples were taken from 40 graduate students along with the MBTI. The experiment utilizes naïve Bayes classifiers and Support Vector Machines (SVMs) in an attempt to guess an individual’s personality type based on their word-choice. Classification is also attempted using emotional, social, cognitive, and psychological dimensions elicited by the analysis software, Linguistic Inquiry and Word Count (LIWC). The classifiers are evaluated with 40 distinct trials (leave-one-out cross validation), and parameters are chosen using leave-one-out cross validation of each trial’s training set. The experiment showed that the naïve Bayes classifiers (word-based and LIWC-based) outperformed the SVMs when guessing Sensing-Intuition (S-N) and Thinking-Feeling (T-F).