Goto

Collaborating Authors

 Genre


Serious challenges before our schools, students and professionals

#artificialintelligence

A third to half the jobs that we are currently employed in would disappear in the next 15 years; and yet your child is being prepared in school for those very same jobs that won't exist by the time they graduate. Our curriculum prepares us for a lifetime career, but a child today can expect to change jobs at least seven times over the course of their lives – and five of those jobs don't exist yet. The coming days would see us pursuing careers that we cannot even imagine today. For instance your child could be an expert licensed drone pilot, or a cyber warrior in the army, a data analyst making sense of the peta bytes of data generated through our social interactions and trying to forecast our behavior. The other big challenge facing students today is that the velocity of technology changes has gained incredible speed; this is making knowledge obsolete faster than before.


Investigating Personalized Search in E-Commerce

AAAI Conferences

Personalized recommendations have become a common feature of many modern online services. In particular on e-commerce sites, one value of such recommendations is that they help consumers find items of interest in large product assortments more quickly. Many of today's sites take advantage of modern recommendation technologies to create personalized item suggestions for consumers navigating the site. However, limited research exists on the use of personalization and recommendation technology when consumers rely on the site's catalog search functionality to discover relevant items. In this work we explore the value of personalizing search results on e-commerce sites using recommendation technology. We design and evaluate different personalization strategies using log data of an online retail site. Our results show that considering several item relevance signals within the recommendation process in parallel leads to the best ranking of the search results. Specifically, the factors taken into account include the users' general interests, their most recent browsing behavior, as well as the consideration of current sales trends.


Dating Tablets in the Garshana Corpus

AAAI Conferences

This paper reports on an effort to assign dates to undated tablets in the Garshana corpus, a fully curated and annotated collection of Sumerian tables from the Ur III era. Of the 1488 tablets, 92% are dated, giving a strong training set for dating those undated. Two approaches were pursued: (1) a naive one which determined the date of a tablet by simply counting the name overlap with tablets of known year and (2) invoking a collection of machine learning algorithms of established robustness. The naive method reports an accuracy of 45-55% and the machine learning algorithms achieve 78.8-84.15% accuracy.


Evaluating Preprocessing Strategies for Time Series Prediction using Deep Learning Architectures

AAAI Conferences

We propose a novel approach to combine state-of-the-art time series data processing methods, such as symbolic aggregate approximation (SAX), with very recently developed deep neural network architectures, such as deep recurrent neural networks (DRNN), for time series data modeling and prediction. Time series data appear extensively in various scientific domains and industrial applications, yet the challenges in accurate modeling and prediction from such data remain open. Deep recurrent neural networks (DRNN) have been proposed as promising approaches to sequence prediction. We extend this research to the new challenge of the time series prediction space, building a system that effectively combines recurrent neural networks (RNN) with time series specific preprocessing techniques. Our experiments show comparisons of model performance with various data preprocessing techniques. We demonstrate that preprocessed inputs can steer us towards simpler (and therefore more computationally efficient) architectures of neural networks (when compared to original inputs).


Recurrence Quantification Analysis: A Technique for the Dynamical Analysis of Student Writing

AAAI Conferences

The current study examined the degree to which the quality and characteristics of students’ essays could be modeled through dynamic natural language processing analyses. Undergraduate students (n = 131) wrote timed, persuasive essays in response to an argumentative writing prompt. Recurrent patterns of the words in the essays were then analyzed using recurrence quantification analysis (RQA). Results of correlation and regression analyses revealed that the RQA indices were significantly related to the quality of students’ essays, at both holistic and sub-scale levels (e.g., organization, cohesion). Additionally, these indices were able to account for between 11% and 43% of the variance in students’ holistic and sub-scale essay scores. Overall, our results suggest that dynamic techniques can be used to improve natural language processing assessments of student essays.


Automatic Authorship Attribution of Noisy Documents

AAAI Conferences

In this survey, we conduct an investigation on the robustness of several features and classifiers in automatic authorship attribution. Our corpus consists in 25 different documents written by 5 different American philosophers in English. The different documents pass throw a digital conversion into grey-scaled images and several levels of noise are added to corrupt those image documents. The noise consists in a “Salt & Pepper” type, which is randomly added on the surface of the images with the following noise levels: 0%, 1%, 2%, 3%, 4%, 5%, 6% and 7%. Thus, each image goes throw an OCR program (Optical Character Recognition) to extract the text from the image. Then, the obtained text document is kept to be used during the experiments of authorship attribution. Several features and classifiers are employed and evaluated with regards to the classification performances. Results are quite interesting and show that the most robust feature in au-thorship attribution is the character-tetragram, which provides a score of 100% even at a noise level of 7%.


Identifying Original Projects in App Inventor

AAAI Conferences

Millions of users use online, open-ended blocks programming environments like App Inventor to learn how to program and to build personally meaningful programs and apps. As part of understanding the computational thinking concepts being learned by these users, we want to distinguish original projects that they create from unoriginal ones that arise from learning activities like tutorials and exercises. Given all the projects of students taking an App Inventor course, we describe how to automatically classify them as original vs. unoriginal using a hierarchical clustering technique. Although our current analysis focuses only on a small group of users (16 students taking a course in our institution) and their 902 projects, our findings establish a foundation for extending this analysis to larger groups of users.


A Text Mining Approach for Anomaly Detection in Application Layer DDoS Attacks

AAAI Conferences

Distributed Denial of Service (DDoS) attacks are a major threat to Internet security, with their use continuing to grow. Attackers are finding more sophisticated methods to attack servers. A lot of defense mechanisms have been proposed for DDoS attacks at IP and TCP layers. Those methods will not work well for application layer DDoS attacks that utilize legitimate application layer requests to overwhelm a webserver. These attacks look legitimate in both packets and protocol characteristics, which makes them harder to detect. In this paper, we propose an anomaly detection method to detect application layer DDoS attacks. We take a text mining approach to extract features which represent a user’s HTTP request sequence using bigrams. We apply the one class Support Vector Machine (SVM) algorithm on the extracted features from normal users’ HTTP request sequences. The one class SVM labels any newly seen instance that deviates from the normal, trained model as an application layer DDoS instance. We apply our experimental analysis on real web server logs collected from a student resource website. Three different variants of HTTP GET flood attacks are implemented on our server, generated via penetration testing. Our results show that the proposed method is able to detect application layer DDoS attacks with very good performance results.


Evolutionary Practice Problems Generation: More Design Guidelines

AAAI Conferences

We propose to further extend preliminary investigations of the nature of the problem of evolving practice problems for learners. Using a refinement of a previous simple model of interaction between learners and practice problems, we examine some of its properties and experimentally highlight the role played by the number of values each gene may take in our encoding of practice problems. We then experimentally compare both a traditional - P-CHC - and Pareto-based - P-PHC - variants of coevolutionary algorithms. Comparisons are conducted with respect to the presence of noise in fitness evaluations, the number of values genes may take, and two distinct fitness functions. Each fitness captures an aspect of the nature of learner-problem interaction but one has been shown to induce overspecialization pathologies. We then summarize our findings in terms of guidelines on how to adapt evolutionary algorithms to tackle the task of evolving practice problems.


Can Word Embeddings Help Find Latent Emotions in Text? Preliminary Results

AAAI Conferences

We report results of several experiments evaluating performance of word embeddings on semantic similarity of emotions. Our experiments suggest that the standard embeddings like GloVe and Word2Vec have very limited applicability in identifying emotions in text. Namely, using the standard arithmetic of emotions as a test, we show the mean reciprocal rank of a correct response is about 0.24, that is, combinations of word vectors are not a good proxy for expressed emotions. For example, the sum vector Joy+Fear, contrary to expectations, is not close to the vector representing Guilt. In addition, the opposite emotions, like Pessimism and Delight, have relatively high similarity to each other as word vectors (on average 0.2-0.44). Another experiment shows relatively low similarity (0.2-0.3) of word embeddings for similar emotions, such as Anger and Envy. Thus the standard methods for producing word embeddings are not adequate to represent relationships between emotion words. We conclude with a few hypotheses about improving the accuracy of embeddings in representing emotions.