Personal Assistant Systems
The Use of Paraphrase Identification in the Retrieval of Appropriate Responses for Script Based Conversational Agents
McClendon, Jerome L. (Clemson University) | Mack, Naja A. (Clemson University) | Hodges, Larry F. (Clemson University)
This paper presents an approach to creating intelligent conversational agents that are capable of returning appropriate responses to natural language input. Our approach consists of using a supervised learning algorithm in combination with different NLP algorithms in training the system to identify paraphrases of the user’s question stored in a database. When tested on a data set consisting of questions and answers for a current conversational agent project, our approach returned an accuracy score of 79.15%, a precision score of 77.58%and a recall score of 78.01%.
Assessing Impacts of a Power User Attack on a Matrix Factorization Collaborative Recommender System
Seminario, Carlos E. (University of North Carolina at Charlotte) | Wilson, David C. (University of North Carolina at Charlotte)
Collaborative Filtering (CF) Recommender Systems (RSs) help users deal with the information overload they face when browsing, searching, or shopping for products and services. Power users are those individuals that are able to exert substantial influence over the recommendations made to other users, and RS operators encourage the existence of power user communities and leverage them to help fellow users make informed purchase decisions, especially on new items. Attacks on RSs occur when malicious users attempt to bias recommendations by introducing fake reviews or ratings; these attacks remain a key problem area for system operators. Thus, the influence wielded by power users can be used for both positive (addressing the "new item" problem) or negative (attack) purposes. Our research is investigating the impact on RS predictions and top-N recommendation lists when attackers emulate power users to provide biased ratings for new items. Previously we showed that power user attacks are effective against user-based CF RSs and that item-based CF RSs are robust to this type of attack. This paper presents the next stage in our investigation: (1) an evaluation of heuristic approaches to power user selection, and (2) evaluation of power user attacks in the context of matrix-factorization (SVD) based recommenders. Results show that social measures of influence such as degree centrality are more effective for selection of power users, and that matrix-factorization approaches are susceptible to power user attacks.
Differential Neighborhood Selection In Memory-Based Group Recommender Systems
Najjar, Nadia A (University of North Carolina at Charlotte) | Wilson, David C (University of North Carolina at Charlotte)
As recommender systems have become commonplace to support individual decision making, a need has also been recognized for systems that tailor and provide recommendations to a group of users together rather than individuals alone. Group recommender research to date has focused on evaluating strategies for aggregating profiles of group members to form a consolidated group profile or for aggregating recommendations to individual group members as a consolidated group recommendation list. This paper presents a novel neighborhood selection approach for group recommendation in the context of a neighborhood-based Collaborative Filtering system. We evaluate the performance of this approach with respect to group characteristics such as size and group member similarity. Results show that this approach can result in more accurate predictions for the group, particularly for groups that are more homogenous.
A Constrained Matrix-Variate Gaussian Process for Transposable Data
Koyejo, Oluwasanmi, Lee, Cheng, Ghosh, Joydeep
Transposable data represents interactions among two sets of entities, and are typically represented as a matrix containing the known interaction values. Additional side information may consist of feature vectors specific to entities corresponding to the rows and/or columns of such a matrix. Further information may also be available in the form of interactions or hierarchies among entities along the same mode (axis). We propose a novel approach for modeling transposable data with missing interactions given additional side information. The interactions are modeled as noisy observations from a latent noise free matrix generated from a matrix-variate Gaussian process. The construction of row and column covariances using side information provides a flexible mechanism for specifying a-priori knowledge of the row and column correlations in the data. Further, the use of such a prior combined with the side information enables predictions for new rows and columns not observed in the training data. In this work, we combine the matrix-variate Gaussian process model with low rank constraints. The constrained Gaussian process approach is applied to the prediction of hidden associations between genes and diseases using a small set of observed associations as well as prior covariances induced by gene-gene interaction networks and disease ontologies. The proposed approach is also applied to recommender systems data which involves predicting the item ratings of users using known associations as well as prior covariances induced by social networks. We present experimental results that highlight the performance of constrained matrix-variate Gaussian process as compared to state of the art approaches in each domain.
Personalized Recommendation of Twitter Lists using Content and Network Information
Rakesh, Vineeth (Wayne State University) | Singh, Dilpreet (Wayne State University) | Vinzamuri, Bhanukiran (Wayne State University) | Reddy, Chandan K (Wayne State University)
Lists in social networks have become popular tools to orga-nize content. This paper proposes a novel framework for rec-ommending lists to users by combining several features thatjointly capture their personal interests. Our contribution is oftwo-fold. First, we develop a ListRec model that leveragesthe dynamically varying tweet content, the network of twitterers and the popularity of lists to collectively model the users’preference towards social lists. Second, we use the topicalinterests of users, and the list network structure to developa novel network-based model called the LIST-PAGERANK.We use this model to recommend auxiliary lists that are morepopular than the lists that are currently subscribed by theusers. We evaluate our ListRec model using the Twitterdataset consisting of 2988 direct list subscriptions. Using au-tomatic evaluation technique, we compare the performanceof the ListRec model with different baseline methods andother competing approaches and show that our model deliversbetter precision in terms of the prediction of the subscribedlists of the twitterers. Furthermore, we also demonstrate the importance of combining different weighting schemes andtheir effect on capturing users’ interest towards Twitter lists.To evaluate the LIST-PAGERANK model, we employ a user-study based evaluation to show that the model is effective inrecommending auxiliary lists that are more authoritative thanthe lists subscribed by the users.
Predicting User Replying Behavior on a Large Online Dating Site
Xia, Peng (University of Massachusetts Lowell) | Jiang, Hua (Baihe.com) | Wang, Xiaodong (Baihe.com) | Chen, Cindy (University of Massachusetts Lowell) | Liu, Benyuan (University of Massachusetts Lowell)
Online dating sites have become popular platforms for people to look for potential romantic partners. Many online dating sites provide recommendations on compatible partners based on their proprietary matching algorithms. It is important that not only the recommended dates match the user's preference or criteria, but also the recommended users are interested in the user and likely to reciprocate when contacted. The goal of this paper is to predict whether an initial contact message from a user will be replied to by the receiver. The study is based on a large scale real-world dataset obtained from a major dating site in China with more than sixty million registered users. We formulate our reply prediction as a link prediction problem of social networks and approach it using a machine learning framework. The availability of a large amount of user profile information and the bipartite nature of the dating network present unique opportunities and challenges to the reply prediction problem. We extract user-based features from user profiles and graph-based features from the bipartite dating network, apply them in a variety of classification algorithms, and compare the utility of the features and performance of the classifiers. Our results show that the user-based and graph-based features result in similar performance, and can be used to effectively predict the reciprocal links. Only a small performance gain is achieved when both feature sets are used. Among the five classifiers we considered, random forests method outperforms the other four algorithms (naive Bayes, logistic regression, KNN, and SVM). Our methods and results can provide valuable guidelines to the design and performance of recommendation engine for online dating sites.
Collaborative Filtering with Information-Rich and Information-Sparse Entities
Zhu, Kai, Wu, Rui, Ying, Lei, Srikant, R.
In this paper, we consider a popular model for collaborative filtering in recommender systems where some users of a website rate some items, such as movies, and the goal is to recover the ratings of some or all of the unrated items of each user. In particular, we consider both the clustering model, where only users (or items) are clustered, and the co-clustering model, where both users and items are clustered, and further, we assume that some users rate many items (information-rich users) and some users rate only a few items (information-sparse users). When users (or items) are clustered, our algorithm can recover the rating matrix with $\omega(MK \log M)$ noisy entries while $MK$ entries are necessary, where $K$ is the number of clusters and $M$ is the number of items. In the case of co-clustering, we prove that $K^2$ entries are necessary for recovering the rating matrix, and our algorithm achieves this lower bound within a logarithmic factor when $K$ is sufficiently large. We compare our algorithms with a well-known algorithms called alternating minimization (AM), and a similarity score-based algorithm known as the popularity-among-friends (PAF) algorithm by applying all three to the MovieLens and Netflix data sets. Our co-clustering algorithm and AM have similar overall error rates when recovering the rating matrix, both of which are lower than the error rate under PAF. But more importantly, the error rate of our co-clustering algorithm is significantly lower than AM and PAF in the scenarios of interest in recommender systems: when recommending a few items to each user or when recommending items to users who only rated a few items (these users are the majority of the total user population). The performance difference increases even more when noise is added to the datasets.
Distributed Online Learning in Social Recommender Systems
Tekin, Cem, Zhang, Simpson, van der Schaar, Mihaela
In this paper, we consider decentralized sequential decision making in distributed online recommender systems, where items are recommended to users based on their search query as well as their specific background including history of bought items, gender and age, all of which comprise the context information of the user. In contrast to centralized recommender systems, in which there is a single centralized seller who has access to the complete inventory of items as well as the complete record of sales and user information, in decentralized recommender systems each seller/learner only has access to the inventory of items and user information for its own products and not the products and user information of other sellers, but can get commission if it sells an item of another seller. Therefore the sellers must distributedly find out for an incoming user which items to recommend (from the set of own items or items of another seller), in order to maximize the revenue from own sales and commissions. We formulate this problem as a cooperative contextual bandit problem, analytically bound the performance of the sellers compared to the best recommendation strategy given the complete realization of user arrivals and the inventory of items, as well as the context-dependent purchase probabilities of each item, and verify our results via numerical examples on a distributed data set adapted based on Amazon data. We evaluate the dependence of the performance of a seller on the inventory of items the seller has, the number of connections it has with the other sellers, and the commissions which the seller gets by selling items of other sellers to its users.
The AAAI-13 Conference Workshops
Agrawal, Vikas (IBM Research-India) | Archibald, Christopher (Mississippi State University) | Bhatt, Mehul (University of Bremen) | Bui, Hung (Nuance) | Cook, Diane J. (Washington State University) | Cortés, Juan (University of Toulouse) | Geib, Christopher (Drexel University) | Gogate, Vibhav (University of Texas at Dallas) | Guesgen, Hans W. (Massey University) | Jannach, Dietmar (TU Dortmund) | Johanson, Michael (University of Alberta) | Kersting, Kristian (University of Bonn) | Konidaris, George (Massachusetts Institute of Technology) | Kotthoff, Lars (University College Cork) | Michalowski, Martin (Adventium Labs) | Natarajan, Sriraam (Indiana University) | O' (University College Cork) | Sullivan, Barry (Naval Research Laboratory) | Pickett, Marc (University of Zagreb) | Podobnik, Vedran (University of British Columbia) | Poole, David (GM Research, India) | Shastri, Lokendra (George Mason University) | Shehu, Amarda (University of Central Florida) | Sukthankar, Gita
Benjamin Grosof (Coherent Knowledge from episodic memory to great progress is being made on methods Systems) on representing activity create semantic memory, using a combination to solve problems related to structure context through semantic rule methods, of semantic memory and prediction, motion simulation, deriving from experience in the episodic memory to guide users?
Online Matrix Completion Through Nuclear Norm Regularisation
Dhanjal, Charanpal, Gaudel, Romaric, Clémençon, Stéphan
It is the main goal of this paper to propose a novel method to perform matrix completion on-line. Motivated by a wide variety of applications, ranging from the design of recommender systems to sensor network localization through seismic data reconstruction, we consider the matrix completion problem when entries of the matrix of interest are observed gradually. Precisely, we place ourselves in the situation where the predictive rule should be refined incrementally, rather than recomputed from scratch each time the sample of observed entries increases. The extension of existing matrix completion methods to the sequential prediction context is indeed a major issue in the Big Data era, and yet little addressed in the literature. The algorithm promoted in this article builds upon the Soft Impute approach introduced in Mazumder et al. (2010). The major novelty essentially arises from the use of a randomised technique for both computing and updating the Singular Value Decomposition (SVD) involved in the algorithm. Though of disarming simplicity, the method proposed turns out to be very efficient, while requiring reduced computations. Several numerical experiments based on real datasets illustrating its performance are displayed, together with preliminary results giving it a theoretical basis.