Personal Assistant Systems
Transfer Learning in Collaborative Filtering for Sparsity Reduction
Pan, Weike (Hong Kong University of Science and Technology) | Xiang, Evan Wei (Hong Kong University of Science and Technology) | Liu, Nathan Nan (Hong Kong University of Science and Technology) | Yang, Qiang (Hong Kong University of Science and Technology)
Data sparsity is a major problem for collaborative filtering (CF) techniques in recommender systems, especially for new users and items. We observe that, while our target data are sparse for CF systems, related and relatively dense auxiliary data may already exist in some other more mature application domains. In this paper, we address the data sparsity problem in a target domain by transferring knowledge about both users and items from auxiliary data sources. We observe that in different domains the user feedbacks are often heterogeneous such as ratings vs. clicks. Our solution is to integrate both user and item knowledge in auxiliary data sources through a principled matrix-based transfer learning framework that takes into account the data heterogeneity. In particular, we discover the principle coordinates of both users and items in the auxiliary data matrices, and transfer them to the target domain in order to reduce the effect of data sparsity. We describe our method, which is known as coordinate system transfer or CST, and demonstrate its effectiveness in alleviating the data sparsity problem in collaborative filtering. We show that our proposed method can significantly outperform several state-of-the-art solutions for this problem.
UserRec: A User Recommendation Framework in Social Tagging Systems
Zhou, Tom Chao (The Chinese University of Hong Kong) | Ma, Hao (The Chinese University of Hong Kong) | Lyu, Michael R. (The Chinese University of Hong Kong) | King, Irwin (The Chinese University of Hong Kong)
Social tagging systems have emerged as an effective way for users to annotate and share objects on the Web. However, with the growth of social tagging systems, users are easily overwhelmed by the large amount of data and it is very difficult for users to dig out information that he/she is interested in. Though the tagging system has provided interest-based social network features to enable the user to keep track of other users' tagging activities, there is still no automatic and effective way for the user to discover other users with common interests. In this paper, we propose a User Recommendation (UserRec) framework for user interest modeling and interest-based user recommendation, aiming to boost information sharing among users with similar interests. Our work brings three major contributions to the research community: (1) we propose a tag-graph based community detection method to model the users' personal interests, which are further represented by discrete topic distributions; (2) the similarity values between users' topic distributions are measured by Kullback-Leibler divergence (KL-divergence), and the similarity values are further used to perform interest-based user recommendation; and (3) by analyzing users' roles in a tagging system, we find users' roles in a tagging system are similar to Web pages in the Internet. Experiments on tagging dataset of Web pages (Yahoo!~Delicious) show that UserRec outperforms other state-of-the-art recommender system approaches.
Learning to Extract Quality Discourse in Online Communities
Brennan, Michael Robert (Drexel University) | Wrazien, Stacy (Drexel University) | Greenstadt, Rachel (Drexel University)
Collaborative filtering systems have been developed to manage information overload and improve discourse in online communities. In such systems, users rank content provided by other users on the validity or usefulness within their particular context. The goal is that "good" content will rise to prominence and "bad" content will fade into obscurity. These filtering mechanisms are not well-understood and have known weaknesses. For example, they depend on the presence of a large crowd to rate content, but such a crowd may not be present. Additionally, the community's decisions determine which voices will reach a large audience and which will be silenced, but it is not known if these decisions represent "the wisdom of crowds" or a "censoring mob." Our approach uses statistical machine learning to predict community ratings. By extracting features that replicate the community's verdict, we can better understand collaborative filtering, improve the way the community uses the ratings of their members, and design agents that augment community decision-making. Slashdot is an example of such a community where peers will rate each others' comments based on their relevance to the post. This work extracts a wide variety of features from the Slashdot metadata and posts' linguistic contents to identify features that can predict the community rating. We find that author reputation, use of pronouns, and author sentiment are salient. We achieve 76% accuracy predicting community ratings as good, neutral, or bad.
A Second Chance to Make a First Impression: Factors Affecting the Longevity of Online Dating Relationships
Taylor, Lindsay Shaw (University of California, Berkeley) | Fiore, Andrew T. (University of California, Berkeley) | Mendelsohn, G. A. (University of California, Berkeley) | Cheshire, Coye (University of California, Berkeley)
This research explored the transition of romantic relationships from meeting online to the first face-to-face date. It is inevitable that impressions of a partner will change to some degree, but how much, and with what consequences? One hundred and fifty users of a popular online dating site participated in the study. They recalled a person whom they had met through the site, reporting their impressions of their partners from both before and after the first face-to-face meeting. We expected, based on prior research demonstrating the importance of physical attractiveness in romantic attraction both on- and offline, that changes in beliefs about partnersโ physical appeal would be the most powerful predictor of relationship longevity. However, they were unrelated to relationship success. Across all the dimensions we examined, impressions were in fact relatively stable, but when respondents said they knew their partners better after meeting face-to-face, relationships lasted longer.v
Star Quality: Aggregating Reviews to Rank Products and Merchants
McGlohon, Mary (Carnegie Mellon University, Google, Inc.) | Glance, Natalie (Google, Inc.) | Reiter, Zach (Google, Inc.)
Given a set of reviews of products or merchants from a wide range of authors and several reviews websites, how can we measure the true quality of the product or merchant?ย How do we remove the bias of individual authors or sources?ย How do we compare reviews obtained from different websites, where ratings may be on different scales (1-5 stars, A/B/C, etc.)?ย How do we filter out unreliable reviews to use only the ones with ``star quality''?ย Taking into account these considerations, we analyze data sets from a variety of different reviews sites (the first paper, to our knowledge, to do this). These data sets include 8 million product reviews and 1.5 million merchant reviews. We explore statistic- and heuristic- based models for estimating the true quality of a product or merchant, and compare the performance of these estimators on the task of ranking pairs of objects.ย We also apply the same models to the task of using Netflix ratings data to rank pairs of movies, and discover that the performance of the different models is surprisingly similar on this data set.
Effective Question Recommendation Based on Multiple Features for Question Answering Communities
Kabutoya, Yutaka (NTT Cyber Solutions Laboratories, NTT Corporation) | Iwata, Tomoharu (NTT Cyber Solutions Laboratories, NTT Corporation) | Shiohara, Hisako (NTT Cyber Solutions Laboratories, NTT Corporation) | Fujimura, Ko (NTT Cyber Solutions Laboratories, NTT Corporation)
We propose a new method of recommending questions to answerers so as to suit the answerersโ knowledge and interests in User-Interactive Question Answering (QA) communities. A question recommender can help answerers select the questions that interest them. This increases the number of answers, which will activate QA communities. An effective question recommender should satisfy the following three requirements: First, its accuracy should be higher than the existing category-based approach; more than 50% of answerers select the questions to answer according a fixed system of categories. Second, it should be able to recommend unanswered questions because more than 2,000 questions are posted every day. Third, it should be able to support even those people who have never answered a question previously, because more than 50% of users in current QA communities have never given any answer. To achieve an effective question recommender, we use question histories as well as the answer histories of each user by combining collaborative filtering schemes and content-base filtering schemes. Experiments on real log data sets of a famous Japanese QA community, Oshiete goo, show that our recommender satisfies the three requirements.
Using Linked Data to Build Open, Collaborative Recommender Systems
Heitmann, Benjamin (Digital Enterprise Research Institute, National University of Ireland, Galway) | Hayes, Conor (Digital Enterprise Research Institute, National University of Ireland, Galway)
While recommender systems can greatly enhance the user experience, the entry barriers in terms of data acquisition are very high, making it hard for new service providers to compete with existing recommendation services. This paper proposes to build open recommender systems which can utilise Linked Data to mitigate the new-user, new-item and sparsity problems of collaborative recommender systems. We describe how to aggregate data about object centred sociality from different sources and how to process it for collaborative recommendation. To demonstrate the validity of our approach, we augment the data from a closed collaborative music recommender system with Linked Data, and significantly improve its precision and recall.
Social Navigation through the Spoken Web: Improving Audio Access through Collaborative Filtering in Gujarat, India
Farrell, Robert (IBM Research) | Das, Rajarshi (IBM Research) | Rajput, Nitendra (IBM India Research Lab)
The rapid uptake of mobile phones, cheaper and more Given the potentially large number of users of the Spoken widespread mobile connectivity, and increasing familiarity Web system and the likelihood of shared information needs with technology are driving Internet adoption in developing and significant user similarities, we expect considerable improvements nations, but major hurdles still remain. First, today's Internet in audio navigation from using CF. is mostly in English and is thus largely inaccessible to A useful distinction among CFbased approaches arises billions of people for whom English is not a native or second from the types of data used to associate users to products language. Second, today's Internet is accessible largely and other items. In some scenarios, users may provide explicit through text-based technologies (web browsing, email, text feedback about their interest in products through ratings.
Efficiently Discovering Hammock Paths from Induced Similarity Networks
Hossain, M. Shahriar, Narayan, Michael, Ramakrishnan, Naren
Similarity networks are important abstractions in many information management applications such as recommender systems, corpora analysis, and medical informatics. For instance, by inducing similarity networks between movies rated similarly by users, or between documents containing common terms, and or between clinical trials involving the same themes, we can aim to find the global structure of connectivities underlying the data, and use the network as a basis to make connections between seemingly disparate entities. In the above applications, composing similarities between objects of interest finds uses in serendipitous recommendation, in storytelling, and in clinical diagnosis, respectively. We present an algorithmic framework for traversing similarity paths using the notion of `hammock' paths which are generalization of traditional paths. Our framework is exploratory in nature so that, given starting and ending objects of interest, it explores candidate objects for path following, and heuristics to admissibly estimate the potential for paths to lead to a desired destination. We present three diverse applications: exploring movie similarities in the Netflix dataset, exploring abstract similarities across the PubMed corpus, and exploring description similarities in a database of clinical trials. Experimental results demonstrate the potential of our approach for unstructured knowledge discovery in similarity networks.
Client-server multi-task learning from distributed datasets
Dinuzzo, Francesco, Pillonetto, Gianluigi, De Nicolao, Giuseppe
A client-server architecture to simultaneously solve multiple learning tasks from distributed datasets is described. In such architecture, each client is associated with an individual learning task and the associated dataset of examples. The goal of the architecture is to perform information fusion from multiple datasets while preserving privacy of individual data. The role of the server is to collect data in real-time from the clients and codify the information in a common database. The information coded in this database can be used by all the clients to solve their individual learning task, so that each client can exploit the informative content of all the datasets without actually having access to private data of others. The proposed algorithmic framework, based on regularization theory and kernel methods, uses a suitable class of mixed effect kernels. The new method is illustrated through a simulated music recommendation system.