Technology
Online Social Capital: Mood, Topical and Psycholinguistic Analysis
Nguyen, Thin (Deakin University) | Dao, Bo (Deakin University) | Phung, Dinh (Deakin University) | Venkatesh, Svetha (Deakin University) | Berk, Michael (Deakin University)
Social media provides rich sources of personal information and community interaction which can be linked to aspect of mental health. In this paper we investigate manifest properties of textual messages, including latent topics, psycholinguistic features, and authors' mood, of a large corpus of blog posts, to analyze the aspect of social capital in social media communities. Using data collected from Live Journal, we find that bloggers with lower social capital have fewer positive moods and more negative moods than those with higher social capital. It is also found that people with low social capital have more random mood swings over time than the people with high social capital. Significant differences are found between low and high social capital groups when characterized by a set of latent topics and psycholinguistic features derived from blogposts, suggesting discriminative features, proved to be useful for classification tasks. Good prediction is achieved when classifying among social capital groups using topic and linguistic features, with linguistic features are found to have greater predictive power than latent topics. The significance of our work lies in the importance of online social capital to potential construction of automatic healthcare monitoring systems. We further establish the link between mood and social capital in online communities, suggesting the foundation of new systems to monitor online mental well-being.
What Yelp Fake Review Filter Might Be Doing?
Mukherjee, Arjun (University of Illinois at Chicago) | Venkataraman, Vivek (University of Illinois at Chicago) | Liu, Bing (University of Illinois at Chicago) | Glance, Natalie (Google Inc.)
Online reviews have become a valuable resource for decision making. However, its usefulness brings forth a curse ‒ deceptive opinion spam. In recent years, fake review detection has attracted significant attention. However, most review sites still do not publicly filter fake reviews. Yelp is an exception which has been filtering reviews over the past few years. However, Yelp’s algorithm is trade secret. In this work, we attempt to find out what Yelp might be doing by analyzing its filtered reviews. The results will be useful to other review hosting sites in their filtering effort. There are two main approaches to filtering: supervised and unsupervised learning. In terms of features used, there are also roughly two types: linguistic features and behavioral features. In this work, we will take a supervised approach as we can make use of Yelp’s filtered reviews for training. Existing approaches based on supervised learning are all based on pseudo fake reviews rather than fake reviews filtered by a commercial Web site. Recently, supervised learning using linguistic n-gram features has been shown to perform extremely well (attaining around 90% accuracy) in detecting crowdsourced fake reviews generated using Amazon Mechanical Turk (AMT). We put these existing research methods to the test and evaluate performance on the real-life Yelp data. To our surprise, the behavioral features perform very well, but the linguistic features are not as effective. To investigate, a novel information theoretic analysis is proposed to uncover the precise psycholinguistic difference between AMT reviews and Yelp reviews (crowdsourced vs. commercial fake reviews). We find something quite interesting. This analysis and experimental results allow us to postulate that Yelp’s filtering is reasonable and its filtering algorithm seems to be correlated with abnormal spamming behaviors.
Properties, Prediction, and Prevalence of Useful User-Generated Comments for Descriptive Annotation of Social Media Objects
Momeni, Elaheh (University of Vienna) | Cardie, Claire (Cornell University) | Ott, Myle (Cornell University)
User-generated comments in online social media have recently been gaining increasing attention as a viable source of general-purpose descriptive annotations for digital objects like photos or videos. Because users have different levels of expertise, however, the quality of their comments can vary from very useful to entirely useless. Our aim is to provide automated support for the curation of useful user-generated comments from public collections of digital objects. After constructing a crowd-sourced gold standard of useful and not useful comments, we use standard machine learning methods to develop a usefulness classifier, exploring the impact of surface-level, syntactic, semantic, and topic-based features in addition to extra-linguistic attributes of the author and his or her social media activity. We then adapt an existing model of prevalence detection that uses the learned classifier to investigate patterns in the commenting culture of two popular social media platforms. We find that the prevalence of useful comments is platform-specific and is further influenced by the entity type of the media object being commented on (person, place, event), its time period (e.g., year of an event), and the degree of polarization among commenters.
Detecting Comments on News Articles in Microblogs
Kothari, Alok (QCRI) | Magdy, Walid (QCRI) | Darwish, Kareem (QCRI) | Mourad, Ahmed (QCRI) | Taei, Ahmed (QCRI)
A reader of a news article would often be interested in the comments of other readers on an article, because comments give insight into popular opinions or feelings toward a given piece of news. In recent years, social media platforms, such as Twitter, have become a social hub for users to communicate and express their thoughts. This includes sharing news articles and commenting on them. In this paper, we propose an approach for identifying “comment-tweets” that comment on news articles. We discuss the nature of comment-tweets and compare them to subjective tweets. We utilize a machine learning-based classification approach for distinguishing between comment-tweets and others that only report the news. Our approach is evaluated on the TREC-2011 Microblog track data after applying additional annotations to tweets containing comments. Results show the effectiveness of our classification approach. Furthermore, we demonstrate the effectiveness of our approach on live news articles.
Exploiting Burstiness in Reviews for Review Spammer Detection
Fei, Geli (The University of Illinois at Chicago) | Mukherjee, Arjun (The University of Illinois at Chicago) | Liu, Bing (The University of Illinois at Chicago) | Hsu, Meichun (HP Labs) | Castellanos, Malu (HP Labs) | Ghosh, Riddhiman (HP Labs)
Online product reviews have become an important source of user opinions. Due to profit or fame, imposters have been writing deceptive or fake reviews to promote and/or to demote some target products or services. Such imposters are called review spammers. In the past few years, several approaches have been proposed to deal with the problem. In this work, we take a different approach, which exploits the burstiness nature of reviews to identify review spammers. Bursts of reviews can be either due to sudden popularity of products or spam attacks. Reviewers and reviews appearing in a burst are often related in the sense that spammers tend to work with other spammers and genuine reviewers tend to appear together with other genuine reviewers. This paves the way for us to build a network of reviewers appearing in different bursts. We then model reviewers and their co-occurrence in bursts as a Markov Random Field (MRF), and employ the Loopy Belief Propagation (LBP) method to infer whether a reviewer is a spammer or not in the graph. We also propose several features and employ feature induced message passing in the LBP framework for network inference. We further propose a novel evaluation method to evaluate the detected spammers automatically using supervised classification of their reviews. Additionally, we employ domain experts to perform a human evaluation of the identified spammers and non-spammers. Both the classification result and human evaluation result show that the proposed method outperforms strong baselines, which demonstrate the effectiveness of the method.
The Where and When of Finding New Friends: Analysis of a Location-based Social Discovery Network
Chen, Terence (National ICT Australia and University of New South Wales) | Kaafar, Mohamed Ali (National ICT Australia and INRIA) | Boreli, Roksana (National ICT Australia and University of New South Wales)
With more people accessing Online Social Networks (OSN) using their mobile devices, location-based features have become an important part of the social networking. In this paper, we present the first measurement study of a new category of location-based online social networking services, a location-based social discovery (LBSD) network, that enables users to discover and communicate with nearby people. Unlike popular check-in-based social networks, LBSD allows users to publicly reveal their locations without being associ- ated to a specific “venue” and their usage is not influenced by the incentive mechanisms of the underlying virtual community. By analyzing over 8 million user profiles and around 150 million location updates collected from a popular new LBSD network, we first present the characteristics of spatial- temporal usage patterns of the observed users, showing that 40% of updates are from the user’s primary location and 80% are from their top 10 locations. We identify events that trigger bursts of growth in subscriber numbers, showing the importance of social media marketing. Finally, we investigate how usage patterns may be utilized to re-identify individuals with e.g. different identifiers or from datasets belonging to different online services. We evaluate re-identification by usage, spatial and spatial-temporal patterns and using a number of metrics and show that the best results can be achieved using location data, with a high accuracy: our experiments demonstrate that we can re-identify up-to 85% of users with a precision of 77% using monitored spatial data. Overall, we find that although users exhibit strong periodic behavior in their usage pattern and movements, the success rate of re-identification is highly dependent on the level of activeness and the lifetime of the users in the network.
Blind Men and the Elephant: Detecting Evolving Groups in Social News
Bandari, Roja (University of California Los Angeles) | Rahmandad, Hazhir (Virginia Polytechnic Institute) | Roychowdhury, Vwani P (University of California Los Angeles)
We propose an automated and unsupervised methodology for a novel summarization of group behavior based on content preference. We show that graph theoretical community evolution (based on similarity of user preference for content) is effective in indexing these dynamics. Combined with text analysis that targets automatically-identified representative content for each community, our method produces a novel multi-layered representation of evolving group behavior. We demonstrate this methodology in the context of political discourse on a social news site with data that spans more than four years and find coexisting political leanings over extended periods and a disruptive external event that lead to a significant reorganization of existing patterns. Finally, where there exists no ground truth, we propose a new evaluation approach by using entropy measures as evidence of coherence along the evolution path of these groups. This methodology is valuable to designers and managers of online forums in need of granular analytics of user activity, as well as to researchers in social and political sciences who wish to extend their inquiries to large-scale data available on the web.
The International General Game Playing Competition
Genesereth, Michael ( Stanford University) | Björnsson, Yngvi (Reykjavik University)
Games have played a prominent role as a test-bed for advancements in the field of Artificial Intelligence ever since its foundation over half a century ago, resulting in highly specialized world-class game-playing systems being developed for various games. The establishment of the International General Game Playing Competition in 2005, however, resulted in a renewed interest in more general problem solving approaches to game playing. In general game playing (GGP) the goal is to create game-playing systems that autonomously learn how to skillfully play a wide variety of games, given only the descriptions of the game rules. In this paper we review the history of the competition, discuss progress made so far, and list outstanding research challenges.
The Annual Computer Poker Competition
Bard, Nolan (University of Alberta) | Hawkin, John (Verafin) | Rubin, Jonathan (PARC) | Zinkevich, Martin (Google)
Now entering its eighth year, the Annual Computer Poker Competition (ACPC) is the premier event within the field of computer poker. With both academic and nonacademic competitors from around the world, the competition provides an open and international venue for benchmarking computer poker agents. We describe the competition’s origins and evolution, current events, and winning techniques.