Country
Around the Water Cooler: Shared Discussion Topics and Contact Closeness in Social Search
Komanduri, Saranga (Carnegie Mellon University) | Fang, Lujun (University of Michigan at Ann Arbor) | Huffaker, David (Google, Inc) | Staddon, Jessica (Google, Inc)
Search engines are now augmenting search results with social annotations, i.e., endorsements from usersโ social network contacts. However, there is currently a dearth of published research on the effects of these annotations on user choice. This work investigates two research questions associated with annotations: 1) do some contacts affect user choice more than others, and 2) are annotations relevant across various information needs. We conduct a controlled experiment with 355 participants, using hypothetical searches and annotations, and elicit usersโ choices. We find that domain contacts are preferred to close contacts, and this preference persists across a variety of information needs. Further, these contacts need not be experts and might be identified easily from conversation data.
Distributional Footprints of Deceptive Product Reviews
Feng, Song (Stony Brook University) | Xing, Longfei (Stony Brook University) | Gogar, Anupam (Stony Brook University) | Choi, Yejin (Stony Brook University)
This paper postulates that there are natural distributions of opinions in product reviews. In particular, we hypothesize that for a given domain, there is a set of representative distributions of review rating scores. A deceptive business entity that hires people to write fake reviews will necessarily distort its distribution of review scores, leaving distributional footprints behind. In order to validate this hypothesis, we introduce strategies to create dataset with pseudo-gold standard that is labeled automatically based on different types of distributional footprints. A range of experiments confirm the hypothesized connection between the distributional anomaly and deceptive reviews. This study also provides novel quantitative insights into the characteristics of natural distributions of opinions in the TripAdvisor hotel review and the Amazon product review domains.
Homophily and Latent Attribute Inference: Inferring Latent Attributes of Twitter Users from Neighbors
Zamal, Faiyaz Al (McGill University) | Liu, Wendy (McGill University) | Ruths, Derek (McGill University)
In this paper, we extend existing work on latent attribute inference by leveraging the principle of homophily: we evaluate the inference accuracy gained by augmenting the user features with features derived from the Twitter profiles and postings of her friends. We consider three attributes which have varying degrees of assortativity: gender, age, and political affiliation. Our approach yields a significant and robust increase in accuracy for both age and political affiliation, indicating that our approach boosts performance for attributes with moderate to high assortativity. Furthermore, different neighborhood subsets yielded optimal performance for different attributes, suggesting that different subsamples of the user's neighborhood characterize different aspects of the user herself. Finally, inferences using only the features of a user's neighbors outperformed those based on the user's features alone. This suggests that the neighborhood context alone carries substantial information about the user.
You Too?! Mixed-Initiative LDA Story Matching to Help Teens in Distress
Dinakar, Karthik (Massachusetts Institute of Technology) | Jones, Birago (Massachusetts Institute of Technology) | Lieberman, Henry (Massachusetts Institute of Technology) | Picard, Rosalind (Massachusetts Institute of Technology) | Rose, Carolyn (Carnegie Mellon University) | Thoman, Matthew (Northeastern University) | Reichart, Roi (Massachusetts Institute of Technology)
Adolescent cyber-bullying on social networks is a phenomenon that has received widespread attention. Recent work by sociologists has examined this phenomenon under the larger context of teenage drama and it's manifestations on social networks. Tackling cyber-bullying involves two key components โ automatic detection of possible cases, and interaction strategies that encourage reflection and emotional support. Key is showing distressed teenagers that they are not alone in their plight. Conventional topic spotting and document classification into labels like "dating" or "sports" are not enough to effectively match stories for this task. In this work, we examine a corpus of 5500 stories from distressed teenagers from a major youth social network. We combine Latent Dirichlet Allocation and human interpretation of its output using principles from sociolinguistics to extract high-level themes in the stories and use them to match new stories to similar ones. A user evaluation of the story matching shows that theme-based retrieval does a better job of finding relevant and effective stories for this application than conventional approaches.
Tweetin' in the Rain: Exploring Societal-Scale Effects of Weather on Mood
Hannak, Aniko (Northeastern University) | Anderson, Eric (Northeastern University) | Barrett, Lisa Feldman (Northeastern University) | Lehmann, Sune (Technical University of Denmark) | Mislove, Alan (Northeastern University) | Riedewald, Mirek (Northeastern University)
There has been significant recent interest in using the aggregate sentiment from social media sites to understand and predict real-world phenomena. However, the data from social media sites also offers a unique and โ so far โ unexplored opportunity to study the impact of external factors on aggregate sentiment, at the scale of a society. Using a Twitter-specific sentiment extraction methodology, we the explore patterns of sentiment present in a corpus of over 1.5 billion tweets. We focus primarily on the effect of the weather and time on aggregate sentiment, evaluating how clearly the well-known individual patterns translate into population-wide patterns. Using machine learning techniques on the Twitter corpus correlated with the weather at the time and location of the tweets, we find that aggregate sentiment follows distinct climate, temporal, and seasonal patterns.
Mixed Membership Models for Exploring User Roles in Online Fora
White, Arthur J. (University College Dublin) | Chan, Jeffrey (University of Melbourne) | Hayes, Conor (National University Ireland Galway) | Murphy, Brendan (University College Dublin)
Discussion boards are a form of social media which allow users to discuss topics and exchange information in a complex manner, in a number of different settings. As the popularity of such message boards has increased, communities of users have emerged, and several prominent types of social role have been identified, such as Question Answerer, Celebrity, Discussion Person and Topic Initiator. Recent studies have noted the structural similarity of the egocentric network of users assigned the same role by qualitative criteria. In this paper a methodology is developed with which to cluster together users with similar ego-centric network structures. This is achieved using a mixed membership formulation which allows for the fact that different groups of users may have characteristics in common. The method is then applied to data taken from boards.ie, a medium sized message boards website. Prominent clusters of users are identified and discussed, and illustrative examples of user behaviour provided. The type of interaction, both locally and globally, taking place within forums is examined.
Grief-Stricken in a Crowd: The Language of Bereavement and Distress in Social Media
Brubaker, Jed R. (University of California, Irvine) | Kivran-Swaine, Funda (Rutgers University) | Taber, Lee (University of California, Irvine) | Hayes, Gillian R. (University of California, Irvine)
People turn to social media to express their emotions surrounding major life events. Death of a loved one is one scenario in which people share their feelings in the semi-public space of social networking sites. In this paper, we present the results of a two-part investigation of grief and distress in the context of messages posted to the profiles of deceased MySpace users. We present coding system for identifying emotion distressed content, followed by a detailed analysis of language use that lays a foundation for natural language processing (NLP) tasks, such as automatic detection of bereavement-related distress. Our findings suggest that in addition to words bearing positive or negative sentiment, linguistic style can be an indicator of messages that demonstrate distress in the space of post-mortem social media content. These results contribute to research in computational linguistics by identifying linguistic features that can be used for automatic classification as well as to research on death and bereavement by enumerating attributes of distressed self-expression in a post-mortem context.
Tutorials
Breslin, John (National University of Ireland, Galway)
The ICWSM 2012 conference tutorials will be How to Analyze Massive Social Network Datasets without a Cluster, presented by Derek Ruths; Charting Collections of Connections in Social Media: Creating Maps and Measures with NodeXL, presented by Marc Smith; Evidenced-Based Social Design of Online Communities: Getting to Critical Mass and Encouraging Contributions, presented by Paul Resnick and Robert Kraut; Sentiment Mining from User Generated Content, presented by Lyle Ungar and Ronen Feldman; and Information Extraction for Social Media Anaylsis, presented by Denilson Barbosa.
Cross-Community Influence in Discussion Fora
Belรกk, Vรกclav (National University of Ireland, Galway) | Lam, Samantha (National University of Ireland, Galway) | Hayes, Conor (National University of Ireland, Galway)
Online discussion fora have become an important cultural and business asset in the context of many services provided by both non-profit organizations and enterprises. In order to keep and eventually increase the value these systems deliver to their users, it is often necessary to moderate or even manage their dynamics. One way to do this efficiently is to focus primarily on the most influential actors in the system. However, identifying such users becomes increasingly hard with systems where there is a continuously growing large user base. We show that analysis and explanation of influence on the cross-community level is a promising way to provide a coarse-grained picture of a potentially very large system and that it may enable its stakeholders to find groups through which the system can be efficiently influenced, or it can help them to identify and avoid activity considered as malicious. In order to achieve that, we present a novel framework for cross-community influence analysis, which is evaluated on 10 years of data from the largest Irish online discussion system Boards.ie.
A Systematic Investigation of Blocking Strategies for Real-Time Classification of Social Media Content into Events
Reuter, Timo (CITEC, Universität Bielefeld) | Cimiano, Philipp (CITEC, Universität Bielefeld)
Events play a prominent role in our lives, such that many social media documents describe or are related to some event. Organizing social media documents with respect to events thus seems a promising approach to better manage and organize the ever-increasing amount of user-generated content in social media applications. It would support the navigation of data by events or allow one to get notified about new postings related to the events one is interested in, just to name two applications. A challenge is to automatize this process so that incoming documents can be assigned to their corresponding event without any user intervention. We present a system that is able to classify a stream of social media data into a growing and evolving set of events. In order to scale up to the data sizes and data rates in social media applications, the use of a candidate retrieval or blocking step is crucial to reduce the number of events that are considered as potential candidates to which the incoming data point could belong to.In this paper we present and experimentally compare different blocking strategies along their cost vs. effectiveness tradeoff.We show that using a blocking strategy that selects the 60 closest events with respect to upload time, we reach F-Measures of about 85.1% while being able to process the incoming documents within 32ms on average. We thus provide a principled approach supporting to scale up classification of social media documents into events and to process the incoming stream of documents in real time.