Country
More Voices Than Ever? Quantifying Media Bias in Networks
Lin, Yu-Ru (Harvard University and Northeastern University) | Bagrow, James P. (Northeastern University and Harvard University) | Lazer, David (Northeastern University and Harvard University)
Social media, such as blogs, are often seen as democratic entities that allow more voices to be heard than the conventional mass or elite media. Some also feel that social media exhibits a balancing force against the arguably slanted elite media. A systematic comparison between social and mainstream media is necessary but challenging due to the scale and dynamic nature of modern communication. Here we propose empirical measures to quantify the extent and dynamics of social (blog) and mainstream (news) media bias. We focus on a particular form of bias--coverage quantity--as applied to stories about the 111th US Congress. We compare observed coverage of Members of Congress against a null model of unbiased coverage, testing for biases with respect to political party, popular front runners, regions of the country, and more. Our measures suggest distinct characteristics in news and blog media. A simple generative model, in agreement with data, reveals differences in the process of coverage selection between the two media.
Find Me the Right Content! Diversity-Based Sampling of Social Media Spaces for Topic-Centric Search
Choudhury, Munmun De (Rutgers, The State University of New Jersey) | Counts, Scott (Microsoft Research) | Czerwinski, Mary (Microsoft Research)
Social media and networking websites, such as Twitter and Facebook, generate large quantities of information and have become mechanisms for real-time content dissipation to users. An important question that arises is: how do we sample such social media information spaces in order to deliver relevant content on a topic to end users? Notice that these large-scale information spaces are inherently diverse, featuring a wide array of attributes such as location, recency, degree of diffusion effects in the network and so on. Naturally, for the end user, different levels of diversity in social media content can significantly impact the information consumption experience: low diversity can provide focused content that may be simpler to understand, while high diversity can increase breadth in the exposure to multiple opinions and perspectives. Hence to address our research question, we turn to diversity as a core concept in our proposed sampling methodology. Here we are motivated by ideas in the "compressive sensing" literature and utilize the notion of sparsity in social media information to represent such large spaces via a small number of basis components. Thereafter we use a greedy iterative clustering technique on this transformed space to construct samples matching a desired level of diversity. Based on Twitter Firehose data, we demonstrate quantitatively that our method is robust, and performs better than other baseline techniques over a variety of trending topics. In a user study, we further show that users find samples generated by our method to be more interesting and subjectively engaging compared to techniques inspired by state-of-the-art systems, with improvements in the range of 15--45%.
Making Project Team Recommendations from Online Information Sources
Earl, Charles C. (Virkaz Technologies) | Johnson, Amos (Morehouse College) | Yelpaala, Kaakpema (Yelpaala Good Advisors) | Good, Travis (Yelpaala Good Advisors)
We are developing an Internet platform called MediaTeam that provides a marketplace connecting media content consumers to communities of media content creators. The platform is enabled by our method for automated assembly of virtual project teams. Media creators use the automated team assembler to quickly identify and team with collaborators. The team assembly platform factors in how the skills, work, and communication styles of team members complement each other into its team recommendation process. We are now testing the teaming and collaboration platforms with video creators and seek to launch by the summer.
Areca: Online Comparison of Research Results
Urbansky, David (Dresden University of Technology) | Muthmann, Klemens (Dresden University of Technology) | Kreisz, Lars (Dresden University of Technology) | Schill, Alexander (Dresden University of Technology)
To experiment properly, scientists from many researchareas need large sets of real world data. Information re-trieval scientists for example often need to evaluate theiralgorithms on a dataset or a gold standard. The availabil-ity of these datasets often is insufficient and authors withthe same goal do not evaluate their approaches on thesame data. To make research results more transparentand comparable, we introduce Areca, an online portalfor sharing datasets and/or the results that were reachedwith the author’s algorithms on these datasets. Havingsuch an online comparison makes it easier to grasp thestate-of-the-art on certain tasks and drive research toimprove the results.
Why do People Retweet? Anti-Homophily Wins the Day!
Macskassy, Sofus A. ( Fetch Technologies ) | Michelson, Matthew (Fetch Technologies)
Twitter and other microblogs have rapidly become a significant means by which people communicate with the world and each other in near realtime. There has been a large number of studies surrounding these social media, focusing on areas such as information spread, various centrality measures, topic detection and more. However, one area which has not received much attention is trying to better understand what information is being spread and why it is being spread. This work looks to get a better understanding of what makes people spread information in tweets or microblogs through the use of retweeting. Several retweet behavior models are presented and evaluated on a Twitter data set consisting of over 768,000 tweets gathered from monitoring over 30,000 users for a period of one month. We evaluate the proposed models against each user and show how people use different retweet behavior models. For example, we find that although users in the majority of cases do not retweet information on topics that they themselves Tweet about as or from people who are "like them" (hence anti-homophily), we do find that models which do take homophily, or similarity, into account fits the observed retweet behaviors much better than other more general models which do not take this into account. We further find that, not surprisingly, people's retweeting behavior is better explained through multiple different models rather than one model.
Automatically Identifying Groups Based on Content and Collective Behavioral Patterns of Group Members
Gregory, Michelle (Pacific Northwest National Laboratory) | Engel, Dave W. (Pacific Northwest National Laboratory) | Bell, Eric (Pacific Northwest National Laboratory) | Piatt, Andy (Pacific Northwest National Laboratory) | Dowson, Scott (Pacific Northwest National Laboratory) | Cowell, Andrew (Pacific Northwest National Laboratory)
For example, on Live Journal1, there are a number of categories, gaming, for The explosion of popularity in social media, such as internet example, that one can categorize themselves and their forums, weblogs (blogs), wikis, etc., in the past decade blogs. While a number of those that self select that category has created a new opportunity to measure public opinion, may interact, there is no explicit requirement to do so. If attitude, and social structures (Agichtein et al. 2008, one is interested in marketing to a gaming crowd, for instance, Qualman 2010). A very common social structure investigated knowing all persons interested in gaming would be is online communities, or groups. There are a number useful, even if they do not interact directly with one another.
Viral Actions: Predicting Video View Counts Using Synchronous Sharing Behaviors
Shamma, David A. (Yahoo! Research) | Yew, Jude (University of Michigan) | Kennedy, Lyndon (Yahoo! Research) | Churchill, Elizabeth F. (Yahoo! Research)
In this article, we present a method for predicting the view count of a YouTube video using a small feature set collected from a synchronous sharing tool. We hypothesize that videos which have a high YouTube view count will exhibit a unique sharing pattern when shared in synchronous environments. Using a one-day sample of 2,188 dyadic sessions from the Yahoo! Zync synchronous sharing tool, we demonstrate how to predict the video's view count on YouTube, specifically if a video has over 10 million views. The prediction model is 95.8% accurate and done with a relatively small training set; only 15% of the videos had more than one session viewing; in effect, the classifier had a precision of 76.4% and a recall of 81%. We describe a prediction model that relies on using implicit social shared viewing behavior such as how many times a video was paused, rewound, or fast-forwarded as well as the duration of the session. Finally, we present some new directions for future virality research and for the design of future social media tools.
Creating Conversations: An Automated Dialog System
Gandy, Lisa (Northwestern University) | Hammond, Kristian (Northwestern University)
Online news sites often include a comments section where readers are allowed to leave their thoughts. These comments often contain interesting and insightful conversations between readers about the news article. However the richness of these conversations is often lost among other meaningless comments, and moreover all comments are found at the bottom of the web page. In this article, we discuss how our system inserts reader conversations into the news article to create a multimedia presentation called Shout Out. Shout Out features two virtual news anchors: one anchor reads the news and when appropriate the anchor pauses to have a conversation about the news with another anchor. This current iteration of Shout Out combines natural language techniques and reader conversations to create an engaging system.
A Machine Learning Approach to Twitter User Classification
Pennacchiotti, Marco (Yahoo! Labs) | Popescu, Ana-Maria (Yahoo! Labs)
This paper addresses the task of user classification in social media, with an application to Twitter. We automatically infer the values of user attributes such as political orientation or ethnicity by leveraging observable information such as the user behavior, network structure and the linguistic content of the user’s Twitter feed. We employ a machine learning approach which relies on a comprehensive set of features derived from such user information. We report encouraging experimental results on 3 tasks with different characteristics: political affiliation detection, ethnicity identification and detecting affinity for a particular business. Finally, our analysis shows that rich linguistic features prove consistently valuable across the 3 tasks and show great promise for additional user classification needs.
Political Polarization on Twitter
Conover, Michael D. (Indiana University) | Ratkiewicz, Jacob (Indiana University) | Francisco, Matthew (Indiana University) | Goncalves, Bruno (Indiana University) | Menczer, Filippo (Indiana University) | Flammini, Alessandro (Indiana University)
In this study we investigate how social media shape the networked public sphere and facilitate communication between communities with different political orientations. We examine two networks of political communication on Twitter, comprised of more than 250,000 tweets from the six weeks leading up to the 2010 U.S. congressional midterm elections. Using a combination of network clustering algorithms and manually-annotated data we demonstrate that the network of political retweets exhibits a highly segregated partisan structure, with extremely limited connectivity between left- and right-leaning users. Surprisingly this is not the case for the user-to-user mention network, which is dominated by a single politically heterogeneous cluster of users in which ideologically-opposed individuals interact at a much higher rate compared to the network of retweets. To explain the distinct topologies of the retweet and mention networks we conjecture that politically motivated individuals provoke interaction by injecting partisan content into information streams whose primary audience consists of ideologically-opposed users. We conclude with statistical evidence in support of this hypothesis.