Media
Leveraging Usage Data for Linked Data Movie Entity Summarization
Thalhammer, Andreas, Toma, Ioan, Roa-Valverde, Antonio, Fensel, Dieter
Novel research in the field of Linked Data focuses on the problem of entity summarization. This field addresses the problem of ranking features according to their importance for the task of identifying a particular entity. Next to a more human friendly presentation, these summarizations can play a central role for semantic search engines and semantic recommender systems. In current approaches, it has been tried to apply entity summarization based on patterns that are inherent to the regarded data. The proposed approach of this paper focuses on the movie domain. It utilizes usage data in order to support measuring the similarity between movie entities. Using this similarity it is possible to determine the k-nearest neighbors of an entity. This leads to the idea that features that entities share with their nearest neighbors can be considered as significant or important for these entities. Additionally, we introduce a downgrading factor (similar to TF-IDF) in order to overcome the high number of commonly occurring features. We exemplify the approach based on a movie-ratings dataset that has been linked to Freebase entities.
Using Web Services and Policies within a Social Platform to Support Collaborative Research
Pignotti, Edoardo (University of Aberdeen) | Edwards, Peter (University of Abeerdeen)
In this paper we present an architecture for provenance policies which can be used to describe and enact behavioural constraints in a system in order to ensure compliance with user and organisational policies. We discuss how this architecture has been used in order to manage the behaviour of the services powering an existing virtual research environment while reasoning about the relationships between users, their social network, their roles in a project, their groups and the provenance of research data.
Sifu: Interactive Crowd-Assisted Language Learning
Chan, Cheng-wei (National Taiwan University) | Hsu, Jane Yung-jen ( National Taiwan University )
This paper introduces SIFU, a system that recruits in real time native speakers as online volunteer tutors to help answer questions from Chinese language learners in reading news articles. SIFU integrates the strengths of two effective online language learning methods: reading online news and communicating with online native speakers. SIFU recruits volunteers from an online social network rather than recruits workers from Amazon Mechanical Turk.Initial experiments showed that the proposed approach is able to effectively recruit online volunteer tutors, adequately answer the learners' questions, and efficiently obtain an answer for the learner. Our field deployment illustrates that SIFU is very useful in assisting Chinese learners in reading Chinese news articles and online volunteer tutors are willing to help Chinese learners when they are on social network service.
Personalisation of Social Web Services in the Enterprise Using Spreading Activation for Multi-Source, Cross-Domain Recommendations
Heitmann, Benjamin (National University of Ireland, Galway) | Dabrowski, Maciej (National University of Ireland, Galway) | Passant, Alexandre (National University of Ireland, Galway) | Hayes, Conor (National University of Ireland, Galway) | Griffin, Keith (Cisco Systems)
Existing personalisation approaches, such as collaborative filtering or content based recommendations, are highly dependent on the domain and/or the source of the data. Therefore, there is a need for more accurate means to capture and model the interests of the user across domains, and to interlink them in a semantically-enhanced interest graph. We propose a new approach for multi-source, cross-genre recommendations that can exploit the heterogeneous nature of user profile data, which has been aggregated from multiple personalised web services, such as blogs, wikis and microblogs. Our approach is based on the Spreading Activation model that exploits intrinsic links between entities across a number of data sources. The proposed method is highly customizable and applicable both to generic and specific recommendation scenarios and use cases. With the growing number of Social Web applications in the enterprise (blogs, wikis, micro blogging, etc.), it becomes difficult for knowledge workers to avoid content overload and to quickly identify relevant people, communities and information. We demonstrate the application of our approach in an industrial use case that involves recommendation of social semantic data across multiple services in a distributed collaborative environment.
SNARE: Social Network Analysis and Reasoning Environment
Riecken, Doug (Columbia University) | Raja, Anita (University of North Carolina Charlotte/Columbia University) | Passonneau, Rebecca J. (Columbia University) | Waltz, David L. (Columbia University)
The importance of diversity in reasoning and learning to successfully address complex problems is examined. We discuss an approach by which a multiagent framework with decentralized control mechanisms provides diverse perspectives and hypotheses addressing a class of complex problems. We introduce the SNARE multiagent system. SNARE performs tasks to gain situational awareness of situations of interest in a Social Media Space. It applies a decentralized control mechanism for each agent; this mechanism enables an agent to interact with other agents to reason and learn. This approach facilitates dynamic agent organizations that adapt the topologies of interactions between agents based on the problem context.
Bayesian exponential family projections for coupled data sources
Klami, Arto, Virtanen, Seppo, Kaski, Samuel
Exponential family extensions of principal component analysis (EPCA) have received a considerable amount of attention in recent years, demonstrating the growing need for basic modeling tools that do not assume the squared loss or Gaussian distribution. We extend the EPCA model toolbox by presenting the first exponential family multi-view learning methods of the partial least squares and canonical correlation analysis, based on a unified representation of EPCA as matrix factorization of the natural parameters of exponential family. The models are based on a new family of priors that are generally usable for all such factorizations. We also introduce new inference strategies, and demonstrate how the methods outperform earlier ones when the Gaussianity assumption does not hold.
Around the Water Cooler: Shared Discussion Topics and Contact Closeness in Social Search
Komanduri, Saranga (Carnegie Mellon University) | Fang, Lujun (University of Michigan at Ann Arbor) | Huffaker, David (Google, Inc) | Staddon, Jessica (Google, Inc)
Search engines are now augmenting search results with social annotations, i.e., endorsements from usersโ social network contacts. However, there is currently a dearth of published research on the effects of these annotations on user choice. This work investigates two research questions associated with annotations: 1) do some contacts affect user choice more than others, and 2) are annotations relevant across various information needs. We conduct a controlled experiment with 355 participants, using hypothetical searches and annotations, and elicit usersโ choices. We find that domain contacts are preferred to close contacts, and this preference persists across a variety of information needs. Further, these contacts need not be experts and might be identified easily from conversation data.
A Temporal Analysis of Posting Behavior in Social Media Streams
Lee, Bumsuk (The Catholic University of Korea)
In this work, we investigated the social media streams to understand their characteristics and their temporal aspects. We assumed that each blogger has different temporal preference for posting. To investigate this hypothesis, we analyzed a massive dataset, nearly 700,000 blog articles, with the consideration of two factors which are day of the week and time of the day. The comparison was done in manifold ways: Blogosphere vs. Twitter, commercial blogs vs. non-commercial blogs, and their individuals. We hope that this work provides a hint to develop a personalized system which can be used for the reduction of the system resources for pull/fetch technology.
Identifying Microblogs for Targeted Contextual Advertising
Dave, Kushal Shailesh (International Institute of Information Technology, Hyderabad) | Varma, Vasudeva (International Institute of Information Technology, Hyderabad)
Micro-blogging sites such as Facebook, Twitter, Google+ present a nice opportunity for targeting advertisements that are contextually related to the microblog content. By virtue of the sparse and noisy text makes identifying the microblogs suitable for advertising a very hard problem. In this work, we approach the problem of identifying the microblogs that could be targeted for advertisements as a two-step classification approach. In the first pass, microblogs suitable for advertising are identified. Next, in the second pass, we build a model to find the sentiment of the advertisable microblog. The systems use features derived from the Part-of-speech tags, the tweet content and uses external resources such as query logs and n-gram dictionaries from previously labeled data.This work aims at providing a thorough insight into the problem and analyzing various features to assess which features contribute the most towards identifying the tweets that can be targeted for advertisements.
Emotional Divergence Influences Information Spreading in Twitter
Pfitzner, Rene (ETH Zurich) | Garas, Antonios (ETH Zurich) | Schweitzer, Frank (ETH Zurich)
We analyze data about the micro-blogging site Twitter using sentiment extraction techniques. From an information perspective, Twitter users are involved mostly in two processes: information creation and subsequent distribution (tweeting), and pure information distribution (retweeting), with pronounced preference to the first. However a rather substantial fraction of tweets are retweeted. Here, we address the role of the sentiment expressed in tweets for their potential aftermath. We find that although the overall sentiment (polarity) does not influence the probability of a tweet to be retweeted, a new measure called "emotional divergence" does have an impact. In general, tweets with high emotional diversity have a better chance of being retweeted, hence influencing the distribution of information.