Media
A Max-Norm Constrained Minimization Approach to 1-Bit Matrix Completion
We consider in this paper the problem of noisy 1-bit matrix completion under a general non-uniform sampling distribution using the max-norm as a convex relaxation for the rank. A max-norm constrained maximum likelihood estimate is introduced and studied. The rate of convergence for the estimate is obtained. Information-theoretical methods are used to establish a minimax lower bound under the general sampling model. The minimax upper and lower bounds together yield the optimal rate of convergence for the Frobenius norm loss. Computational algorithms and numerical performance are also discussed.
Listening to the Crowd: Automated Analysis of Events via Aggregated Twitter Sentiment
Hu, Yuheng (Arizona State University) | Wang, Fei (Arizona State University) | Kambhampati, Subbarao (Arizona State University,)
Individuals often express their opinions on social media platforms like Twitter and Facebook during public events such as the U.S. Presidential debate and the Oscar awards ceremony. Gleaning insights from these posts is of importance to analyzing the impact of the event. In this work, we consider the problem of identifying the segments and topics of an event that garnered praise or criticism, according to aggregated Twitter responses. We propose a flexible factorization framework, SocSent, to learn factors about segments, topics, and sentiments. To regulate the learning process, several constraints based on prior knowledge on sentiment lexicon, sentiment orientations (on a few tweets) as well as tweets alignments to the event are enforced. We implement our approach using simple update rules to get the optimal solution. We evaluate the proposed method both quantitatively and qualitatively on two large-scale tweet datasets associated with two events from different domains to show that it improves significantly over baseline models.
Topic Segmentation and Labeling in Asynchronous Conversations
Joty, S., Carenini, G., Ng, R. T.
Topic segmentation and labeling is often considered a prerequisite for higher-level conversation analysis and has been shown to be useful in many Natural Language Processing (NLP) applications. We present two new corpora of email and blog conversations annotated with topics, and evaluate annotator reliability for the segmentation and labeling tasks in these asynchronous conversations. We propose a complete computational framework for topic segmentation and labeling in asynchronous conversations. Our approach extends state-of-the-art methods by considering a fine-grained structure of an asynchronous conversation, along with other conversational features by applying recent graph-based methods for NLP. For topic segmentation, we propose two novel unsupervised models that exploit the fine-grained conversational structure, and a novel graph-theoretic supervised model that combines lexical, conversational and topic features. For topic labeling, we propose two novel (unsupervised) random walk models that respectively capture conversation specific clues from two different sources: the leading sentences and the fine-grained conversational structure. Empirical evaluation shows that the segmentation and the labeling performed by our best models beat the state-of-the-art, and are highly correlated with human annotations.
Rotunde — A Smart Meeting Cinematography Initiative — Tools, Datasets, and Benchmarks for Cognitive Interpretation and Control
Bhatt, Mehul (University of Bremen) | Suchan, Jakob (University of Bremen) | Freksa, Christian ( Spatial Cognition Research Center (SFB/TR 8), University of Bremen, Germany )
The cognitive interpretation of perceptual data (e.g., from video, depth, motion sensors) requires the representational and inferential mediation of commonsense and qualitative abstractions of space, actions, events, change, and interaction. General methods and benchmarks for high-level cognitive interpretation, and their seamless integration and access within large-scale projects concerned with cognitive vision, robotics, hybrid-intelligent systems are necessary. We present the Rotunde initiative as a particular instance of a challenging smart meeting cinematography concept primarily concerning human activity interpretation. The Rotunde initiative aims to release general tools (e.g., for reasoning and control), methodological and performance benchmarks, and developmental aids (e.g., management and visualisation of complex spatio-temporal data) for the cognitive interpretation of interaction.
Personalized Text-Based Music Retrieval
Hariri, Negar (DePaul University) | Mobasher, Bamshad (DePaul University) | Burke, Robin (DePaul University)
We consider the problem of personalized text-based music retrieval where users' history of preferences are taken into account in addition to their issued textual queries.Current retrieval methods mostly rely on songs meta-data. This limits the query vocabulary. Moreover, it is very costly to gather this information in large collections of music. Alternatively, we use music annotations retrieved from social tagging Websites such as last.fm and use them as textual descriptions of songs. Considering a user's profile and using preference patterns of music among all users, as in collaborative filtering approaches, can be useful in providing personalized and more satisfactory results. The main challenge is how to include both users' profiles and the songs meta-data in the retrieval model. In this paper, we propose a hierarchical probabilistic model that takes into account the users' preference history as well as tag co-occurrences in songs. Our model is an extension of LDA where topics are formed as joint clusterings of songs and tags. These topics capture the tag associations and user preferences and correspond to different music tastes. Each user's profile is represented as a distribution over topics which shows the user's interests in different types of music.We will explain how our model can be used for contextual retrieval. Our experimental results show significant improvement in retrieval when user profiles are taken into account.
A Comparison of Playlist Generation Strategies for Music Recommendation and a New Baseline Scheme
Bonnin, Geoffray (Technische Universität Dortmund) | Jannach, Dietmar (Technische Universität Dortmund)
The digitalization of music and the instant availability of millions of tracks on the Internet require new approaches to support the user in the exploration of these huge music collections. One possible approach to address this problem, which can also be found on popular online music platforms, is the use of user-created or automatically generated playlists (mixes). The automated generation of such playlists represents a particular type of the music recommendation problem with two special characteristics. First, the tracks of the list are usually consumed immediately at recommendation time; secondly, songs are listened to mostly in consecutive order so that the sequence of the recommended tracks can be relevant. In the past years, a number of different approaches for playlist generation have been proposed in the literature. In this paper, we review the existing core approaches to playlist generation, discuss aspects of appropriate offline evaluation designs and report the results of a comparative evaluation based on different datasets. Based on the insights from these experiments, we propose a comparably simple and computationally tractable new baseline algorithm for future comparisons, which is based on track popularity and artist information and is competitive with more sophisticated techniques in our evaluation settings.
Enforcing Meter in Finite-Length Markov Sequences
Roy, Pierre (Associate Researcher) | Pachet, Francois (Sony CSL Paris)
Markov processes are increasingly used to generate finite-length sequences that imitate a given style. However, Markov processes are notoriously difficult to control. Recently, Markov constraints have been introduced to give users some control on generated sequences. Markov constraints reformulate finite-length Markov sequence generation in the framework of constraint satisfaction (CSP). However, in practice, this approach is limited to local constraints and its performance is low for global constraints, such as cardinality or arithmetic constraints. This limitation prevents generated sequences to follow structural properties which are independent of the style, but inherent to the domain, such as meter. In this article, we introduce meter, a constraint that ensures a sequence is 1) Markovian with regards to a given corpus and 2) follows metrical rules expressed as cumulative cost functions. Additionally, meter can simultaneously enforce cardinality constraints. We propose a domain consistency algorithm whose complexity is pseudo-polynomial. This result is obtained thanks to a theorem on the growth of sumsets by Khovanskii. We illustrate our constraint on meter-constrained music generation problems that were so far not solvable by any other technique.
A Hierarchical Aspect-Sentiment Model for Online Reviews
Kim, Suin (KAIST) | Zhang, Jianwen (Microsoft Research Asia) | Chen, Zheng (Microsoft Research Asia) | Oh, Alice (KAIST) | Liu, Shixia (Microsoft Research Asia)
To help users quickly understand the major opinions from massive online reviews, it is important to automatically reveal the latent structure of the aspects, sentiment polarities, and the association between them. However, there is little work available to do this effectively. In this paper, we propose a hierarchical aspect sentiment model (HASM) to discover a hierarchical structure of aspect-based sentiments from unlabeled online reviews. In HASM, the whole structure is a tree. Each node itself is a two-level tree, whose root represents an aspect and the children represent the sentiment polarities associated with it. Each aspect or sentiment polarity is modeled as a distribution of words. To automatically extract both the structure and parameters of the tree, we use a Bayesian nonparametric model, recursive Chinese Restaurant Process (rCRP), as the prior and jointly infer the aspect-sentiment tree from the review texts. Experiments on two real datasets show that our model is comparable to two other hierarchical topic models in terms of quantitative measures of topic trees. It is also shown that our model achieves better sentence-level classification accuracy than previously proposed aspect-sentiment joint models.
Don’t Be Spoiled by Your Friends: Spoiler Detection in TV Program Tweets
Jeon, Sungho (The Attached Institute of Electronics and Telecommunications Research Institute) | Kim, Sungchul (POSTECH) | Yu, Hwanjo (POSTECH)
Providing a convenient mechanism for accessing the Internet, smartphones have led to the rapid growth of Social Networking Services (SNSs) such as Twitter and have served as a major platform for SNSs. Nowadays, people are able to check conveniently the SNS messages posted by their friends and followers via their smartphones. As a consequence, people are exposed to spoilers of TV programs that they follow. So far, there are two previous works that explored the detection of spoilers in texts, not SNS: (1) keyword matching method and (2) machine-learning method based on Latent Dirichlet Allocation (LDA). The keyword matching method evaluates most tweets as spoilers; hence its poor recall performance. The other method based on LDA, although successful on large text, works poorly on short segments of text such as those found on Twitter and evaluates most tweets as non-spoilers. This paper presents four features that are significant in the classification of spoiler tweets. Using those features, we classified spoiler tweets pertaining to a reality TV show (“Dancing with the Stars”). We experimentally compared our method with previous methods, with our method achieving substantially higher precision compared to the keyword matching and LDA-based methods while maintaining comparable recalls.
A Virtual Archive for the History of AI
Buchanan, Bruce G. (University of Pittsburgh) | Eckroth, Joshua (The Ohio State University) | Smith, Reid (Marathon Oil Corporation)
Publications that have influenced the growth of artificial intelligence are often difficult to obtain. We first collected titles of several thousand publications from many well-known sources and then selected about 1850 titles considered to be especially influential. We have identified, and in a few cases created, online versions of about half of these “classics in AI.” Searchable text of the documents enables additional analysis of trends and influences. Integration into the rest of the AITopics information portal contextualizes the classic publications.