Information Retrieval
Supervised Hashing via Uncorrelated Component Analysis
Sohn, SungRyull (Electronics and Telecommunications Research Institute and Korea Advanced Institute of Science and Technology) | Kim, Hyunwoo (Kakao Corp.) | Kim, Junmo (Korea Advanced Institute of Science and Technology)
The Approximate Nearest Neighbor (ANN) search problem is important in applications such as information retrieval. Several hashing-based search methods that provide effective solutions to the ANN search problem have been proposed. However, most of these focus on similarity preservation and coding error minimization, and pay little attention to optimizing the precision-recall curve or receiver operating characteristic curve. In this paper, we propose a novel projection-based hashing method that attempts to maximize the precision and recall. We first introduce an uncorrelated component analysis (UCA) by examining the precision and recall, and then propose a UCA-based hashing method. The proposed method is evaluated with a variety of datasets. The results show that UCA-based hashing outperforms state-of-the-art methods, and has computationally efficient training and encoding processes.
Social Role-Aware Emotion Contagion in Image Social Networks
Yang, Yang (Tsinghua University) | Jia, Jia (Tsinghua University) | Wu, Boya (Tsinghua Univeristy) | Tang, Jie (Tsinghua University)
Psychological theories suggest that emotion represents the state of mind and instinctive responses of one’s cognitive system (Cannon 1927). Emotions are a complex state of feeling that results in physical and psychological changes that influence our behavior. In this paper, we study an interesting problem of emotion contagion in social networks. In particular, by employing an image social network (Flickr) as the basis of our study, we try to unveil how users’ emotional statuses influence each other and how users’ positions in the social network affect their influential strength on emotion. We develop a probabilistic framework to formalize the problem into a role-aware contagion model. The model is able to predict users’ emotional statuses based on their historical emotional statuses and social structures. Experiments on a large Flickr dataset show that the proposed model significantly outperforms (+31% in terms of F1-score) several alternative methods in predicting users’ emotional status. We also discover several intriguing phenomena. For example, the probability that a user feels happy is roughly linear to the number of friends who are also happy; but taking a closer look, the happiness probability is superlinear to the number of happy friends who act as opinion leaders (Page et al. 1999) in the network and sublinear in the number of happy friends who span structural holes (Burt 2001). This offers a new opportunity to understand the underlying mechanism of emotional contagion in online social networks.
California Inc.: Anyone in the market for a slightly used search engine?
Welcome to California Inc., the weekly newsletter of the L.A. Times Business Section. Expect financial markets to face headwinds today after the Federal Reserve reported Friday that U.S. industrial production fell more than expected in March. This is the latest sign that economic growth slowed significantly in the first quarter. On the plus side, though, many economists still forecast a rebound in growth as the year plods ahead. Tax deadline: Monday is the deadline for most Americans to submit their tax returns.
EU wants Google, Microsoft to be more transparent about ads in search results
The European Union's digital chief wants search engines such as Alphabet Inc's Google and Microsoft's Bing to be more transparent about advertising in web search results but ruled out a separate law for web platforms. European Commission vice-president Andrus Ansip, who is overseeing a wide-ranging inquiry into how web platforms conduct their business, said on Friday the EU executive would not take a horizontal approach to regulating online services. "We will take a problem-driven approach," Ansip said. "It's practically impossible to regulate all the platforms with one really good single solution." Related: Do Google's'unprofessional hair' results show it is racist?
Have You Tried Using a 'Nearest Neighbor Search'?
Roughly a year and a half ago, I had the privelage of taking a graduate "Introduction to Machine Learning" course under the tutelage of the fantastic Professor Leslie Kaelbling. While I learned a great deal over the course of the semester, there was one minor point that she made to the class which stuck with me more than I expected it to at the time: before using a really fancy or sophisticated or "in-vogue" machine learning algorithm to solve your problem, try a simple Nearest Neighbor Search first. Let's say I gave you a bunch of data points, each with a location in space and a value, and then asked you to predict the value of a new point in space. Perhaps the values of you data are binary (just s and -s) and you've heard of Support Vector Machines. Should you give that a shot?
Automatic Summary Generation for Scientific Data Charts
Al-Zaidy, Rabah A. (The Pennsylvania State University) | Choudhury, Sagnik Ray (The Pennsylvania State University) | Giles, C. Lee (The Pennsylvania State University)
Scientific charts in the web, whether as images or embedded in digital documents, contain valuable information that is not fully available to information retrieval tools. The information used to describe these charts is typically extracted from the image metadata rather than the information the graphic was initially designed to express. The problem of understanding digital charts found in scholarly documents, and inferring useful textual information from their graphical components is the focus of this study. We present an approach to automatically read the chart data, specifically bar charts, and provide the user with a textual summary of the chart. The proposed method follows a knowledge discovery approach that relies on a versatile graph representation of the chart. This representation is derived from analyzing a chart's original data values, from which useful features are extracted. The data features are in turn used to construct a semantic-graph. To generate a summary, the semantic-graph of the chart is mapped to appropriately crafted protoforms, which are constructs based on fuzzy logic. We verify the effectiveness of our framework by conducting experiments on bar charts extracted from over 1,000 PDF documents. Our preliminary results show that, under certain assumptions, 83% of the produced summaries provide plausible descriptions of the bar charts.
Encoding Lineage in Scholarly Articles
Naim, Sheikh Motahar (University of Texas at El Paso) | Kader, Md Abdul (University of Texas at El Paso) | Boedihardjo, Arnold P. (US Army Corps of Engineers) | Hossain, M. Shahriar (University of Texas at El Paso)
The development of new scientific concepts today is an outcome of the accumulated knowledge built over time. Every scientific domain requires understanding of the trends of the dependencies between its subdomains. Analyses of trends to capture such dependencies using conventional document modeling techniques is a challenging task due to two reasons: (1) conventional vector-space modeling based representation of documents does not realize the history of the content, and (2) neither feature-level nor document-level causality is provided with any digital library metadata or citation network. In this paper, we propose an intuitive temporal representation of a scientific article that encodes inherent historic characteristics of the content. This intuitive representation of each document is then leveraged to discover causal relationships between scientific articles. In addition, we provide a mechanism to explore the lineage of each document in terms of other previously published documents, which illustrates how the theme of the document under analysis evolved over time. Empirical studies reported in the paper show that the proposed technique identifies meaningful causal relationships and discovers meaningful lineage in the scientific literature that could not be discovered through the citation network of the articles.
Automatic Construction of Evaluation Sets and Evaluation of Document Similarity Models in Large Scholarly Retrieval Systems
Krstovski, Kriste (Harvard-Smithsonian Center for Astrophysics) | Smith, David A. (Northeastern University) | Kurtz, Michael J. (Harvard-Smithsonian Center for Astrophysics)
Retrieval systems for scholarly literature offer the ability for the scientific community to search, explore and download scholarly articles across various scientific disciplines. Mostly used by the experts in the particular field, these systems contain user community logs including information on user specific downloaded articles. In this paper we present a novel approach for automatically evaluating document similarity models in large collections of scholarly publications. Unlike typical evaluation settings that use test collections consisting of query documents and human annotated relevance judgments, we use download logs to automatically generate pseudo-relevant set of similar document pairs. More specifically we show that consecutively downloaded document pairs, extracted from a scholarly information retrieval (IR) system, could be utilized as a test collection for evaluating document similarity models. Another novel aspect of our approach lies in the method that we employ for evaluating the performance of the model by comparing the distribution of consecutively downloaded document pairs and random document pairs in log space. Across two families of similarity models, that represent documents in the term vector and topic spaces, we show that our evaluation approach achieves very high correlation with traditional performance metrics such as Mean Average Precision (MAP), while being more efficient to compute.
Creating Content for Google's RankBrain
Google revealed in October that it uses artificial intelligence to help with 15% of search queries. Named RankBrain, the system analyzes vague, ambiguous queries and matches them with the most relevant results. In fact, Google's Greg Corrado told Bloomberg that RankBrain is now the third-highest signal contributing to a search-query result. Google – and similar search-engine services – are getting smarter. As marketers, we no longer can rely solely on traditional digital strategies such as link-building or social-media signaling.