Goto

Collaborating Authors

 Information Retrieval


How to build a search engine: Part 3

@machinelearnbot

Assuming the dataset is named "people_wiki.csv", Executing this script will result in steaming logs which is ultimately leading to the data getting indexed in elasticsearch. That's how easy it is! Let's spend the next few lines on what actually happened. We declare our elasticsearch object configured on our local machine. Once that object is initialized we will use it to index all of our data.


A General Framework for Robust Interactive Learning

Neural Information Processing Systems

We propose a general framework for interactively learning models, such as (binary or non-binary) classifiers, orderings/rankings of items, or clusterings of data points. Our framework is based on a generalization of Angluin's equivalence query model and Littlestone's online learning model: in each iteration, the algorithm proposes a model, and the user either accepts it or reveals a specific mistake in the proposal. The feedback is correct only with probability p > 1/2 (and adversarially incorrect with probability 1 - p), i.e., the algorithm must be able to learn in the presence of arbitrary noise. The algorithm's goal is to learn the ground truth model using few iterations. Our general framework is based on a graph representation of the models and user feedback. To be able to learn efficiently, it is sufficient that there be a graph G whose nodes are the models, and (weighted) edges capture the user feedback, with the property that if s, s* are the proposed and target models, respectively, then any (correct) user feedback s' must lie on a shortest s-s* path in G. Under this one assumption, there is a natural algorithm, reminiscent of the Multiplicative Weights Update algorithm, which will efficiently learn s* even in the presence of noise in the user's feedback. From this general result, we rederive with barely any extra effort classic results on learning of classifiers and a recent result on interactive clustering; in addition, we easily obtain new interactive learning algorithms for ordering/ranking.


Query Complexity of Clustering with Side Information

Neural Information Processing Systems

Suppose, we are given a set of $n$ elements to be clustered into $k$ (unknown) clusters, and an oracle/expert labeler that can interactively answer pair-wise queries of the form, ``do two elements $u$ and $v$ belong to the same cluster?''. The goal is to recover the optimum clustering by asking the minimum number of queries. In this paper, we provide a rigorous theoretical study of this basic problem of query complexity of interactive clustering, and give strong information theoretic lower bounds, as well as nearly matching upper bounds. Most clustering problems come with a similarity matrix, which is used by an automated process to cluster similar points together. To improve accuracy of clustering, a fruitful approach in recent years has been to ask a domain expert or crowd to obtain labeled data interactively. Many heuristics have been proposed, and all of these use a similarity function to come up with a querying strategy. Even so, there is a lack systematic theoretical study. Our main contribution in this paper is to show the dramatic power of side information aka similarity matrix on reducing the query complexity of clustering. A similarity matrix represents noisy pair-wise relationships such as one computed by some function on attributes of the elements. A natural noisy model is where similarity values are drawn independently from some arbitrary probability distribution $f_+$ when the underlying pair of elements belong to the same cluster, and from some $f_-$ otherwise. We show that given such a similarity matrix, the query complexity reduces drastically from $\Theta(nk)$ (no similarity matrix) to $O(\frac{k^2\log{n}}{\cH^2(f_+\|f_-)})$ where $\cH^2$ denotes the squared Hellinger divergence. Moreover, this is also information-theoretic optimal within an $O(\log{n})$ factor. Our algorithms are all efficient, and parameter free, i.e., they work without any knowledge of $k, f_+$ and $f_-$, and only depend logarithmically with $n$.


A Game Theoretic Analysis of the Adversarial Retrieval Setting

Journal of Artificial Intelligence Research

The main goal of search engines is ad hoc retrieval: ranking documents in a corpus by their relevance to the information need expressed by a query. The Probability Ranking Principle (PRP) --- ranking the documents by their relevance probabilities --- is the theoretical foundation of most existing ad hoc document retrieval methods. A key observation that motivates our work is that the PRP does not account for potential post-ranking effects; specifically, changes to documents that result from a given ranking. Yet, in adversarial retrieval settings such as the Web, authors may consistently try to promote their documents in rankings by changing them. We prove that, indeed, the PRP can be sub-optimal in adversarial retrieval settings. We do so by presenting a novel game theoretic analysis of the adversarial setting. The analysis is performed for different types of documents (single-topic and multi-topic) and is based on different assumptions about the writing qualities of documents' authors. We show that in some cases, introducing randomization into the document ranking function yields an overall user utility that transcends that of applying the PRP.


Solving Google Cache Issue Multilingual Search Engine Optimization

#artificialintelligence

After Google Penguin 6 ( aka Penguin 3.0) update, some webmasters asked me why Google does not cache their updated pages. Before talking about Google Cache issue let's have an overview on the differences between Google index and Google Cache. An "indexed" webpage is a site that has been crawled by a search engine spider and filed away in their index for later use, and a "cached" page is one that may show up in search results. Submitting pages listed in the SERPs is a multi-step task and process. First, the crawlers need to access your website and scan the page.


Search engine optimization tips and tricks for 2018

@machinelearnbot

Search engine optimization tips and tricks for 2018: SEO stands for SEARCH ENGINE OPTIMIZATION. It is the technique to appear your website on top of Search Engine's results. It is the part of the DIGITAL MARKETING. In another word you can say it is the process to drive the huge traffic to your website. Traffic should be organic, editorial, natural.


A Gentle Introduction to Applied Machine Learning as a Search Problem - Machine Learning Mastery

#artificialintelligence

Applied machine learning is challenging because the designing of a perfect learning system for a given problem is intractable. There is no best training data or best algorithm for your problem, only the best that you can discover. The application of machine learning is best thought of as search problem for the best mapping of inputs to outputs given the knowledge and resources available to you for a given project. In this post, you will discover the conceptualization of applied machine learning as a search problem. A Gentle Introduction to Applied Machine Learning as a Search Problem Photo by tonko43, some rights reserved.


Ella brings smart searching to home security cameras

PCWorld

Her name is Ella and she promises to bring smart searching to your home security cameras. Ella is an AI-powered search engine, developed by IC Realtime, that augments residential and commercial surveillance systems with natural language search capabilities. The average continuously recording security camera captures less than two minutes of noteworthy footage in a 24-hour period, IC Realtime CEO Matt Sailor said during an embargoed briefing last week. Scrubbing through a day's worth of video just to retrieve that data is a tedious time-suck. Even with the time- and date-sorting parameters typically offered by most consumer cameras, reviewing video can be arduous.


Multilingual Topic Models

arXiv.org Machine Learning

Scientific publications have evolved several features for mitigating vocabulary mismatch when indexing, retrieving, and computing similarity between articles. These mitigation strategies range from simply focusing on high-value article sections, such as titles and abstracts, to assigning keywords, often from controlled vocabularies, either manually or through automatic annotation. Various document representation schemes possess different cost-benefit tradeoffs. In this paper, we propose to model different representations of the same article as translations of each other, all generated from a common latent representation in a multilingual topic model. We start with a methodological overview on latent variable models for parallel document representations that could be used across many information science tasks. We then show how solving the inference problem of mapping diverse representations into a shared topic space allows us to evaluate representations based on how topically similar they are to the original article. In addition, our proposed approach provides means to discover where different concept vocabularies require improvement.


microsoft-bing-reddit-search-engine-partnership-ama-subreddit-content

#artificialintelligence

Microsoft today announced a partnership with Reddit to surface the social news site's content high up in Bing search results, so that searching for specific Reddit communities, pages, and info will spit back information gleaned from subreddits, AMAs, and other threads in a dedicated section. The news, part of a suite of new artificial intelligence-powered features underlying Bing's search function, works by way of what Microsoft calls intelligent search. The company says intelligent search uses natural language processing to pair Reddit content with the appropriate search terms. The Reddit integration works in three ways. First, by typing in the name of a subreddit on Bing, you'll get a live snapchat of that subreddit's top threads displayed in the search results with links to the individual pages.