Goto

Collaborating Authors

 Information Retrieval


How to build a search engine - Part 2: Configuring elasticsearch

@machinelearnbot

In this post we will focus on configuring the elasticsearch bit. I have chosen the Wikipedia people dump for the dataset. This is the wiki pages of a subset of people on Wikipedia. This dataset consists of three columns – URI, name, text. As the column names suggest, URI is the actual wiki link to that person's page, name is the person's name.


Gradient Augmented Information Retrieval with Autoencoders and Semantic Hashing

arXiv.org Machine Learning

This paper will explore the use of autoencoders for semantic hashing in the context of Information Retrieval. This paper will summarize how to efficiently train an autoencoder in order to create meaningful and low-dimensional encodings of data. This paper will demonstrate how computing and storing the closest encodings to an input query can help speed up search time and improve the quality of our search results. The novel contributions of this paper involve using the representation of the data learned by an auto-encoder in order to augment our search query in various ways. I present and evaluate the new gradient search augmentation (GSA) approach, as well as the more well-known pseudo-relevance-feedback (PRF) adjustment. I find that GSA helps to improve the performance of the TF-IDF based information retrieval system, and PRF combined with GSA works best overall for the systems compared in this paper.


What real-world problems can AI really solve? An interview with YITU Technology · TechNode

#artificialintelligence

You might have heard about those ATMs that use facial recognition instead of cards and PIN numbers for authentication. You might also have seen on the news a smart security algorithm that helps police identify suspects and cracks criminal cases. Artificial intelligence (AI), the wiz behind these advanced technologies, is permeating our daily lives--everything from financial services to public safety to healthcare and transportation. YITU Technology, one of China's front-running AI startups, has developed solutions that help solve real-world problems. YITU now has the ability to enable accurate facial recognition with a large database of over 1 billion faces in just one second, and their technology has in fact assisted Chinese law enforcement in criminal investigations.


Beyond Keywords and Relevance: A Personalized Ad Retrieval Framework in E-Commerce Sponsored Search

arXiv.org Machine Learning

On most sponsored search platforms, advertisers bid on some keywords for their advertisements (ads). Given a search request, ad retrieval module rewrites the query into bidding keywords, and uses these keywords as keys to select Top N ads through inverted indexes. In this way, an ad will not be retrieved even if queries are related when the advertiser does not bid on corresponding keywords. Moreover, most ad retrieval approaches regard rewriting and ad-selecting as two separated tasks, and focus on boosting relevance between search queries and ads. Recently, in e-commerce sponsored search more and more personalized information has been introduced, such as user profiles, long-time and real-time clicks. Personalized information makes ad retrieval able to employ more elements (e.g. real-time clicks) as search signals and retrieval keys, however it makes ad retrieval more difficult to measure ads retrieved through different signals. To address these problems, we propose a novel ad retrieval framework beyond keywords and relevance in e-commerce sponsored search. Firstly, we employ historical ad click data to initialize a hierarchical network representing signals, keys and ads, in which personalized information is introduced. Then we train a model on top of the hierarchical network by learning the weights of edges. Finally we select the best edges according to the model, boosting RPM/CTR. Experimental results on our e-commerce platform demonstrate that our ad retrieval framework achieves good performance.


Bias already exists in search engine results, and it's only going to get worse

#artificialintelligence

The internet might seem like a level playing field, but it isn't. Safiya Umoja Noble came face to face with that fact one day when she used Google's search engine to look for subjects her nieces might find interesting. She entered the term "black girls" and came back with pages dominated by pornography. Noble, a USC Annenberg communications professor, was horrified but not surprised. For years she has been arguing that the values of the web reflect its builders--mostly white, Western men--and do not represent minorities and women.


Building a Content-Based Search Engine I: Quantifying Similarity - deep ideas

#artificialintelligence

The explosion of user-generated content on the internet during the last decades has left the world of querying multimedia data with unprecedented challenges. There is a demand for this data to be processed and indexed in order to make it available for different types of queries, whilst ensuring acceptable response times. An arguably important task is the retrieval of multimedia objects (e.g. We define two multimedia objects to be visually similar if they depict contents that "look similar" to humans. So far, this task has gained comparatively little research recognition.


The Low Hanging SEO Fruit You're Missing Out On

#artificialintelligence

This article covers several different methods to help you identify and seize opportunities to promote your site quickly and efficiently. I believe that, when it comes down to it, website promotion works best when based on the Pareto principle – that is, 20% of the pages create 80% of the traffic. This article will help you identify, optimize, and promote your best pages, the 20% that generate most of your traffic. As you probably know already, more and more business owners understand the importance of their website's ranking in Google's search results. As it is the go-to search engine for almost everyone and everything, Google has become the effective reality of the business world.



Google launched its own job search engine – here's how it works

The Independent - Tech

Google launched its own job search engine -- here's how it works The tech giant recently launched its own job search feature, Google for Jobs. As Business Insider's Matt Weinberger reports, the new feature employs machine learning-trained algorithms to sort and organise job listings from a range of employment sites including LinkedIn, Monster and Glassdoor. So if you decide to find your next gig on Google, you'll have a streamlined place to search and AI technology on your side. Here's 13 tips on how to get started using Google for Jobs: Follow Business Insider UK on Twitter. The Independent's bitcoin group on Facebook is the best place to follow the latest discussions and developments in cryptocurrency.


Artificial Intelligence and the Future of Search Engines

#artificialintelligence

It was not long ago that Artificial Intelligence (AI) was only in the realm of science fiction. Today, it has become a reality and is only growing more prominent in many different industries every day. This includes the internet as AI in search engine technology has been around for a few years. The algorithms used to rank pages have been affected considerably by AI already and that trend will continue into the foreseeable future. Currently, Google's RankBrain, an AI process used help set search engine rankings, is having a major impact which is only expected to expand.