Goto

Collaborating Authors

 Information Retrieval


A Clustering-Based Combinatorial Approach to Unsupervised Matching of Product Titles

arXiv.org Machine Learning

The constant growth of the e-commerce industry has rendered the problem of product retrieval particularly important. As more enterprises move their activities on the Web, the volume and the diversity of the product-related information increase quickly. These factors make it difficult for the users to identify and compare the features of their desired products. Recent studies proved that the standard similarity metrics cannot effectively identify identical products, since similar titles often refer to different products and vice-versa. Other studies employed external data sources (search engines) to enrich the titles; these solutions are rather impractical mainly because the external data fetching is slow. In this paper we introduce UPM, an unsupervised algorithm for matching products by their titles. UPM is independent of any external sources, since it analyzes the titles and extracts combinations of words out of them. These combinations are evaluated according to several criteria, and the most appropriate of them constitutes the cluster where a product is classified into. UPM is also parameter-free, it avoids product pairwise comparisons, and includes a post-processing verification stage which corrects the erroneous matches. The experimental evaluation of UPM demonstrated its superiority against the state-of-the-art approaches in terms of both efficiency and effectiveness.


Voyageur: An Experiential Travel Search Engine

arXiv.org Artificial Intelligence

We describe Voyageur, which is an application of experiential search to the domain of travel. Unlike traditional search engines for online services, experiential search focuses on the experiential aspects of the service under consideration. In particular, Voyageur needs to handle queries for subjective aspects of the service (e.g., quiet hotel, friendly staff) and combine these with objective attributes, such as price and location. Voyageur also highlights interesting facts and tips about the services the user is considering to provide them with further insights into their choices.


On Application of Learning to Rank for E-Commerce Search

arXiv.org Machine Learning

E-Commerce (E-Com) search is an emerging important new application of information retrieval. Learning to Rank (LETOR) is a general effective strategy for optimizing search engines, and is thus also a key technology for E-Com search. While the use of LETOR for web search has been well studied, its use for E-Com search has not yet been well explored. In this paper, we discuss the practical challenges in applying learning to rank methods to E-Com search, including the challenges in feature representation, obtaining reliable relevance judgments, and optimally exploiting multiple user feedback signals such as click rates, add-to-cart ratios, order rates, and revenue. We study these new challenges using experiments on industry data sets and report several interesting findings that can provide guidance on how to optimally apply LETOR to E-Com search: First, popularity-based features defined solely on product items are very useful and LETOR methods were able to effectively optimize their combination with relevance-based features. Second, query attribute sparsity raises challenges for LETOR, and selecting features to reduce/avoid sparsity is beneficial. Third, while crowdsourcing is often useful for obtaining relevance judgments for Web search, it does not work as well for E-Com search due to difficulty in eliciting sufficiently fine grained relevance judgments. Finally, among the multiple feedback signals, the order rate is found to be the most robust training objective, followed by click rate, while add-to-cart ratio seems least robust, suggesting that an effective practical strategy may be to initially use click rates for training and gradually shift to using order rates as they become available.


Does The Meta Description Tag Affect SEO & Search Engine Rankings?

#artificialintelligence

There's a lot of confusion when it comes to meta descriptions and SEO. Do they affect search engine rankings? Is it worth spending the time to write a good meta description? Well, in theory, meta descriptions do not affect SEO. This is an official statement from Google, released in 2009. However, since meta descriptions show in the search engine results, they can affect CTRs (click through rates), which are linked to SEO & rankings. So, in practice, meta descriptions might have an impact on SEO.


Which Social Network Drives the Most Traffic to Your Website? [POLL] - Search Engine Journal

#artificialintelligence

Facebook typically drives around 60 percent of all the traffic Search Engine Journal get from social media networks. This isn't too surprising, considering how big Facebook is with over 2.2 billion monthly active users. Even if social media won't directly help your organic search rankings, posting engaging content that attracts lots of traffic, shares, likes, and comments all helps you increase your reach, visibility, and linking potential. So which social networks have the highest potential to send lots of traffic to your website? We asked our Twitter community what trends they're seeing.


Ranking in Genealogy: Search Results Fusion at Ancestry

arXiv.org Machine Learning

Genealogy research is the study of family history using available resources such as historical records. Ancestry provides its customers with one of the world's largest online genealogical index with billions of records from a wide range of sources, including vital records such as birth and death certificates, census records, court and probate records among many others. Search at Ancestry aims to return relevant records from various record types, allowing our subscribers to build their family trees, research their family history, and make meaningful discoveries about their ancestors from diverse perspectives. In a modern search engine designed for genealogical study, the appropriate ranking of search results to provide highly relevant information represents a daunting challenge. In particular, the disparity in historical records makes it inherently difficult to score records in an equitable fashion. Herein, we provide an overview of our solutions to overcome such record disparity problems in the Ancestry search engine. Specifically, we introduce customized coordinate ascent (customized CA) to speed up ranking within a specific record type. We then propose stochastic search (SS) that linearly combines ranked results federated across contents from various record types. Furthermore, we propose a novel information retrieval metric, normalized cumulative entropy (NCE), to measure the diversity of results. We demonstrate the effectiveness of these two algorithms in terms of relevance (by NDCG) and diversity (by NCE) if applicable in the offline experiments using real customer data at Ancestry.


Product Search Engines - vs - Product Recommendation Engines

#artificialintelligence

However similar they may seem in terms of underlying technology, conventional search engines are quite different from product recommendation engines. Behaviour is typically at the center of a recommender system, while it is an added dimension for the search engine.


Google's John Mueller on Ranking for Featured Snippets - Search Engine Journal

#artificialintelligence

Someone asked John Mueller in a Webmaster Hangout about Schema structured data and ranking for featured snippets. Structured Data is useful for communicating a deep amount of precise data. John Mueller answered the question by describing what it takes to make it easier for Google to use your page for featured snippets. The question asked was about the use of structured data for ranking in featured snippets. It was also about showing up for voice search via the Google Assistant.


Empowering Elasticsearch with Exact and Fast $r$-Neighbor Search in Hamming Space

arXiv.org Machine Learning

A growing interest has been witnessed recently in building nearest neighbor search solutions within Elasticsearch--one of the most popular full-text search engines. In this paper, we focus specifically on Hamming space nearest neighbor search using Elasticsearch. By combining three techniques: bit operation, substring filtering and data preprocessing with permutation, we develop a novel approach called FENSHSES (Fast Exact Neighbor Search in Hamming Space on Elasticsearch), which achieves dramatic speed-ups over the existing term match baseline. This will empower Elasticsearch with the capability of fast information retrieval even when documents (e.g., texts, images and sounds) are represented with binary codes--a common practice in nowadays semantic representation learning.


LDA for Text Summarization and Topic Detection - DZone AI

#artificialintelligence

Machine learning clustering techniques are not the only way to extract topics from a text data set. Text mining literature has proposed a number of statistical models, known as probabilistic topic models, to detect topics from an unlabeled set of documents. One of the most popular models is the latent Dirichlet allocation (LDA) algorithm developed by Blei, Ng, and Jordan [i]. LDA is a generative unsupervised probabilistic algorithm that isolates the top K topics in a data set as described by the most relevant N keywords. In other words, the documents in the data set are represented as random mixtures of latent topics, where each topic is characterized by a Dirichlet distribution over a fixed vocabulary.