Goto

Collaborating Authors

 Information Retrieval


PageRank algorithm for Directed Hypergraph

arXiv.org Machine Learning

With the huge amount of information inflowing the World Wide Web every second, it becomes more difficult and more difficult to retrieve information from the Web. This explains why the existence of a search engine is as important as the existence of the web itself. Since the appearance of the web, there has been a fundamental talk in the web research communit y to develop the rapid, effective, and precise search engines. This paper will be chiefly discussing about the most common search engine nowadays which is Google. The mathematical theory behind the Google search engine is the PageRank algorithm, which was presented by Sergey Brin and Lawrence Page [1].


Enterprise Search Software & Semantic Search Engine

#artificialintelligence

Welcome to the era of Big Data where data-driven insights have the power to transform your business. You're about to discover the solution: a powerful, innovative and adaptive platform power packed with every feature you need for Search, Discovery & Analytics of your data. We have named it 3RDi "Third Eye". It's the semantic search engine your enterprise needs to help you take action, boost revenues and cut costs! Powered by NLP and semantic search, it is designed for multidimensional information analysis and easy search relevancy management.


Building an Image Hashing Search Engine with VP-Trees and OpenCV - PyImageSearch

#artificialintelligence

In this tutorial, you will learn how to build a scalable image hashing search engine using OpenCV, Python, and VP-Trees. Back in 2017, I wrote a tutorial on image hashing with OpenCV and Python (which is required reading for this tutorial). That guide showed you how to find identical/duplicate images in a given dataset. However, there was a scalability problem with that original tutorial -- namely that it did not scale! To find near-duplicate images, our original image hashing method would require us to perform a linear search, comparing the query hash to each individual image hash in our dataset. In a practical, real-world application that's far too slow -- we need to find a way to reduce that search to sub-linear time complexity. But how can we reduce search time so dramatically?


Google Says it Doesn't Require Fixing Structured Data Warnings - Search Engine Journal

#artificialintelligence

In a Webmaster Hangout, an eCommerce publisher complained about structured data warnings regarding data fields that are inappropriate to their product. They refused to create fake information to get a passing score. John Mueller responded that there's a difference between warnings and errors. The person asking the question sold custom hand made products. They did not have a global identifier.


Real-world Conversational AI for Hotel Bookings

arXiv.org Machine Learning

Hussein Fazal SnapTravel Toronto, Canada hussein@snaptravel.com Abstract --In this paper, we present a real-world conversational AI system to search for and book hotels through text messaging. Our architecture consists of a frame-based dialogue management system, which calls machine learning models for intent classification, named entity recognition, and information retrieval subtasks. Our chatbot has been deployed on a commercial scale, handling tens of thousands of hotel searches every day. We describe the various opportunities and challenges of developing a chatbot in the travel industry. Index T erms--conversational AI, task-oriented chatbot, named entity recognition, information retrieval I. I NTRODUCTION Task-oriented chatbots have recently been applied to many areas in e-commerce.


Nearest Neighbor Search-Based Bitwise Source Separation Using Discriminant Winner-Take-All Hashing

arXiv.org Artificial Intelligence

We propose an iteration-free source separation algorithm based on Winner-Take-All (WTA) hash codes, which is a faster, yet accurate alternative to a complex machine learning model for single-channel source separation in a resource-constrained environment. We first generate random permutations with WTA hashing to encode the shape of the multidimensional audio spectrum to a reduced bitstring representation. A nearest neighbor search on the hash codes of an incoming noisy spectrum as the query string results in the closest matches among the hashed mixture spectra. Using the indices of the matching frames, we obtain the corresponding ideal binary mask vectors for denoising. Since both the training data and the search operation are bitwise, the procedure can be done efficiently in hardware implementations. Experimental results show that the WTA hash codes are discriminant and provide an affordable dictionary search mechanism that leads to a competent performance compared to a comprehensive model and oracle masking.


WordPress 3 Search Engine Optimization - Programmer Books

#artificialintelligence

WordPress is a powerful platform for creating feature-rich and attractive websites and blogs; but with a little extra tweaking and effort your WordPress site can dominate the search engines and bring thousands of new customers to your blog or business. WordPress3.0 Search Engine Optimization will show you the secrets that professional SEO companies use to take websites to the top of search results and proliferate their business. You'll be able to take your WordPress blog/site to the next level, as well as brush aside even the stiffest competition with this book in hand. We'll begin with a typical WordPress installation and with a variety of simple techniques, turn it into a powerful website that search engines will reward with high rankings. We'll go further: with advanced plug-ins we'll connect your WordPress site to popular social media sites and expand the reach of your site to bring more visitors.


Automatic Language Identification in Texts: A Survey

Journal of Artificial Intelligence Research

Language identification ("LI") is the problem of determining the natural language that a document or part thereof is written in. Automatic LI has been extensively researched for over fifty years. Today, LI is a key part of many text processing pipelines, as text processing techniques generally assume that the language of the input text is known. Research in this area has recently been especially active. This article provides a brief history of LI research, and an extensive survey of the features and methods used in the LI literature. We describe the features and methods using a unified notation, to make the relationships between methods clearer. We discuss evaluation methods, applications of LI, as well as off-the-shelf LI systems that do not require training by the end user. Finally, we identify open issues, survey the work to date on each issue, and propose future directions for research in LI.


Google's John Mueller Answers if Linking Out Good for SEO - Search Engine Journal

#artificialintelligence

Google launched a new video series that answers a single question. The first episode was about links but in my opinion it did not adequately answer the question. "Does linking to other websites help or hurt SEO?" The SEO community has thought of outbound links as ranking signals since at least 2002. I hope to show you how and why outbound links for SEO was invented.


As Search Engines Increasingly Turn To AI They Are Harming Search

#artificialintelligence

For more than half a century our digital search engines have relied upon the humble keyword. Yet over the past few years, search engines of all kinds have increasingly turned to deep learning-powered categorization and recommendation algorithms to augment and slowly replace the traditional keyword search. Behavioral and interest-based personalization has further eroded the impact of keyword searches, meaning that if ten people all search for the same thing, they may all get different results. As search engines depreciate traditional raw "search" in favor of AI-assisted navigation, the concept of informational access is being harmed and our digital world is being redefined by the limitations of today's AI. At first glance, the evolution of search from simple TF-IDF keyword queries into today's AI-powered personalized digital navigation is a positive step towards making the digital world more accessible to the general public.