Goto

Collaborating Authors

 Information Retrieval


Why artificial intelligence still needs a human touch

#artificialintelligence

Multiple examples have recently gone viral, each demonstrating the shortcomings of Google's search methods and the ability of the AI utilised to generate factual information. Ask Google if Obama is planning a coup, and you'll get the answer that he might be. Ask Google if women are evil and Google delivers the answer that all women have a "degree of prostitute in them". Ask Google if all republicans are fascists, and Google produces an answer that includes the suggestion they're all Nazis. Most of these answers are derisory and inherently damage Google's brand. More dangerous, however, is the potential to warp users' perception of the world by providing incorrect and unfortunately inflammatory information. By giving patently false narratives top billing on Google lends these stories unwarranted credibility. Google's utilisation of AI is helping to usher in what has been called the "post-truth" era.


Comparison Based Nearest Neighbor Search

arXiv.org Machine Learning

We consider machine learning in a comparison-based setting where we are given a set of points in a metric space, but we have no access to the actual distances between the points. Instead, we can only ask an oracle whether the distance between two points $i$ and $j$ is smaller than the distance between the points $i$ and $k$. We are concerned with data structures and algorithms to find nearest neighbors based on such comparisons. We focus on a simple yet effective algorithm that recursively splits the space by first selecting two random pivot points and then assigning all other points to the closer of the two (comparison tree). We prove that if the metric space satisfies certain expansion conditions, then with high probability the height of the comparison tree is logarithmic in the number of points, leading to efficient search performance. We also provide an upper bound for the failure probability to return the true nearest neighbor. Experiments show that the comparison tree is competitive with algorithms that have access to the actual distance values, and needs less triplet comparisons than other competitors.


Probabilistic Search for Structured Data via Probabilistic Programming and Nonparametric Bayes

arXiv.org Machine Learning

Databases are widespread, yet extracting relevant data can be difficult. Without substantial domain knowledge, multivariate search queries often return sparse or uninformative results. This paper introduces an approach for searching structured data based on probabilistic programming and nonparametric Bayes. Users specify queries in a probabilistic language that combines standard SQL database search operators with an information theoretic ranking function called predictive relevance. Predictive relevance can be calculated by a fast sparse matrix algorithm based on posterior samples from CrossCat, a nonparametric Bayesian model for high-dimensional, heterogeneously-typed data tables. The result is a flexible search technique that applies to a broad class of information retrieval problems, which we integrate into BayesDB, a probabilistic programming platform for probabilistic data analysis. This paper demonstrates applications to databases of US colleges, global macroeconomic indicators of public health, and classic cars. We found that human evaluators often prefer the results from probabilistic search to results from a standard baseline.


26 Experts On How AI Will Change The Way We Do SEO

#artificialintelligence

Things change pretty much on a daily basis in the world of SEO. Since the announcement of Google's AI machine learning algorithm โ€“ RankBrain โ€“ in 2015, one of the most discussed topics in SEO galleries is: With Google admitting RankBrain being one of the top three ranking factors, these discussions have become even more worthwhile. In past 3-4 months, we also saw a spike in the number of SERPed members asking the same question. And, multiple posts claiming 2017 as the year of AI and Voice Search, we think it is the right time to dive deeper to understand more about it. To get more clarity on this topic, we decided to go straight to the big guns and find out what they think about it. The responses from each expert are compiled below. Fasten your seat belts and get ready for an awesome ride. Albert Mora is the CEO and co-founder of Seolution, an SEO agency for Shopify e-commerce sites. He has been doing SEO from 1997 and has around 20 years of experience. Follow Albert on Twitter here. Since the beginning of the Internet, artificial intelligence has played a relevant role in the operation of search engines. Logically, the algorithms have been evolving, but the fundamental underlying principle remains the same: search engines want to deliver quality search results to the users. For this reason, if you want a long term sustainable SEO results, you must think about the users first, not about the search engines. Alex has more than 15 years of experience in Digital Marketing, and he is working online since 2002.


The Death of Organic Search (As We Know It) - Search Engine Journal

#artificialintelligence

It would be easier to count all the stars in the night sky than the number of articles written about the death of SEO. I've never written one personally but I was having a discussion with the author of a great piece here on Search Engine Journal on AI and its impact on search and the question came up: Between machine learning and the limited space available for organic search, is it on its death spiral? The most interesting thing about this question may not be the answer but the journey in understanding the question itself, as it's therein that we understand the strategies that will make it either true or false. Between machine learning, the limited space available for organic search, and the growth of both voice search and personal assistants, is it on its death spiral? To explore this question, we're going to look at each of these three areas individually, what they mean together, and finally (and what you likely most want to know), what you need to do about it.


Twitter found to block certain words in search engine

Daily Mail - Science & tech

Twitter has quietly started blocking certain words on the platform's built-in search engine. Words such as'porn', 'nsfw', 'sex' and similar terms will no longer appear when searched under'Latest' tab โ€“ but, racial slurs and the word'jihad' have not been removed. Although Twitter has blocked these words from being found in the Latest tab, users can still find some of the'forbidden' terms by searching in the'Top' tab. Twitter has quietly started blocking certain words on the platform's built-in search engine. Words such as'porn', 'nsfw', 'sex' and similar terms will no longer appear when searched under'Latest' tab Twitter says it'prohibits the promotion of hate content, sensitive topics, and violence globally.' But this policy does not apply to news and information that calls attention to hate, sensitive topics, or violence, but does not advocate for it.


How Zocdoc's New Machine Learning Search Engine Makes Medicine More Human

#artificialintelligence

And, wait, are internists doctors who specialize in internal medicine - or just their interns? Some well-versed readers may indeed know the answer to all three of these questions (for the record: ophthalmologists are the ones who can perform eye surgery, psychiatrists can prescribe medications, and internists are your good old-fashioned general medicine practitioners). But the complexities of medical jargon can make an already complicated U.S. health system even more convoluted for millions of people seeking care. Click here to subscribe to Brainstorm Health Daily, our brand new newsletter about health innovations. That's why Zocdoc, the online doctor-locating and medical appointments platform, launched a new feature today on desktop and mobile devices that it dubs the "Patient-Powered Search." The firm describes this new engine as a "more intuitive search experience, built specifically to bridge the gap between healthcare industry and human speak."


A Survey of Available Corpora for Building Data-Driven Dialogue Systems

arXiv.org Artificial Intelligence

During the past decade, several areas of speech and language understanding have witnessed substantial breakthroughs from the use of data-driven models. In the area of dialogue systems, the trend is less obvious, and most practical systems are still built through significant engineering and expert knowledge. Nevertheless, several recent results suggest that data-driven approaches are feasible and quite promising. To facilitate research in this area, we have carried out a wide survey of publicly available datasets suitable for data-driven learning of dialogue systems. We discuss important characteristics of these datasets, how they can be used to learn diverse dialogue strategies, and their other potential uses. We also examine methods for transfer learning between datasets and the use of external knowledge. Finally, we discuss appropriate choice of evaluation metrics for the learning objective.


How Search Engines Will Become More Integrated In The Near Future

Forbes - Tech

How would search engine evolve in the next 10 years? When we talk about search engines today, search boxes and search results come to our minds. What might future search engines look like? But we would be happy to have a much more powerful search engine that we may see, hear and even feel in different scenarios, different products or different interfaces. Firstly, deeper understanding of user's intent, deeper understanding of content and more accurate matching of intent and content would empower the search engine. The understanding of user's intent will depend not only on a single query, but also on more comprehensive search contexts, including query sessions, time, location, device, and the user's personalization features.


A Visual Search Engine for the Entire Planet

The Atlantic - Technology

At this moment in history, there are more satellites photographing Earth from orbit than just about anyone knows what to do with. Planet, Inc., has more than 150 orbiting cameras, each the size of a shoebox. And more startups are planning to launch their own. What should we do with all that imagery? How can we search it and process it?