Information Retrieval
Nearest Neighbor Search Under Uncertainty
Mason, Blake, Tripathy, Ardhendu, Nowak, Robert
Nearest Neighbor Search (NNS) is a central task in knowledge representation, learning, and reasoning. There is vast literature on efficient algorithms for constructing data structures and performing exact and approximate NNS. This paper studies NNS under Uncertainty (NNSU). Specifically, consider the setting in which an NNS algorithm has access only to a stochastic distance oracle that provides a noisy, unbiased estimate of the distance between any pair of points, rather than the exact distance. This models many situations of practical importance, including NNS based on human similarity judgements, physical measurements, or fast, randomized approximations to exact distances. A naive approach to NNSU could employ any standard NNS algorithm and repeatedly query and average results from the stochastic oracle (to reduce noise) whenever it needs a pairwise distance. The problem is that a sufficient number of repeated queries is unknown in advance; e.g., a point maybe distant from all but one other point (crude distance estimates suffice) or it may be close to a large number of other points (accurate estimates are necessary). This paper shows how ideas from cover trees and multi-armed bandits can be leveraged to develop an NNSU algorithm that has optimal dependence on the dataset size and the (unknown)geometry of the dataset.
The AI Index 2021 Annual Report
Zhang, Daniel, Mishra, Saurabh, Brynjolfsson, Erik, Etchemendy, John, Ganguli, Deep, Grosz, Barbara, Lyons, Terah, Manyika, James, Niebles, Juan Carlos, Sellitto, Michael, Shoham, Yoav, Clark, Jack, Perrault, Raymond
Welcome to the fourth edition of the AI Index Report. This year we significantly expanded the amount of data available in the report, worked with a broader set of external organizations to calibrate our data, and deepened our connections with the Stanford Institute for Human-Centered Artificial Intelligence (HAI). The AI Index Report tracks, collates, distills, and visualizes data related to artificial intelligence. Its mission is to provide unbiased, rigorously vetted, and globally sourced data for policymakers, researchers, executives, journalists, and the general public to develop intuitions about the complex field of AI. The report aims to be the most credible and authoritative source for data and insights about AI in the world.
Privacy-First Browser Brave Is Launching a Search Engine
Google's grip on the web has never been stronger. Its Chrome web browser has almost 70 percent of the market and its search engine a whopping 92 percent share. This story originally appeared on WIRED UK. But Google's dominance is being challenged. Regulators are questioning its monopoly position and claim the company has used anticompetitive tactics to strengthen its dominance.
The RLR-Tree: A Reinforcement Learning Based R-Tree for Spatial Data
Gu, Tu, Feng, Kaiyu, Cong, Gao, Long, Cheng, Wang, Zheng, Wang, Sheng
Despite the success of these learned indices in improving the performance Learned indices have been proposed to replace classic index structures of some types of queries, they still have various limitations, like B-Tree with machine learning (ML) models. They require e.g., they can only handle spatial point objects and limited types to replace both the indices and query processing algorithms currently of spatial queries, some only return approximate query results, deployed by the databases, and such a radical departure is and they either cannot handle updates or need a periodic rebuild likely to encounter challenges and obstacles. In contrast, we propose to retain high query efficiency (Detailed discussions are in Section a fundamentally different way of using ML techniques to 2). These limitations, together with the requirement that the improve on the query performance of the classic R-Tree without learned indices need a replacement of the index structures and the need of changing its structure or query processing algorithms.
The 3 Unexpected Benefits of Search Strategy
Let's be honest, you've probably heard a thousand times just how important search engine optimization (SEO) is for your business. If you want to gain higher page rankings on search engines like Google and drive more targeted traffic to your site, a winning search strategy is a must. Well, it turns out that there are a few additional unexpected benefits to SEO that should give you all the more reason to make it a cornerstone of your marketing strategy. No matter how laggy or slow-loading it was, people used to stick around. The attention span of users has shrunk to six seconds on average.
BERT based patent novelty search by training claims to their own description
Freunek, Michael, Bodmer, Andrรฉ
In this paper we present a method to concatenate patent claims to their own description. By applying this method, BERT trains suitable descriptions for claims. Such a trained BERT (claim-to-description- BERT) could be able to identify novelty relevant descriptions for patents. In addition, we introduce a new scoring scheme, relevance scoring or novelty scoring, to process the output of BERT in a meaningful way. We tested the method on patent applications by training BERT on the first claims of patents and corresponding descriptions. BERT's output has been processed according to the relevance score and the results compared with the cited X documents in the search reports. The test showed that BERT has scored some of the cited X documents as highly relevant.
Brave Search is a privacy-first search engine
Browser privacy is a big deal, as Google and other companies use your search data to serve you ads while you surf the web. While most users accept that tradeoff, others who believe strongly in maintaining their own data privacy. If you're one of these, Brave Software can help. On Wednesday the company said it's launching a search engine to compete with Google and Bing, with privacy as its first priority. Brave is buying Tailcat, an open search engine, and will add it to what it's calling Brave Search, a forthcoming search engine.
Brave is developing its own privacy-focused search engine
Privacy-focused browser Brave is working on its own search engine. It has bought Tailcat, an open-source engine created by a team who worked on the defunct anti-tracking browser and search engine Cliqz, to power Brave Search. The company will allow others to use Brave Search tech to build their own search engines. Brave says the search engine will provide an alternative to Google Search and Chrome. It's developing Brave Search using the same principles as its browser, which now has more than 25 million monthly active users.
Data Augmentation for Abstractive Query-Focused Multi-Document Summarization
Pasunuru, Ramakanth, Celikyilmaz, Asli, Galley, Michel, Xiong, Chenyan, Zhang, Yizhe, Bansal, Mohit, Gao, Jianfeng
The progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training datasets, which we construct using two data augmentation methods: (1) transferring the commonly used single-document CNN/Daily Mail summarization dataset to create the QMDSCNN dataset, and (2) mining search-query logs to create the QMDSIR dataset. These two datasets have complementary properties, i.e., QMDSCNN has real summaries but queries are simulated, while QMDSIR has real queries but simulated summaries. To cover both these real summary and query aspects, we build abstractive end-to-end neural network models on the combined datasets that yield new state-of-the-art transfer results on DUC datasets. We also introduce new hierarchical encoders that enable a more efficient encoding of the query together with multiple documents. Empirical results demonstrate that our data augmentation and encoding methods outperform baseline models on automatic metrics, as well as on human evaluations along multiple attributes.
Making Enterprise Search Personal - Coruzant Technologies
Knowledge management providers are now looking to build systems that are more tailored to the needs of their customers. In technical parlance, this is known as the behavioral model for information retrieval system design. With these models, users search for a product or service, and the results often include related offerings that are better matched to the user's intent. Honing in on this kind of personalization is at the crux of the new experience economy of customer service and the forefront of Enterprise Search advancements. One of the key requirements for forward-looking knowledge management is the capacity to extract data from the typically hundreds and thousands of data silos scattered throughout a company and crawl them to create meaningful insights.