Goto

Collaborating Authors

 Information Retrieval


Risk-based Adaptive Deep Learning for Entity Resolution

arXiv.org Artificial Intelligence

The state-of-the-art performance on entity resolution (ER) has been achieved by deep learning. However, deep models are usually trained on large quantities of accurately labeled training data, and can not be easily tuned towards a target workload. Unfortunately, in real scenarios, there may not be sufficient labeled training data, and even worse, their distribution is usually more or less different from the target workload even when they come from the same domain. To alleviate the said limitations, this paper proposes a novel risk-based approach to tune a deep model towards a target workload by its particular characteristics. Built on the recent advances on risk analysis for ER, the proposed approach first trains a deep model on labeled training data, and then fine-tunes it by minimizing its estimated misprediction risk on unlabeled target data. Our theoretical analysis shows that risk-based adaptive training can correct the label status of a mispredicted instance with a fairly good chance. We have also empirically validated the efficacy of the proposed approach on real benchmark data by a comparative study. Our extensive experiments show that it can considerably improve the performance of deep models. Furthermore, in the scenario of distribution misalignment, it can similarly outperform the state-of-the-art alternative of transfer learning by considerable margins. Using ER as a test case, we demonstrate that risk-based adaptive training is a promising approach potentially applicable to various challenging classification tasks.


information retrieval document search using vector space model in R

#artificialintelligence

Now calculate cosine similarity between each document and each query. For each query sort the cosine similarity scores for all the documents and take top-3 documents having high scores.


Airbus AI Introduces Natural Language QA System for Flight Crews

#artificialintelligence

Airbus AI researchers have developed a system that uses natural language understanding to improve question answering (QA) performance when flight crews search for aircraft operating information. The aerospace industry relies on technical documents such as Aircraft Operating Manuals (AOM), Aircraft Operating Instructions and particularly Flight Crew Operating Manuals (FCOM) to guide flight crews on aircraft operations under normal, abnormal, and emergency conditions. FCOMs are issued by aircraft manufacturers and cover system descriptions, procedures, techniques, and performance data. They are the references used to develop standard operating procedures to improve safety and efficiency. Most government aviation administrations have authorized the use of tablet computers by commercial carrier pilots and flight crews to access FCOM information. The Airbus AI researchers note however that existing electronic flight bag (EFB) systems used for this purpose are in practice little more than pdf viewers with keyword search functionality.


End-to-End QA on COVID-19: Domain Adaptation with Synthetic Training

arXiv.org Artificial Intelligence

End-to-end question answering (QA) requires both information retrieval (IR) over a large document collection and machine reading comprehension (MRC) on the retrieved passages. Recent work has successfully trained neural IR systems using only supervised question answering (QA) examples from open-domain datasets. However, despite impressive performance on Wikipedia, neural IR lags behind traditional term matching approaches such as BM25 in more specific and specialized target domains such as COVID-19. Furthermore, given little or no labeled data, effective adaptation of QA systems can also be challenging in such target domains. In this work, we explore the application of synthetically generated QA examples to improve performance on closed-domain retrieval and MRC. We combine our neural IR and MRC systems and show significant improvements in end-to-end QA on the CORD-19 collection over a state-of-the-art open-domain QA baseline.


ClimaText: A Dataset for Climate Change Topic Detection

arXiv.org Artificial Intelligence

Climate change communication in the mass media and other textual sources may affect and shape public perception. Extracting climate change information from these sources is an important task, e.g., for filtering content and e-discovery, sentiment analysis, automatic summarization, question-answering, and fact-checking. However, automating this process is a challenge, as climate change is a complex, fast-moving, and often ambiguous topic with scarce resources for popular text-based AI tasks. In this paper, we introduce \textsc{ClimaText}, a dataset for sentence-based climate change topic detection, which we make publicly available. We explore different approaches to identify the climate change topic in various text sources. We find that popular keyword-based models are not adequate for such a complex and evolving task. Context-based algorithms like BERT \cite{devlin2018bert} can detect, in addition to many trivial cases, a variety of complex and implicit topic patterns. Nevertheless, our analysis reveals a great potential for improvement in several directions, such as, e.g., capturing the discussion on indirect effects of climate change. Hence, we hope this work can serve as a good starting point for further research on this topic.


Search Engine Optimization Complete Specialization Course

#artificialintelligence

Welcome to the World's best specialized SEO course ever. This is the only course in the world where you woll also learn about the technicalities of SEO and how to handle them. The content of this course is based on real world practices and checklists used by professionals in the SEO world. How to get a job in SEO? How to start your own digital marketing company?


Think Search Is Solved? Think Again

#artificialintelligence

Search is one of the oldest technologies around. Ever since the dawn of the World Wide Web, a search engine has been the portal through which we obtain information. The search for a better search engine index kick started the Hadoop craze, and it continues to drive Google to push the limits of technology. But don't for a second think that search has been solved. "Search is far from being solved. It's the hardest thing we do. It's the hardest thing everybody does."


Can Artificial Intelligence be friends of Humans - OnPassive

#artificialintelligence

A question arises that how will it become self-aware and realize that humans stand in its way? Artificial Intelligence is the capability of a digital computer or computer-controlled robot that performs a task commonly associated with intelligent beings. Robots and AI allow producing things faster, better, and cheaper with higher consistency. AI is very disruptive for low-cost countries that provide low-cost manufacturing for international companies since robots do this cheaply. It is also disruptive to countries with higher salary levels, but not at the same level as low-cost countries. Our forefathers had the same concern with industrial revolutions.


Code Search Intent Classification Using Weak Supervision

arXiv.org Artificial Intelligence

Developers use search for various tasks such as finding code, documentation, debugging information, etc. In particular, web search is heavily used by developers for finding code examples and snippets during the coding process. Recently, natural language based code search has been an active area of research. However, the lack of real-world large-scale datasets is a significant bottleneck. In this work, we propose a weak supervision based approach for detecting code search intent in search queries for C# and Java programming languages. We evaluate the approach against several baselines on a real-world dataset comprised of over 1 million queries mined from Bing web search engine and show that the CNN based model can achieve an accuracy of 77% and 76% for C# and Java respectively. Furthermore, we are also releasing the first large-scale real-world dataset of code search queries mined from Bing web search engine. We hope that the dataset will aid future research on code search.


Ghostery's New Search Engine Will Be Entirely Ad-Free

WIRED

The internet runs on advertising, and that includes search engines. Google brought in $26 billion of search revenue in the most recent quarter alone. As that business has grown, it's reshaped what search looks like. Year after year, ads have gobbled up more space on its results pages, pushing organic results further out of view. Which is why using Ghostery's new ad-free search engine and desktop browser, even in their pre-beta form, feels at once like a throwback to a simpler internet and a glimpse of a future where browsing that puts results ahead of revenue is once again possible.