Goto

Collaborating Authors

 Information Retrieval


BioADAPT-MRC: Adversarial Learning-based Domain Adaptation Improves Biomedical Machine Reading Comprehension Task

arXiv.org Artificial Intelligence

Biomedical machine reading comprehension (biomedical-MRC) aims to comprehend complex biomedical narratives and assist healthcare professionals in retrieving information from them. The high performance of modern neural network-based MRC systems depends on high-quality, large-scale, human-annotated training datasets. In the biomedical domain, a crucial challenge in creating such datasets is the requirement for domain knowledge, inducing the scarcity of labeled data and the need for transfer learning from the labeled general-purpose (source) domain to the biomedical (target) domain. However, there is a discrepancy in marginal distributions between the general-purpose and biomedical domains due to the variances in topics. Therefore, direct-transferring of learned representations from a model trained on a general-purpose domain to the biomedical domain can hurt the model's performance. We present an adversarial learning-based domain adaptation framework for the biomedical machine reading comprehension task (BioADAPT-MRC), a neural network-based method to address the discrepancies in the marginal distributions between the general and biomedical domain datasets. BioADAPT-MRC relaxes the need for generating pseudo labels for training a well-performing biomedical-MRC model. We extensively evaluate the performance of BioADAPT-MRC by comparing it with the best existing methods on three widely used benchmark biomedical-MRC datasets -- BioASQ-7b, BioASQ-8b, and BioASQ-9b. Our results suggest that without using any synthetic or human-annotated data from the biomedical domain, BioADAPT-MRC can achieve state-of-the-art performance on these datasets. Availability: BioADAPT-MRC is freely available as an open-source project at \url{https://github.com/mmahbub/BioADAPT-MRC}.


Buffer Pool Aware Query Scheduling via Deep Reinforcement Learning

arXiv.org Artificial Intelligence

One could imagine many simple heuristics, query scheduling with the explicit goal of reducing disk reads such as greedily selecting the next query with the highest and thus implicitly increasing query performance. We introduce expected buffer usage, to solve this problem. However, a SmartQueue, a learned scheduler that leverages overlapping hand-designed policy to handle the complexity of the entire data reads among incoming queries and learns a problem, including different buffer sizes, shifting query scheduling strategy that improves cache hits. SmartQueue workloads, heterogeneous data types (e.g., index files vs base relies on deep reinforcement learning to produce workloadspecific relations), and balancing short-term gains against long-term scheduling strategies that focus on long-term performance strategy is much more difficult to conceive.


What you Know about Keywords and their Importance in SEO

#artificialintelligence

Modern internet growth has created the need for many skills, which are called digital skills. One of these skills is search engine optimization. So people can improve their skills by reading online content from our website. In this article, although our intended readers are beginners, professionals can also refresh their knowledge. If you know about keywords, then you must know about search engine optimization.


Facing Changes: Continual Entity Alignment for Growing Knowledge Graphs

arXiv.org Artificial Intelligence

Entity alignment is a basic and vital technique in knowledge graph (KG) integration. Over the years, research on entity alignment has resided on the assumption that KGs are static, which neglects the nature of growth of real-world KGs. As KGs grow, previous alignment results face the need to be revisited while new entity alignment waits to be discovered. In this paper, we propose and dive into a realistic yet unexplored setting, referred to as continual entity alignment. To avoid retraining an entire model on the whole KGs whenever new entities and triples come, we present a continual alignment method for this task. It reconstructs an entity's representation based on entity adjacency, enabling it to generate embeddings for new entities quickly and inductively using their existing neighbors. It selects and replays partial pre-aligned entity pairs to train only parts of KGs while extracting trustworthy alignment for knowledge augmentation. As growing KGs inevitably contain non-matchable entities, different from previous works, the proposed method employs bidirectional nearest neighbor matching to find new entity alignment and update old alignment. Furthermore, we also construct new datasets by simulating the growth of multilingual DBpedia. Extensive experiments demonstrate that our continual alignment method is more effective than baselines based on retraining or inductive learning.


Meta AI introduces Sphere, a model designed to verify citations on Wikipedia - Actu IA

#artificialintelligence

When we do a search on the Internet, the search engine very often suggests the site of the community encyclopedia Wikipedia. It contains about 6.5 million articles by volunteer contributors, but how can we know if these are reliable, even though the sources of the articles are cited? Meta relied on Meta AI's research and advances to develop SPHERE, an open source model capable of automatically analyzing hundreds of thousands of citations at a time to check whether they actually support the corresponding claims, it recently published it on the Github platform. Meta said it is not partnering with Wikimedia, the foundation that runs Wikipedia, on this project. Its goal is to create a platform to help Wikipedia editors systematically spot citation problems and quickly correct the citation or the corresponding article content.


Important digital skill-what are the basics of SEO(search engine optimization)?

#artificialintelligence

With the growth of the internet, the world is becoming digital. Millions of websites are being created every day and digital content is being uploaded to them. Every site's purpose is to reach its potential audience at an earlier base. The search Engine Optimization (SEO) technique is used for this purpose. SEO is the method by which it brings organic traffic to your website.


Consistent Polyhedral Surrogates for Top-$k$ Classification and Variants

arXiv.org Artificial Intelligence

Top-$k$ classification is a generalization of multiclass classification used widely in information retrieval, image classification, and other extreme classification settings. Several hinge-like (piecewise-linear) surrogates have been proposed for the problem, yet all are either non-convex or inconsistent. For the proposed hinge-like surrogates that are convex (i.e., polyhedral), we apply the recent embedding framework of Finocchiaro et al. (2019; 2022) to determine the prediction problem for which the surrogate is consistent. These problems can all be interpreted as variants of top-$k$ classification, which may be better aligned with some applications. We leverage this analysis to derive constraints on the conditional label distributions under which these proposed surrogates become consistent for top-$k$. It has been further suggested that every convex hinge-like surrogate must be inconsistent for top-$k$. Yet, we use the same embedding framework to give the first consistent polyhedral surrogate for this problem.


MIA 2022 Shared Task Submission: Leveraging Entity Representations, Dense-Sparse Hybrids, and Fusion-in-Decoder for Cross-Lingual Question Answering

arXiv.org Artificial Intelligence

We describe our two-stage system for the Multilingual Information Access (MIA) 2022 Shared Task on Cross-Lingual Open-Retrieval Question Answering. The first stage consists of multilingual passage retrieval with a hybrid dense and sparse retrieval strategy. The second stage consists of a reader which outputs the answer from the top passages returned by the first stage. We show the efficacy of using a multilingual language model with entity representations in pretraining, sparse retrieval signals to help dense retrieval, and Fusion-in-Decoder. On the development set, we obtain 43.46 F1 on XOR-TyDi QA and 21.99 F1 on MKQA, for an average F1 score of 32.73. On the test set, we obtain 40.93 F1 on XOR-TyDi QA and 22.29 F1 on MKQA, for an average F1 score of 31.61. We improve over the official baseline by over 4 F1 points on both the development and test sets.


Heuristic-free Optimization of Force-Controlled Robot Search Strategies in Stochastic Environments

arXiv.org Artificial Intelligence

In both industrial and service domains, a central benefit of the use of robots is their ability to quickly and reliably execute repetitive tasks. However, even relatively simple peg-in-hole tasks are typically subject to stochastic variations, requiring search motions to find relevant features such as holes. While search improves robustness, it comes at the cost of increased runtime: More exhaustive search will maximize the probability of successfully executing a given task, but will significantly delay any downstream tasks. This trade-off is typically resolved by human experts according to simple heuristics, which are rarely optimal. This paper introduces an automatic, data-driven and heuristic-free approach to optimize robot search strategies. By training a neural model of the search strategy on a large set of simulated stochastic environments, conditioning it on few real-world examples and inverting the model, we can infer search strategies which adapt to the time-variant characteristics of the underlying probability distributions, while requiring very few real-world measurements. We evaluate our approach on two different industrial robots in the context of spiral and probe search for THT electronics assembly.


You.com raises $25M to fuel its AI-powered search engine – TechCrunch

#artificialintelligence

At least, that's the crux of the argument Richard Socher, the former chief scientist at Salesforce, likes to make. In 2020, Socher co-founded You, a search engine that uses AI to understand search queries, rank the results and parse the queries into different languages (including programming languages). You summarizes information from across the web and offers built-in apps, like search tools for Twitter, that allow users to complete tasks without having to leave the results page. It seems there's some truth to his words. Socher claims that You has hundreds of thousands of users, with 70% growth in sign-ups last month and 30% growth in unique searches month over month.