Goto

Collaborating Authors

 Information Retrieval


Sub-GMN: The Neural Subgraph Matching Network Model

arXiv.org Artificial Intelligence

As one of the most fundamental tasks in graph theory, subgraph matching is a crucial task in many fields, ranging from information retrieval, computer vision, biology, chemistry and natural language processing. Yet subgraph matching problem remains to be an NP-complete problem. This study proposes an end-to-end learning-based approximate method for subgraph matching task, called subgraph matching network (Sub-GMN). The proposed Sub-GMN firstly uses graph representation learning to map nodes to node-level embedding. It then combines metric learning and attention mechanisms to model the relationship between matched nodes in the data graph and query graph. To test the performance of the proposed method, we applied our method on two databases. We used two existing methods, GNN and FGNN as baseline for comparison. Our experiment shows that, on dataset 1, on average the accuracy of Sub-GMN are 12.21\% and 3.2\% higher than that of GNN and FGNN respectively. On average running time Sub-GMN runs 20-40 times faster than FGNN. In addition, the average F1-score of Sub-GMN on all experiments with dataset 2 reached 0.95, which demonstrates that Sub-GMN outputs more correct node-to-node matches. Comparing with the previous GNNs-based methods for subgraph matching task, our proposed Sub-GMN allows varying query and data graphes in the test/application stage, while most previous GNNs-based methods can only find a matched subgraph in the data graph during the test/application for the same query graph used in the training stage. Another advantage of our proposed Sub-GMN is that it can output a list of node-to-node matches, while most existing end-to-end GNNs based methods cannot provide the matched node pairs.


Pre-training for Information Retrieval: Are Hyperlinks Fully Explored?

arXiv.org Artificial Intelligence

Recent years have witnessed great progress on applying pre-trained language models, e.g., BERT, to information retrieval (IR) tasks. Hyperlinks, which are commonly used in Web pages, have been leveraged for designing pre-training objectives. For example, anchor texts of the hyperlinks have been used for simulating queries, thus constructing tremendous query-document pairs for pre-training. However, as a bridge across two web pages, the potential of hyperlinks has not been fully explored. In this work, we focus on modeling the relationship between two documents that are connected by hyperlinks and designing a new pre-training objective for ad-hoc retrieval. Specifically, we categorize the relationships between documents into four groups: no link, unidirectional link, symmetric link, and the most relevant symmetric link. By comparing two documents sampled from adjacent groups, the model can gradually improve its capability of capturing matching signals. We propose a progressive hyperlink predication ({PHP}) framework to explore the utilization of hyperlinks in pre-training. Experimental results on two large-scale ad-hoc retrieval datasets and six question-answering datasets demonstrate its superiority over existing pre-training methods.


How to Transform Your Data into a Voice AI Knowledge Assistant - Coruzant Technologies

#artificialintelligence

Almost every enterprise believes data to be one of their most important assets, but most would admit they are not leveraging their data to its full potential. That's because making data easily accessible to employees is surprisingly hard work. It requires a concerted, ongoing effort to gather, structure, and tag data in order to turn it into knowledge, capable of being found in the moments when it can be most useful, using a voice AI knowledge assistant. Of course, this has always been true regardless of the state of technology, from ancient libraries indexing vast physical volumes to today's cloud-based search engines crawling millions of gigabytes of data. We can take for granted the magic of the now-ubiquitous digital keyword search, which has delivered us the power to have any data point just a few mouse and keyboard clicks away.


Y Combinator-backed Andi taps AI to build a better search engine

#artificialintelligence

It's difficult to convince users to switch search engines. That's one reason why public search service startups rarely succeed. Another is that it's expensive to index a huge number of websites (Google has an estimated tens of billions of pages indexed), but one Y Combinator-backed company, Andi, is undeterred -- forging ahead to build an AI assistant that provides answers instead of links when searching online. Andi was founded by Angela Hoover, who registered for YC's Startup School after dropping out of college and got into YC's Winter 2022 Batch. After working overseas in construction and with Microsoft as a data center project administrator, Hoover met Andi's co-founder, Jed White, at the Denver airport upon her return to the U.S. Hoover and White -- who had a background in AI and search, specifically content quality ranking, querying and classification -- talked about how bad web search had become for things like travel and what it would take to build a new type of search engine from scratch.


An Embedding-Based Grocery Search Model at Instacart

arXiv.org Artificial Intelligence

The key to e-commerce search is how to best utilize the large yet noisy log data. In this paper, we present our embedding-based model for grocery search at Instacart. The system learns query and product representations with a two-tower transformer-based encoder architecture. To tackle the cold-start problem, we focus on content-based features. To train the model efficiently on noisy data, we propose a self-adversarial learning method and a cascade training method. AccOn an offline human evaluation dataset, we achieve 10% relative improvement in RECALL@20, and for online A/B testing, we achieve 4.1% cart-adds per search (CAPS) and 1.5% gross merchandise value (GMV) improvement. We describe how we train and deploy the embedding based search model and give a detailed analysis of the effectiveness of our method.


Representing Social Networks as Dynamic Heterogeneous Graphs

arXiv.org Artificial Intelligence

Graph representations for real-world social networks in the past have missed two important elements: the multiplexity of connections as well as representing time. To this end, in this paper, we present a new dynamic heterogeneous graph representation for social networks which includes time in every single component of the graph, i.e., nodes and edges, each of different types that captures heterogeneity. We illustrate the power of this representation by presenting four time-dependent queries and deep learning problems that cannot easily be handled in conventional homogeneous graph representations commonly used. As a proof of concept we present a detailed representation of a new social media platform (Steemit), which we use to illustrate both the dynamic querying capability as well as prediction tasks using graph neural networks (GNNs). The results illustrate the power of the dynamic heterogeneous graph representation to model social networks. Given that this is a relatively understudied area we also illustrate opportunities for future work in query optimization as well as new dynamic prediction tasks on heterogeneous graph structures.


Large-scale Evaluation of Transformer-based Article Encoders on the Task of Citation Recommendation

arXiv.org Artificial Intelligence

Recently introduced transformer-based article encoders (TAEs) designed to produce similar vector representations for mutually related scientific articles have demonstrated strong performance on benchmark datasets for scientific article recommendation. However, the existing benchmark datasets are predominantly focused on single domains and, in some cases, contain easy negatives in small candidate pools. Evaluating representations on such benchmarks might obscure the realistic performance of TAEs in setups with thousands of articles in candidate pools. In this work, we evaluate TAEs on large benchmarks with more challenging candidate pools. We compare the performance of TAEs with a lexical retrieval baseline model BM25 on the task of citation recommendation, where the model produces a list of recommendations for citing in a given input article. We find out that BM25 is still very competitive with the state-of-the-art neural retrievers, a finding which is surprising given the strong performance of TAEs on small benchmarks. As a remedy for the limitations of the existing benchmarks, we propose a new benchmark dataset for evaluating scientific article representations: Multi-Domain Citation Recommendation dataset (MDCR), which covers different scientific fields and contains challenging candidate pools.


How AI Writing Tools are Revolutionizing Content Creation (2022)

#artificialintelligence

Search engines are constantly evolving, and as a result, the way we create and consume content is also changing. In particular, the rise of artificial intelligence (AI) writing tools is revolutionizing the content creation process. AI writing software is now being used by bloggers and businesses to create high-quality content quickly and easily. This software can analyze data and find trends to help you write about what's popular right now. It can also help you come up with catchy headlines and create drafts that are ready for publishing. AI writing tools have improved a great deal over the past few years and now they can help with writing articles, digital ad copy, blog post ideas, youtube video descriptions, and Google ads all fast and in multiple languages. In this article, we'll discuss how AI writing software is changing the way bloggers and businesses create content and answer some frequently asked questions about this technology. AI writing tools are computer programs that can generate written content. AI tools can be used to create blog articles, website content, or even sales letters. Most AI tools use natural language processing (NLP) to understand the topic and then generate relevant content. AI writing tools can save you a lot of time by quickly generating high-quality content. Just enter a few keywords and the AI tool will do the rest.


Code Compliance Assessment as a Learning Problem

arXiv.org Artificial Intelligence

Manual code reviews and static code analyzers are the traditional mechanisms to verify if source code complies with coding policies. However, these mechanisms are hard to scale. We formulate code compliance assessment as a machine learning (ML) problem, to take as input a natural language policy and code, and generate a prediction on the code's compliance, non-compliance, or irrelevance. This can help scale compliance classification and search for policies not covered by traditional mechanisms. We explore key research questions on ML model formulation, training data, and evaluation setup. The core idea is to obtain a joint code-text embedding space which preserves compliance relationships via the vector distance of code and policy embeddings. As there is no task-specific data, we re-interpret and filter commonly available software datasets with additional pre-training and pre-finetuning tasks that reduce the semantic gap. We benchmarked our approach on two listings of coding policies (CWE and CBP). This is a zero-shot evaluation as none of the policies occur in the training set. On CWE and CBP respectively, our tool Policy2Code achieves classification accuracies of (59%, 71%) and search MRR of (0.05, 0.21) compared to CodeBERT with classification accuracies of (37%, 54%) and MRR of (0.02, 0.02). In a user study, 24% Policy2Code detections were accepted compared to 7% for CodeBERT.


Apple could lose $15B if DOJ forces Google to stop paying to be iPhone's default search engine

Daily Mail - Science & tech

Apple stands to lose up to $15 billion a year if the Justice Department forces Google to stop paying the company to be the default search engine on all iPhones - as regulators question the legality of the longtime arrangement. Anytime iPhone users open a web browser to enter a search query, it always defaults to Google. Even though anyone can change this setting, almost no one does, resulting in a huge amount of traffic (and ad revenue) to Google from over a billion iPhone users worldwide. Analysts from Bernstein estimated that Google's payment to Apple would increase to $15 billion in 2021 and as high as $18-$20 billion this year, reports 9to5Mac. The contracts are the basis of the DOJ's antitrust against the California-based company, which began in the closing days of the Trump administration and won't head to trial until sometime in 2023 Last year, Apple's total gross profit was over $152 billion - so losing the Google payments would shave at least 10% off.