Information Retrieval
A Clarifying Question Selection System from NTES_ALONG in Convai3 Challenge
This paper presents the participation of NTES\_ALONG team for the ClariQ challenge at Search-oriented Conversational AI (SCAI) EMNLP workshop in 2020. The challenge asks for a complete conversational information retrieval system that can understanding and generating clarification questions. We propose a clarifying question selection system which consists of response understanding, candidate question recalling and clarifying question ranking. We fine-tune a RoBERTa model to understand user's responses and use an enhanced BM25 model to recall the candidate questions. In clarifying question ranking stage, we reconstruct the training dataset and propose two models based on ELECTRA. Finally we ensemble the models by summing up their output probabilities and choose the question with the highest probability as the clarification question. Experiments show that our ensemble ranking model outperforms in the document relevance task and achieves the best recall@[20,30] metrics in question relevance task.
QBSUM: a Large-Scale Query-Based Document Summarization Dataset from Real-world Applications
Zhao, Mingjun, Yan, Shengli, Liu, Bang, Zhong, Xinwang, Hao, Qian, Chen, Haolan, Niu, Di, Long, Bowei, Guo, Weidong
Query-based document summarization aims to extract or generate a summary of a document which directly answers or is relevant to the search query. It is an important technique that can be beneficial to a variety of applications such as search engines, document-level machine reading comprehension, and chatbots. Currently, datasets designed for query-based summarization are short in numbers and existing datasets are also limited in both scale and quality. Moreover, to the best of our knowledge, there is no publicly available dataset for Chinese query-based document summarization. In this paper, we present QBSUM, a high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization. We also propose multiple unsupervised and supervised solutions to the task and demonstrate their high-speed inference and superior performance via both offline experiments and online A/B tests. The QBSUM dataset is released in order to facilitate future advancement of this research field.
Query Complexity of k-NN based Mode Estimation
Singhal, Anirudh, Pirojiwala, Subham, Karamchandani, Nikhil
Motivated by the mode estimation problem of an unknown multivariate probability density function, we study the problem of identifying the point with the minimum k-th nearest neighbor distance for a given dataset of n points. We study the case where the pairwise distances are apriori unknown, but we have access to an oracle which we can query to get noisy information about the distance between any pair of points. For two natural oracle models, we design a sequential learning algorithm, based on the idea of confidence intervals, which adaptively decides which queries to send to the oracle and is able to correctly solve the problem with high probability. We derive instance-dependent upper bounds on the query complexity of our proposed scheme and also demonstrate significant improvement over the performance of other baselines via extensive numerical evaluations.
A Survey of Embedding Space Alignment Methods for Language and Knowledge Graphs
Kalinowski, Alexander, An, Yuan
The purpose of this survey is to explore the core techniques and categorizations of methods for aligning low-dimensional embedding spaces. Projecting sparse, high-dimensional data sets into compact, lower-dimensional spaces allows not only for a significant reduction in storage space, but also builds dense representations with many applications. These embedding spaces have become a staple in representation learning ever since their heralded application to natural language in a technique called word2vec, and have replaced traditional machine learning features as easy-to-build, high-quality representations of the source objects. There has been a wealth of study around techniques for embedding objects, such as images, natural language and knowledge graphs, and many research agendas focused on mapping one embedding space to another, either for the purpose of aligning and unifying to a common space, applications to joint downstream tasks or ease of transfer learning. In order to fully leverage these dense representations and translate them across domains and problem spaces, techniques for establishing alignments between them must be developed and understood.
Chile's New Interdisciplinary Institute for Foundational Research on Data
The Millennium Institute for Foundational Research on Dataa (IMFD) started its operations in June 2018, funded by the Millennium Science Initiative of the Chilean National Agency of Research and Development.b IMFD is a joint initiative led by Universidad de Chile and Universidad Catรณlica de Chile, with the participation of five other Chilean universities: Universidad de Concepciรณn, Universidad de Talca, Universidad Tรฉcnica Federico Santa Marรญa, Universidad Diego Portales, and Universidad Adolfo Ibรกรฑez. IMFD aims to be a reference center in Latin America related to state-of-the-art research on the foundational problems with data, as well as its applications to tackling diverse issues ranging from scientific challenges to complex social problems. As tasks of this kind are interdisciplinary by nature, IMFD gathers a large number of researchers in several areas that include traditional computer science areas such as data management, Web science, algorithms and data structures, privacy and verification, information retrieval, data mining, machine learning, and knowledge representation, as well as some areas from other fields, including statistics, political science, and communication studies. IMFD currently hosts 36 researchers, seven postdoctoral fellows, and more than 100 students.
Keyphrase Extraction with Dynamic Graph Convolutional Networks and Diversified Inference
Zhang, Haoyu, Long, Dingkun, Xu, Guangwei, Xie, Pengjun, Huang, Fei, Wang, Ji
Keyphrase extraction (KE) aims to summarize a set of phrases that accurately express a concept or a topic covered in a given document. Recently, Sequence-to-Sequence (Seq2Seq) based generative framework is widely used in KE task, and it has obtained competitive performance on various benchmarks. The main challenges of Seq2Seq methods lie in acquiring informative latent document representation and better modeling the compositionality of the target keyphrases set, which will directly affect the quality of generated keyphrases. In this paper, we propose to adopt the Dynamic Graph Convolutional Networks (DGCN) to solve the above two problems simultaneously. Concretely, we explore to integrate dependency trees with GCN for latent representation learning. Moreover, the graph structure in our model is dynamically modified during the learning process according to the generated keyphrases. To this end, our approach is able to explicitly learn the relations within the keyphrases collection and guarantee the information interchange between encoder and decoder in both directions. Extensive experiments on various KE benchmark datasets demonstrate the effectiveness of our approach.
Google Paid Apple Billions To Dominate Search On iPhones, Justice Department Says
The Justice Department says Google CEO Sundar Pichai (left) met privately with Apple chief Tim Cook in 2018 to discuss how their two companies could collaborate. The Justice Department says Google CEO Sundar Pichai (left) met privately with Apple chief Tim Cook in 2018 to discuss how their two companies could collaborate. Buried on page 36 of the Justice Department lawsuit accusing Google of abusing its monopoly power is this remarkable figure: $8 billion to $12 billion. That's the hefty sum Google allegedly paid Apple for one of the most prized pieces of real estate in the world of online search: default status on iPhones and all other Apple devices. Justice Department investigators say Apple, which does not have its own search engine, hammered out a multiyear deal making Google the default search engine on all iPhones and other Apple products.
Migratable AI: Personalizing Dialog Conversations with migration context
Tejwani, Ravi, Katz, Boris, Breazeal, Cynthia
The migration of conversational AI agents across different embodiments in order to maintain the continuity of the task has been recently explored to further improve user experience. However, these migratable agents lack contextual understanding of the user information and the migrated device during the dialog conversations with the user. This opens the question of how an agent might behave when migrated into an embodiment for contextually predicting the next utterance. We collected a dataset from the dialog conversations between crowdsourced workers with the migration context involving personal and non-personal utterances in different settings (public or private) of embodiment into which the agent migrated. We trained the generative and information retrieval models on the dataset using with and without migration context and report the results of both qualitative metrics and human evaluation. We believe that the migration dataset would be useful for training future migratable AI systems.
3Diligent Expands ProdEX and Shopsight Applications
Its Shopsight application provides users access to project opportunities from ProdEX and enables remote assessment, quoting, and project management. Both systems incorporate 3Diligent's Connect interface which enables customers and manufacturers to communicate directly using a secure online portal and Zoom video conferencing tools. Operating similarly to traditional search engine marketing, manufacturers can create text ads that will display based on a customer's material and technology requirements. However, unlike traditional search engines, Connect is driven by RFQ inputs rather than generic keyword searches. As a result, manufacturers can customize their bids and visibility on dimensions such as material, technology, and program size to drive higher ROI.
ColloQL: Robust Cross-Domain Text-to-SQL Over Search Queries
Radhakrishnan, Karthik, Srikantan, Arvind, Lin, Xi Victoria
Translating natural language utterances to executable queries is a helpful technique in making the vast amount of data stored in relational databases accessible to a wider range of non-tech-savvy end users. Prior work in this area has largely focused on textual input that is linguistically correct and semantically unambiguous. However, real-world user queries are often succinct, colloquial, and noisy, resembling the input of a search engine. In this work, we introduce data augmentation techniques and a sampling-based content-aware BERT model (ColloQL) to achieve robust text-to-SQL modeling over natural language search (NLS) questions. Due to the lack of evaluation data, we curate a new dataset of NLS questions and demonstrate the efficacy of our approach. ColloQL's superior performance extends to well-formed text, achieving 84.9% (logical) and 90.7% (execution) accuracy on the WikiSQL dataset, making it, to the best of our knowledge, the highest performing model that does not use execution guided decoding.