Information Retrieval
Fluent answers from AI search engines are more likely to be wrong
If you think search engines powered by artificial intelligence, such as Microsoft's Bing Chat, are providing you with useful-sounding answers, it is more likely that they are wrong, researchers have found. "In these current systems, accuracy is inversely correlated with perceived utility," says Nelson Liu at Stanford University. "The things that look better end up being worse."
Visual Diagrammatic Queries in ViziQuer: Overview and Implementation
Ovฤiลลikiva, Jลซlija, ล ostaks, Agris, ฤerฤns, Kฤrlis
Knowledge graphs (KG) have become an important data organization paradigm. The available textual query languages for information retrieval from KGs, as SPARQL for RDF-structured data, do not provide means for involving non-technical experts in the data access process. Visual query formalisms, alongside form-based and natural language-based ones, offer means for easing user involvement in the data querying process. ViziQuer is a visual query notation and tool offering visual diagrammatic means for describing rich data queries, involving optional and negation constructs, as well as aggregation and subqueries. In this paper we review the visual ViziQuer notation from the end-user point of view and describe the conceptual and technical solutions (including abstract syntax model, followed by a generation model for textual queries) that allow mapping of the visual diagrammatic query notation into the textual SPARQL language, thus enabling the execution of rich visual queries over the actual knowledge graphs. The described solutions demonstrate the viability of the model-based approach in translating complex visual notation into a complex textual one; they serve as semantics by implementation description of the ViziQuer language and provide building blocks for further services in the ViziQuer tool context.
BiTimeBERT: Extending Pre-Trained Language Representations with Bi-Temporal Information
Wang, Jiexin, Jatowt, Adam, Yoshikawa, Masatoshi, Cai, Yi
Time is an important aspect of documents and is used in a range of Temporal signals constitute significant features in various types NLP and IR tasks. In this work, we investigate methods for incorporating of text documents such as news articles or biographies. They can temporal information during pre-training to further improve be leveraged to understand chronology, causalities, developments, the performance on time-related tasks. Compared with common and ramifications of events, being helpful in a range of different pre-trained language models like BERT which utilize synchronic NLP tasks. Utilizing temporal signals in information retrieval has received document collections (e.g., BookCorpus and Wikipedia) as the training considerable attention recently, too. For example, researchers corpora, we use long-span temporal news article collection for have addressed time-sensitive queries in search leading to the formation building word representations. We introduce BiTimeBERT, a novel of a subset of Information Retrieval called Temporal Information language representation model trained on a temporal collection Retrieval [8, 26] in which both query and document of news articles via two new pre-training tasks, which harnesses temporal aspects are of key concern. Event detection and ordering two distinct temporal signals to construct time-aware language [14, 47], timeline summarization [2, 10, 36, 46, 50], event occurrence representations. The experimental results show that BiTimeBERT time prediction [54], temporal clustering [9], question answering consistently outperforms BERT and other existing pre-trained models [39, 52] and semantic change detection [41, 42] are other example with substantial gains on different downstream NLP tasks and tasks where utilizing temporal information has proven beneficial.
Multivariate Representation Learning for Information Retrieval
Zamani, Hamed, Bendersky, Michael
Dense retrieval models use bi-encoder network architectures for learning query and document representations. These representations are often in the form of a vector representation and their similarities are often computed using the dot product function. In this paper, we propose a new representation learning framework for dense retrieval. Instead of learning a vector for each query and document, our framework learns a multivariate distribution and uses negative multivariate KL divergence to compute the similarity between distributions. For simplicity and efficiency reasons, we assume that the distributions are multivariate normals and then train large language models to produce mean and variance vectors for these distributions. We provide a theoretical foundation for the proposed framework and show that it can be seamlessly integrated into the existing approximate nearest neighbor algorithms to perform retrieval efficiently. We conduct an extensive suite of experiments on a wide range of datasets, and demonstrate significant improvements compared to competitive dense retrieval models.
PREME: Preference-based Meeting Exploration through an Interactive Questionnaire
Arabzadeh, Negar, Ahmadvand, Ali, Kiseleva, Julia, Liu, Yang, Awadallah, Ahmed Hassan, Zhong, Ming, Shokouhi, Milad
The recent increase in the volume of online meetings necessitates automated tools for managing and organizing the material, especially when an attendee has missed the discussion and needs assistance in quickly exploring it. In this work, we propose a novel end-to-end framework for generating interactive questionnaires for preference-based meeting exploration. As a result, users are supplied with a list of suggested questions reflecting their preferences. Since the task is new, we introduce an automatic evaluation strategy. Namely, it measures how much the generated questions via questionnaire are answerable to ensure factual correctness and covers the source meeting for the depth of possible exploration.
Microsoft shares up 8.3% as AI features give a boost to sales
Microsoft Corp beat Wall Street's quarterly revenue and profit estimates on Tuesday, driven by growth in its cloud computing and Office productivity software businesses, and the company said artificial intelligence products were stimulating sales. The company forecast that revenue in its main segments for the current quarter would match or top Wall Street targets. Shares gained 8.3% in after-market trading following a report by the Redmond, Washington-based technology company that profits were $2.45 a share in the fiscal third quarter, beating Wall Street estimates of $2.23, according to data from Refinitiv and up 10% from the same quarter last year. In regular trading, fears about earnings had sent Microsoft down 2.2%, making it the biggest drag on the S&P 500 on Tuesday ahead of its report. Revenue rose 7% to $52.9bn in the quarter ended March, inching past the average analyst estimate of $51.02bn, according to Refinitiv.
Explain like I am BM25: Interpreting a Dense Model's Ranked-List with a Sparse Approximation
Llordes, Michael, Ganguly, Debasis, Bhatia, Sumit, Agarwal, Chirag
Neural retrieval models (NRMs) have been shown to outperform their statistical counterparts owing to their ability to capture semantic meaning via dense document representations. These models, however, suffer from poor interpretability as they do not rely on explicit term matching. As a form of local per-query explanations, we introduce the notion of equivalent queries that are generated by maximizing the similarity between the NRM's results and the result set of a sparse retrieval system with the equivalent query. We then compare this approach with existing methods such as RM3-based query expansion and contrast differences in retrieval effectiveness and in the terms generated by each approach.
Hitachi at SemEval-2023 Task 3: Exploring Cross-lingual Multi-task Strategies for Genre and Framing Detection in Online News
Koreeda, Yuta, Yokote, Ken-ichi, Ozaki, Hiroaki, Yamaguchi, Atsuki, Tsunokake, Masaya, Sogawa, Yasuhiro
This paper explains the participation of team Hitachi to SemEval-2023 Task 3 "Detecting the genre, the framing, and the persuasion techniques in online news in a multi-lingual setup.'' Based on the multilingual, multi-task nature of the task and the low-resource setting, we investigated different cross-lingual and multi-task strategies for training the pretrained language models. Through extensive experiments, we found that (a) cross-lingual/multi-task training, and (b) collecting an external balanced dataset, can benefit the genre and framing detection. We constructed ensemble models from the results and achieved the highest macro-averaged F1 scores in Italian and Russian genre categorization subtasks.
Modeling Spoken Information Queries for Virtual Assistants: Open Problems, Challenges and Opportunities
Virtual assistants are becoming increasingly important speech-driven Information Retrieval platforms that assist users with various tasks. We discuss open problems and challenges with respect to modeling spoken information queries for virtual assistants, and list opportunities where Information Retrieval methods and research can be applied to improve the quality of virtual assistant speech recognition. We discuss how query domain classification, knowledge graphs and user interaction data, and query personalization can be helpful to improve the accurate recognition of spoken information domain queries. Finally, we also provide a brief overview of current problems and challenges in speech recognition.
GARCIA: Powering Representations of Long-tail Query with Multi-granularity Contrastive Learning
Wang, Weifan, Hu, Binbin, Peng, Zhicheng, Zhong, Mingjie, Zhang, Zhiqiang, Liu, Zhongyi, Zhang, Guannan, Zhou, Jun
Recently, the growth of service platforms brings great convenience to both users and merchants, where the service search engine plays a vital role in improving the user experience by quickly obtaining desirable results via textual queries. Unfortunately, users' uncontrollable search customs usually bring vast amounts of long-tail queries, which severely threaten the capability of search models. Inspired by recently emerging graph neural networks (GNNs) and contrastive learning (CL), several efforts have been made in alleviating the long-tail issue and achieve considerable performance. Nevertheless, they still face a few major weaknesses. Most importantly, they do not explicitly utilize the contextual structure between heads and tails for effective knowledge transfer, and intention-level information is commonly ignored for more generalized representations. To this end, we develop a novel framework GARCIA, which exploits the graph based knowledge transfer and intention based representation generalization in a contrastive setting. In particular, we employ an adaptive encoder to produce informative representations for queries and services, as well as hierarchical structure aware representations of intentions. To fully understand tail queries and services, we equip GARCIA with a novel multi-granularity contrastive learning module, which powers representations through knowledge transfer, structure enhancement and intention generalization. Subsequently, the complete GARCIA is well trained in a pre-training&fine-tuning manner. At last, we conduct extensive experiments on both offline and online environments, which demonstrates the superior capability of GARCIA in improving tail queries and overall performance in service search scenarios.