Goto

Collaborating Authors

 Information Retrieval


Human memory search as a random walk in a semantic network

Neural Information Processing Systems

The human mind has a remarkable ability to store a vast amount of information in memory, and an even more remarkable ability to retrieve these experiences when needed. Understanding the representations and algorithms that underlie human memory search could potentially be useful in other information retrieval settings, including internet search. Psychological studies have revealed clear regularities in how people search their memory, with clusters of semantically related items tending to be retrieved together. These findings have recently been taken as evidence that human memory search is similar to animals foraging for food in patchy environments, with people making a rational decision to switch away from a cluster of related information as it becomes depleted. We demonstrate that the results that were taken as evidence for this account also emerge from a random walk on a semantic network, much like the random web surfer model used in internet search engines. This offers a simpler and more unified account of how people search their memory, postulating a single process rather than one process for exploring a cluster and one process for switching between clusters.


Contextual Clarity: Generating Sentences with Transformer Models using Context-Reverso Data

arXiv.org Artificial Intelligence

To create a dataset for training the T5 model, we harness the power of that provides usage examples for words. We prepared a dataset in the form of (query word, context or example usage) by parsing Context-Reverso webpages based on a query word. Additionally, we trained t5-small, and t5-base models for generating context-sentences based on input words. This resource enables us to obtain diverse and contextually rich sentences that incorporate the target keywords. We have also developed an application for learning new English words with a generated context [Telegram bot]. Our method aims to address the challenges of generating extremely short contexts and mitigating ambiguity in sentence construction. Objective: To develop a model that can generate informative and contextually relevant sentence-contexts for a given set of keywords, benefiting natural language understanding and generation applications such as search engines, personal assistants, and content summarization.


MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection

arXiv.org Artificial Intelligence

The prevalence of fake news across various online sources has had a significant influence on the public. Existing Chinese fake news detection datasets are limited to news sourced solely from Weibo. However, fake news originating from multiple sources exhibits diversity in various aspects, including its content and social context. Methods trained on purely one single news source can hardly be applicable to real-world scenarios. Our pilot experiment demonstrates that the F1 score of the state-of-the-art method that learns from a large Chinese fake news detection dataset, Weibo-21, drops significantly from 0.943 to 0.470 when the test data is changed to multi-source news data, failing to identify more than one-third of the multi-source fake news. To address this limitation, we constructed the first multi-source benchmark dataset for Chinese fake news detection, termed MCFEND, which is composed of news we collected from diverse sources such as social platforms, messaging apps, and traditional online news outlets. Notably, such news has been fact-checked by 14 authoritative fact-checking agencies worldwide. In addition, various existing Chinese fake news detection methods are thoroughly evaluated on our proposed dataset in cross-source, multi-source, and unseen source ways. MCFEND, as a benchmark dataset, aims to advance Chinese fake news detection approaches in real-world scenarios.



Copeland Dueling Bandits Zohar Karnin Informatics Institute

Neural Information Processing Systems

A version of the dueling bandit problem is addressed in which a Condorcet winner may not exist. Two algorithms are proposed that instead seek to minimize regret with respect to the Copeland winner, which, unlike the Condorcet winner, is guaranteed to exist. The first, Copeland Confidence Bound (CCB), is designed for small numbers of arms, while the second, Scalable Copeland Bandits (SCB), works better for large-scale problems. We provide theoretical results bounding the regret accumulated by CCB and SCB, both substantially improving existing results.


Science Checker Reloaded: A Bidirectional Paradigm for Transparency and Logical Reasoning

arXiv.org Artificial Intelligence

Information retrieval is a rapidly evolving field. However it still faces significant limitations in the scientific and industrial vast amounts of information, such as semantic divergence and vocabulary gaps in sparse retrieval, low precision and lack of interpretability in semantic search, or hallucination and outdated information in generative models. In this paper, we introduce a two-block approach to tackle these hurdles for long documents. The first block enhances language understanding in sparse retrieval by query expansion to retrieve relevant documents. The second block deepens the result by providing comprehensive and informative answers to the complex question using only the information spread in the long document, enabling bidirectional engagement. At various stages of the pipeline, intermediate results are presented to users to facilitate understanding of the system's reasoning. We believe this bidirectional approach brings significant advancements in terms of transparency, logical thinking, and comprehensive understanding in the field of scientific information retrieval.


Foundation Models and Information Retrieval in Digital Pathology

arXiv.org Artificial Intelligence

The surge in adoption of digital pathology has the potential to revolutionize medical diagnosis by allowing computerized analysis of tissue images (Pantanowitz 2010; Aljanabi 2012; Hanna2020). Central to this technology is the digitization of formalin-fixed, paraffin-embedded (FFPE) tissue sections mounted on glass slides. This process converts physical tissue samples into high-resolution, gigapixel digital images called whole slide images (WSIs) (Kumar2020; Evans2022). These WSI files contain detailed patterns of tissue morphology, enabling the application of computer-vision algorithms in diagnostic pathology. Pathologists can now analyze tissue images seamlessly on computer screens at various magnifications (Griffin2017). This shift from light microscopes to digital displays allows for easier visual inspection of anatomic clues that may indicate specific diseases.


COSTREAM: Learned Cost Models for Operator Placement in Edge-Cloud Environments

arXiv.org Artificial Intelligence

In this work, we present COSTREAM, a novel learned cost model for Distributed Stream Processing Systems that provides accurate predictions of the execution costs of a streaming query in an edge-cloud environment. The cost model can be used to find an initial placement of operators across heterogeneous hardware, which is particularly important in these environments. In our evaluation, we demonstrate that COSTREAM can produce highly accurate cost estimates for the initial operator placement and even generalize to unseen placements, queries, and hardware. When using COSTREAM to optimize the placements of streaming operators, a median speed-up of around 21x can be achieved compared to baselines.


Flexible Models for with Application to Entity Resolution

Neural Information Processing Systems

Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman-Yor process mixture models make this assumption, as do all other infinitely exchangeable clustering models. However, for some applications, this assumption is inappropriate. For example, when performing entity resolution, the size of each cluster should be unrelated to the size of the data set, and each cluster should contain a negligible fraction of the total number of data points. These applications require models that yield clusters whose sizes grow sublinearly with the size of the data set. We address this requirement by defining the microclustering property and introducing a new class of models that can exhibit this property. We compare models within this class to two commonly used clustering models using four entity-resolution data sets.


SPLADE-v3: New baselines for SPLADE

arXiv.org Artificial Intelligence

A companion to the release of the latest version of the SPLADE library. We describe changes to the training structure and present our latest series of models -- SPLADE-v3. We compare this new version to BM25, SPLADE++, as well as re-rankers, and showcase its effectiveness via a meta-analysis over more than 40 query sets. SPLADE-v3 further pushes the limit of SPLADE models: it is statistically significantly more effective than both BM25 and SPLADE++, while comparing well to cross-encoder re-rankers. Specifically, it gets more than 40 MRR@10 on the MS MARCO dev set, and improves by 2% the out-of-domain results on the BEIR benchmark.