Information Retrieval
Why killing your content marketing makes the most sense - Search Engine Watch
The problem is, simply put, out of control. Just because a company or individual can create and distribute content on a platform, doesn't mean they should. I've had the opportunity to analyze content marketing strategies from huge brands, desperately trying to build audiences online leveraging content marketing. In almost every case, each one made the same mistake. When an organization decides to fund a content marketing strategy, the initial stages are always exciting.
A Unified Transferable Model for ML-Enhanced DBMS
Wu, Ziniu, Yang, Peilun, Yu, Pei, Zhu, Rong, Han, Yuxing, Li, Yaliang, Lian, Defu, Zeng, Kai, Zhou, Jingren
Recently, the database management system (DBMS) community has witnessed the power of machine learning (ML) solutions for DBMS tasks. Despite their promising performance, these existing solutions can hardly be considered satisfactory. First, these ML-based methods in DBMS are not effective enough because they are optimized on each specific task, and cannot explore or understand the intrinsic connections between tasks. Second, the training process has serious limitations that hinder their practicality, because they need to retrain the entire model from scratch for a new DB. Moreover, for each retraining, they require an excessive amount of training data, which is very expensive to acquire and unavailable for a new DB. We propose to explore the transferabilities of the ML methods both across tasks and across DBs to tackle these fundamental drawbacks. In this paper, we propose a unified model MTMLF that uses a multi-task training procedure to capture the transferable knowledge across tasks and a pretrain finetune procedure to distill the transferable meta knowledge across DBs. We believe this paradigm is more suitable for cloud DB service, and has the potential to revolutionize the way how ML is used in DBMS. Furthermore, to demonstrate the predicting power and viability of MTMLF, we provide a concrete and very promising case study on query optimization tasks. Last but not least, we discuss several concrete research opportunities along this line of work.
Retrieving Complex Tables with Multi-Granular Graph Representation Learning
Wang, Fei, Sun, Kexuan, Chen, Muhao, Pujara, Jay, Szekely, Pedro
The task of natural language table retrieval (NLTR) seeks to retrieve semantically relevant tables based on natural language queries. Existing learning systems for this task often treat tables as plain text based on the assumption that tables are structured as dataframes. However, tables can have complex layouts which indicate diverse dependencies between subtable structures, such as nested headers. As a result, queries may refer to different spans of relevant content that is distributed across these structures. Moreover, such systems fail to generalize to novel scenarios beyond those seen in the training set. Prior methods are still distant from a generalizable solution to the NLTR problem, as they fall short in handling complex table layouts or queries over multiple granularities. To address these issues, we propose Graph-based Table Retrieval (GTR), a generalizable NLTR framework with multi-granular graph representation learning. In our framework, a table is first converted into a tabular graph, with cell nodes, row nodes and column nodes to capture content at different granularities. Then the tabular graph is input to a Graph Transformer model that can capture both table cell content and the layout structures. To enhance the robustness and generalizability of the model, we further incorporate a self-supervised pre-training task based on graph-context matching. Experimental results on two benchmarks show that our method leads to significant improvements over the current state-of-the-art systems. Further experiments demonstrate promising performance of our method on cross-dataset generalization, and enhanced capability of handling complex tables and fulfilling diverse query intents. Code and data are available at https://github.com/FeiWang96/GTR.
One Model to Rule them All: Towards Zero-Shot Learning for Databases
Hilprecht, Benjamin, Binnig, Carsten
And unfortunately, the training data collection needs to be repeated for every new database that needs to be supported. In this paper, we present our vision of so called zero-shot learning To reduce the high cost of training data collection, reinforcement for databases which is a new learning approach for database learning (RL) has been used to execute training queries [10, 17, 18, components. Zero-shot learning for databases is inspired by recent 34] in a more targeted manner (i.e., letting the RL agent decide advances in transfer learning of models such as GPT-3 and can which queries to execute next). However, even with reinforcement support a new database out-of-the box without the need to train a learning still a large amount of training queries needs to be executed new model. As a first concrete contribution in this paper, we show for learning a model. Moreover, training the model is not a onetime the feasibility of zero-shot learning for the task of physical cost effort since similar to workload-driven approaches the learning estimation and present very promising initial results. Moreover, procedure needs to be repeated for every new database at hand. as a second contribution we discuss the core challenges related to A different direction that has thus been proposed to avoid the zero-shot learning for databases and present a roadmap to extend expensive training data collection by running queries on a new zero-shot learning towards many other tasks beyond cost estimation database are so called data-driven approaches [11, 31, 32] that learn or even beyond classical database systems and workloads.
MathBERT: A Pre-Trained Model for Mathematical Formula Understanding
Peng, Shuai, Yuan, Ke, Gao, Liangcai, Tang, Zhi
Large-scale pre-trained models like BERT, have obtained a great success in various Natural Language Processing (NLP) tasks, while it is still a challenge to adapt them to the math-related tasks. Current pre-trained models neglect the structural features and the semantic correspondence between formula and its context. To address these issues, we propose a novel pre-trained model, namely \textbf{MathBERT}, which is jointly trained with mathematical formulas and their corresponding contexts. In addition, in order to further capture the semantic-level structural features of formulas, a new pre-training task is designed to predict the masked formula substructures extracted from the Operator Tree (OPT), which is the semantic structural representation of formulas. We conduct various experiments on three downstream tasks to evaluate the performance of MathBERT, including mathematical information retrieval, formula topic classification and formula headline generation. Experimental results demonstrate that MathBERT significantly outperforms existing methods on all those three tasks. Moreover, we qualitatively show that this pre-trained model effectively captures the semantic-level structural information of formulas. To the best of our knowledge, MathBERT is the first pre-trained model for mathematical formula understanding.
Natural Language in Search Engine Optimization (SEO) -- How, What, When, And Why
For information to be processed, it is necessary to understand the data behind it by going into the essence of such data [1]. When it comes to natural language processing, the "what" of natural language is discussed. For instance, we can describe a native or evolved language as a natural language. Consequently, we can think of any spoken language as a natural language. We can use natural language to describe ordinary non-artificial speaking and writing language for natural language.
IITP in COLIEE@ICAIL 2019: Legal Information Retrieval using BM25 and BERT
Gain, Baban, Bandyopadhyay, Dibyanayan, Saikh, Tanik, Ekbal, Asif
Natural Language Processing (NLP) and Information Retrieval (IR) in the judicial domain is an essential task. With the advent of availability domain-specific data in electronic form and aid of different Artificial intelligence (AI) technologies, automated language processing becomes more comfortable, and hence it becomes feasible for researchers and developers to provide various automated tools to the legal community to reduce human burden. The Competition on Legal Information Extraction/Entailment (COLIEE-2019) run in association with the International Conference on Artificial Intelligence and Law (ICAIL)-2019 has come up with few challenging tasks. The shared defined four sub-tasks (i.e. Task1, Task2, Task3 and Task4), which will be able to provide few automated systems to the judicial system. The paper presents our working note on the experiments carried out as a part of our participation in all the sub-tasks defined in this shared task. We make use of different Information Retrieval(IR) and deep learning based approaches to tackle these problems. We obtain encouraging results in all these four sub-tasks.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Thakur, Nandan, Reimers, Nils, Rücklé, Andreas, Srivastava, Abhishek, Gurevych, Iryna
Neural IR models have often been studied in homogeneous and narrow settings, which has considerably limited insights into their generalization capabilities. To address this, and to allow researchers to more broadly establish the effectiveness of their models, we introduce BEIR (Benchmarking IR), a heterogeneous benchmark for information retrieval. We leverage a careful selection of 17 datasets for evaluation spanning diverse retrieval tasks including open-domain datasets as well as narrow expert domains. We study the effectiveness of nine state-of-the-art retrieval models in a zero-shot evaluation setup on BEIR, finding that performing well consistently across all datasets is challenging. Our results show BM25 is a robust baseline and Reranking-based models overall achieve the best zero-shot performances, however, at high computational costs. In contrast, Dense-retrieval models are computationally more efficient but often underperform other approaches, highlighting the considerable room for improvement in their generalization capabilities. In this work, we extensively analyze different retrieval models and provide several suggestions that we believe may be useful for future work. BEIR datasets and code are available at https://github.com/UKPLab/beir.
[P] Entity Embed: fuzzy and scalable Entity Resolution using Approximate Nearest Neighbors
Entity Embed is based on and is a special case of the AutoBlock model described by Amazon. It allows you to transform entities like companies, products, etc. into vectors to support scalable Record Linkage / Entity Resolution using Approximate Nearest Neighbors. Using Entity Embed, you can train a deep learning model to transform records into vectors in an N-dimensional embedding space. Thanks to a contrastive loss, those vectors are organized to keep similar records close and dissimilar records far apart in this embedding space. Embedding records enables scalable ANN search, which means finding thousands of candidate duplicate pairs of records per second per CPU.
How to Get the Most out of Excel with Machine Learning
Excel is perhaps the most well known data analysis tool out there. It's used to store and organize data such as sales numbers, profit rates, expenditures or revenues. Some businesses even use it to store text data. However, Excel is unable to organize text data without the help of machine learning. Machine learning algorithms can automatically analyze hundreds and thousands of rows of text data in a fast, consistent and scalable way. In other words, machine learning algorithms are able to quantify words and phrases in Excel, by assigning topics, keywords, entities, and even sentiment to each row of text.