Goto

Collaborating Authors

 Query Processing


Representing Social Networks as Dynamic Heterogeneous Graphs

arXiv.org Artificial Intelligence

Graph representations for real-world social networks in the past have missed two important elements: the multiplexity of connections as well as representing time. To this end, in this paper, we present a new dynamic heterogeneous graph representation for social networks which includes time in every single component of the graph, i.e., nodes and edges, each of different types that captures heterogeneity. We illustrate the power of this representation by presenting four time-dependent queries and deep learning problems that cannot easily be handled in conventional homogeneous graph representations commonly used. As a proof of concept we present a detailed representation of a new social media platform (Steemit), which we use to illustrate both the dynamic querying capability as well as prediction tasks using graph neural networks (GNNs). The results illustrate the power of the dynamic heterogeneous graph representation to model social networks. Given that this is a relatively understudied area we also illustrate opportunities for future work in query optimization as well as new dynamic prediction tasks on heterogeneous graph structures.


Share the Tensor Tea: How Databases can Leverage the Machine Learning Ecosystem

arXiv.org Artificial Intelligence

We demonstrate Tensor Query Processor (TQP): a query processor that automatically compiles relational operators into tensor programs. By leveraging tensor runtimes such as PyTorch, TQP is able to: (1) integrate with ML tools (e.g., Pandas for data ingestion, Tensorboard for visualization); (2) target different hardware (e.g., CPU, GPU) and software (e.g., browser) backends; and (3) end-to-end accelerate queries containing both relational and ML operators. TQP is generic enough to support the TPC-H benchmark, and it provides performance that is comparable to, and often better than, that of specialized CPU and GPU query processors.


Merchandise Recommendation for Retail Events with Word Embedding Weighted Tf-idf and Dynamic Query Expansion

arXiv.org Artificial Intelligence

We rank all we rely on item retrieval from marketplace inventory. With retrieved items by the sum of tf-idf scores from matched words, feedback to expand query scope, we discuss keyword expansion and keep the items with total tf-idf scores above a threshold. The candidate selection using word embedding similarity, and an retrieval based system works well to discover relevant enhanced tf-idf formula for expanded words in search ranking.


Multi-agent Databases via Independent Learning

arXiv.org Artificial Intelligence

Machine learning is rapidly being used in database research to improve the effectiveness of numerous tasks included but not limited to query optimization, workload scheduling, physical design, etc. Currently, the research focus has been on replacing a single database component responsible for one task by its learning-based counterpart. However, query performance is not simply determined by the performance of a single component, but by the cooperation of multiple ones. As such, learning based database components need to collaborate during both training and execution in order to develop policies that meet end performance goals. Thus, the paper attempts to address the question "Is it possible to design a database consisting of various learned components that cooperatively work to improve end-to-end query latency?". To answer this question, we introduce MADB (Multi-Agent DB), a proof-of-concept system that incorporates a learned query scheduler and a learned query optimizer. MADB leverages a cooperative multi-agent reinforcement learning approach that allows the two components to exchange the context of their decisions with each other and collaboratively work towards reducing the query latency. Preliminary results demonstrate that MADB can outperform the non-cooperative integration of learned components.


Buffer Pool Aware Query Scheduling via Deep Reinforcement Learning

arXiv.org Artificial Intelligence

One could imagine many simple heuristics, query scheduling with the explicit goal of reducing disk reads such as greedily selecting the next query with the highest and thus implicitly increasing query performance. We introduce expected buffer usage, to solve this problem. However, a SmartQueue, a learned scheduler that leverages overlapping hand-designed policy to handle the complexity of the entire data reads among incoming queries and learns a problem, including different buffer sizes, shifting query scheduling strategy that improves cache hits. SmartQueue workloads, heterogeneous data types (e.g., index files vs base relies on deep reinforcement learning to produce workloadspecific relations), and balancing short-term gains against long-term scheduling strategies that focus on long-term performance strategy is much more difficult to conceive.


Flutter/XCode - iOS App Retailer Join Operation Error - Channel969

#artificialintelligence

I've made some fundamental modifications to my app which I've distributed via the Archive methodology in Xcode a number of instances earlier than. Nevertheless it isn't working this night as I'm offered with the beneath error: I've ran flutter construct iOS –release with no warnings or errors earlier than making an attempt to Archive in Xcode. I've additionally tried doing a flutter clear too. After I run Validate on the archive earlier than making an attempt to distribute the bundle it comes again with 0 errors. It is simply when I attempt to Distribute the bundle, it comes again with the above error and I am undecided why and even how one can go about diagnosing it. Can anybody please assist level me in the appropriate course?


4Bn rows/sec query benchmark: Clickhouse vs QuestDB vs Timescale

#artificialintelligence

QuestDB 6.2, our previous minor version release, introduced JIT (Just-in-Time) compiler for SQL filters. As we mentioned last time, the next step would be to parallelize the query execution when suitable to improve the execution time even further and that's what we're going to discuss and benchmark today. QuestDB 6.3 enables JIT compiled filters by default and, what's even more noticeable, includes parallel SQL filter execution optimization allowing us to reduce both cold and hot query execution times quite dramatically. Prior to diving into the implementation details and running some before/after benchmarks for QuestDB, we'll be having a friendly competition with two popular time series and analytical databases, TimescaleDB and ClickHouse. The purpose of the competition is nothing more but an attempt to understand whether our parallel filter execution is worth the hassle or not.


Distributed Reconstruction of Noisy Pooled Data

arXiv.org Machine Learning

In the pooled data problem we are given a set of $n$ agents, each of which holds a hidden state bit, either $0$ or $1$. A querying procedure returns for a query set the sum of the states of the queried agents. The goal is to reconstruct the states using as few queries as possible. In this paper we consider two noise models for the pooled data problem. In the noisy channel model, the result for each agent flips with a certain probability. In the noisy query model, each query result is subject to random Gaussian noise. Our results are twofold. First, we present and analyze for both error models a simple and efficient distributed algorithm that reconstructs the initial states in a greedy fashion. Our novel analysis pins down the range of error probabilities and distributions for which our algorithm reconstructs the exact initial states with high probability. Secondly, we present simulation results of our algorithm and compare its performance with approximate message passing (AMP) algorithms that are conjectured to be optimal in a number of related problems.


8 Best SQL Courses on Coursera

#artificialintelligence

If you want to gain the skills necessary to query big data with modern distributed SQL engines, then this specialization is for you. The best part of this course is that it will teach you a newer breed of SQL engine: distributed query engines Hive and Impala. Hive and Impala are open-source SQL engines capable of querying enormous datasets. Another advantage of this specialization program is that this program provides excellent preparation for the Cloudera Certified Associate (CCA) Data Analyst certification exam. This Specialization program consists of 3 Courses.


OCR quality affects perceived usefulness of historical newspaper clippings -- a user study

arXiv.org Artificial Intelligence

Effects of Optical Character Recognition (OCR) quality on historical information retrieval have so far been studied in data-oriented scenarios regarding the effectiveness of retrieval results. Such studies have either focused on the effects of artificially degraded OCR quality (see, e.g., [1-2]) or utilized test collections containing texts based on authentic low quality OCR data (see, e.g., [3]). In this paper the effects of OCR quality are studied in a user-oriented information retrieval setting. Thirty-two users evaluated subjectively query results of six topics each (out of 30 topics) based on pre-formulated queries using a simulated work task setting. To the best of our knowledge our simulated work task experiment is the first one showing empirically that users' subjective relevance assessments of retrieved documents are affected by a change in the quality of optically read text. Users of historical newspaper collections have so far commented effects of OCR'ed data quality mainly in impressionistic ways, and controlled user environments for studying effects of OCR quality on users' relevance assessments of the retrieval results have so far been missing. To remedy this The National Library of Finland (NLF) set up an experimental query environment for the contents of one Finnish historical newspaper, Uusi Suometar 1869-1918, to be able to compare users' evaluation of search results of two different OCR qualities for digitized newspaper articles. The query interface was able to present the same underlying document for the user based on two alternatives: either based on the lower OCR quality, or based on the higher OCR quality, and the choice was randomized. The users did not know about quality differences in the article texts they evaluated. The main result of the study is that improved optical character recognition quality affects perceived usefulness of historical newspaper articles significantly. The mean average evaluation score for the improved OCR results was 7.94% higher than the mean average evaluation score of the old OCR results.