Information Retrieval
Adaptive Cardinality Estimation
Ivanov, Oleg, Bartunov, Sergey
In this paper we address cardinality estimation problem which is an important subproblem in query optimization. Query optimization is a part of every relational DBMS responsible for finding the best way of the execution for the given query. These ways are called plans. The execution time of different plans may differ by several orders, so query optimizer has a great influence on the whole DBMS performance. We consider cost-based query optimization approach as the most popular one. It was observed that cost-based optimization quality depends much on cardinality estimation quality. Cardinality of the plan node is the number of tuples returned by it. In the paper we propose a novel cardinality estimation approach with the use of machine learning methods. The main point of the approach is using query execution statistics of the previously executed queries to improve cardinality estimations. We called this approach adaptive cardinality estimation to reflect this point. The approach is general, flexible, and easy to implement. The experimental evaluation shows that this approach significantly increases the quality of cardinality estimation, and therefore increases the DBMS performance for some queries by several times or even by several dozens of times.
How Fruit Fly Brains Are Improving Smart Phone Apps
What do a fruit fly and a search engine have in common? Search engine algorithms go through great pains to match items you've clicked on or purchased, songs you've listened to, or things searched for, to similar ones. As a result, we constantly need ever faster and more efficient search engines, and so computer scientists must work tirelessly to keep up. They have to constantly tackle what they call "a fundamental machine learning problem: approximate similarity (or nearest-neighbors) search." Turns out, fruit fly brains go through a similar matching process, and the way they do it is fast, efficient, and dare I say, elegant.
8 Ways to Measure Social with Google Analytics - Search Engine Journal
The maturity and widespread acceptance of social media marketing, combined with the expectation of being able to track everything in the era of big data, has created a lot of expectations. It has also raised deeper questions about performance. Using Google Analytics, we have the power to go deeper to prove the impact and value of our marketing efforts. We should never start the answer to a question months into a social media campaign with "I think" when we have the capacity to know the impact of digital marketing activities for sure. Google Analytics can be a great source of deeper insights for the social media marketer.
A visual search engine for Bangladeshi laws
Mandal, Manash Kumar, Nath, Pinku Deb, Mizan, Arpeeta Shams, Saquib, Nazmus
Browsing and finding relevant information for Bangladeshi laws is a challenge faced by all law students and researchers in Bangladesh, and by citizens who want to learn about any legal procedure. Some law archives in Bangladesh are digitized, but lack proper tools to organize the data meaningfully. We present a text visualization tool that utilizes machine learning techniques to make the searching of laws quicker and easier. Using Doc2Vec to layout law article nodes, link mining techniques to visualize relevant citation networks, and named entity recognition to quickly find relevant sections in long law articles, our tool provides a faster and better search experience to the users. Qualitative feedback from law researchers, students, and government officials show promise for visually intuitive search tools in the context of governmental, legal, and constitutional data in developing countries, where digitized data does not necessarily pave the way towards an easy access to information.
Information Retrieval Document Search Engine in R
In this post, we learn about building a basic search engine or document retrieval system using Vector space model. This use case is widely used in information retrieval systems. Given a set of documents and search term(s)/query we need to retrieve relevant documents that are similar to the search query.
Fruit Fly Brain Patterns Can Improve Algorithms that Power Netflix, Youtube Recommendations
Researchers have ventured into uncharted territory to find ways to improve computer algorithms -- the brains of fruit flies. While search algorithms work by analyzing users' previous searches, a fruit fly searches for fruits by remembering the odor of the fruit they have fed on. "This is a problem that pretty much every technology company with any kind of information retrieval system has to solve, so it's been something that computer scientists have studied for years. Now, we have this new approach to similarity searches thanks to the fly," said Saket Navlakha, assistant professor at Salk's Integrative Biology Laboratory and lead author of the research paper titled "A neural algorithm for a fundamental computing problem." The paper was published in the Science Journal on Thursday.
A neural algorithm for a fundamental computing problem
Similarity search--for example, identifying similar images in a database or similar documents on the web--is a fundamental computing problem faced by large-scale information retrieval systems. We discovered that the fruit fly olfactory circuit solves this problem with a variant of a computer science algorithm (called locality-sensitive hashing). The fly circuit assigns similar neural activity patterns to similar odors, so that behaviors learned from one odor can be applied when a similar odor is experienced. The fly algorithm, however, uses three computational strategies that depart from traditional approaches. These strategies can be translated to improve the performance of computational similarity searches.
5 Ways Machine Learning Can Improve Access to Enterprise Data - insideBIGDATA
In this special guest feature, Grant Ingersoll, Founder and CTO of Lucidworks, discusses how machine learning is helping companies manage big data and make sense of it for their customers and employees. With smarter search tools, business leaders can more quickly retrieve information and deliver a better user experience for customers. Here are 5 ways that machine learning is powering more intuitive enterprise search. Grant is an active member of the Lucene community. He is a Lucene and Solr committer, co-founder of the Apache Mahout machine learning project, and a longstanding member of the Apache Software Foundation. Grant's prior experience includes work at the Center for Natural Language Processing at Syracuse University in natural language processing and information retrieval.
Schema Independent Relational Learning
Picado, Jose, Termehchy, Arash, Fern, Alan, Ataei, Parisa
Learning novel concepts and relations from relational databases is an important problem with many applications in database systems and machine learning. Relational learning algorithms learn the definition of a new relation in terms of existing relations in the database. Nevertheless, the same data set may be represented under different schemas for various reasons, such as efficiency, data quality, and usability. Unfortunately, the output of current relational learning algorithms tends to vary quite substantially over the choice of schema, both in terms of learning accuracy and efficiency. This variation complicates their off-the-shelf application. In this paper, we introduce and formalize the property of schema independence of relational learning algorithms, and study both the theoretical and empirical dependence of existing algorithms on the common class of (de) composition schema transformations. We study both sample-based learning algorithms, which learn from sets of labeled examples, and query-based algorithms, which learn by asking queries to an oracle. We prove that current relational learning algorithms are generally not schema independent. For query-based learning algorithms we show that the (de) composition transformations influence their query complexity. We propose Castor, a sample-based relational learning algorithm that achieves schema independence by leveraging data dependencies. We support the theoretical results with an empirical study that demonstrates the schema dependence/independence of several algorithms on existing benchmark and real-world datasets under (de) compositions.