Goto

Collaborating Authors

 Scientific Discovery


Predicting customer lifecycle outcomes with machine learning

#artificialintelligence

In our last article, Lifecycle mapping: uncovering rich, predictive data sources, we discussed the importance of mapping out your customer lifecycle to better understand where your most predictive customer data is hiding. Lifecycle mapping is the first step to using artificial intelligence (AI) to optimize your customer lifecycle marketing initiatives. Now, we'll pose some questions to help identify your predictive customer attributes and lifecycle events, pinpoint where that data is located, and recognize patterns to predict outcomes for future prospects, leads, and customers. Data discovery is the second stage in the customer lifecycle optimization (CLO) process. The primary task of this stage is to expand on your lifecycle map to identify authoritative data sources that establish progress.


The alien-hunting Kepler telescope has discovered something big

#artificialintelligence

NASA has called a press conference to reveal a breakthrough discovery from its alien-hunting Kepler telescope. The discovery was driven by Google's machine-learning artificial intelligence software. The announcement will be live-streamed on NASA's website, according to a press release. It will take place Thursday, December 14, at 1 p.m. EST. NASA's Kepler space telescope has been searching for habitable planets since 2009.


Hypothesis Testing for High-Dimensional Multinomials: A Selective Review

arXiv.org Machine Learning

The statistical analysis of discrete data has been the subject of extensive statistical research dating back to the work of Pearson. In this survey we review some recently developed methods for testing hypotheses about high-dimensional multinomials. Traditional tests like the $\chi^2$ test and the likelihood ratio test can have poor power in the high-dimensional setting. Much of the research in this area has focused on finding tests with asymptotically Normal limits and developing (stringent) conditions under which tests have Normal limits. We argue that this perspective suffers from a significant deficiency: it can exclude many high-dimensional cases when - despite having non Normal null distributions - carefully designed tests can have high power. Finally, we illustrate that taking a minimax perspective and considering refinements of this perspective can lead naturally to powerful and practical tests.


Your Guide to Master Hypothesis Testing in Statistics

@machinelearnbot

I started my career as a MIS professional and then made my way into Business Intelligence (BI) followed by Business Analytics, Statistical modeling and more recently machine learning. Each of these transition has required me to do a change in mind set on how to look at the data. But, one instance sticks out in all these transitions. This was when I was working as a BI professional creating management dashboards and reports. Due to some internal structural changes in the Organization I was working with, our team had to start reporting to a team of Business Analysts (BA).


MicroStrategy 10.10 Empowers Enterprises with Massive Update to Data Discovery - DATAVERSITY

@machinelearnbot

According to a new press release, "MicroStrategy Incorporated, a leading worldwide provider of enterprise analytics and mobility software, today announced the general availability of MicroStrategy 10.10, the newest feature release to the company's MicroStrategy 10 platform. This feature release empowers business teams to confidently embrace an enterprise-wide, data-driven culture by introducing two exciting products -- a completely redesigned and more powerful MicroStrategy Desktop and the new MicroStrategy Workstation. 'We are incredibly excited to release MicroStrategy 10.10, which empowers business teams to confidently author, promote and certify analytics content, operationalize dossiers, and deliver the agility a business needs, along with the governance that IT requires,' said Tim Lang, Senior Executive Vice President and Chief Technology Officer, MicroStrategy Incorporated. 'The latest capabilities in MicroStrategy 10.10 are part of MicroStrategy's commitment to deliver the next generation of enterprise analytics to our customers so they can discover growth opportunities, solve complex business problems, and drive real results'."


Listen To Whistler Waves NASA Recorded From Space

International Business Times

Researches have made a breakthrough discovery about the impulsive electron loss that happens in the Earth's upper atmosphere. A paper on the research was published in the Geophysical Review Letters on Wednesday and details the scientific discoveries two spacecraft made about the loss and its cause, according to NASA. The Cubesat FIREBIRD II was one of those craft that recorded the electron microburst when it happened. The craft observed the microbursts from its place orbiting 310 miles above Earth while one of the Van Allen Probes that orbits a bit higher up was able to capture a rising-tone lower band chorus. That chorus of waves had the duration and cadence highly similar to those of the microburst that the FIREBIRD had captured.


Kernel Two-Sample Hypothesis Testing Using Kernel Set Classification

arXiv.org Machine Learning

The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification problem. In the set classification problem the goal is to classify a set of data points that are assumed to have a common class. We prove that the average probability of error given a set is less than or equal to the Bayes error and decreases as a power of $n$ number of sample data points in the set. We use the positive definite Set Kernel for directly mapping sets of data to an associated Reproducing Kernel Hilbert Space, without the need to learn a probability distribution. We specifically solve the two-sample hypothesis testing problem using a one-class SVM in conjunction with the proposed Set Kernel. We compare the proposed method with the Maximum Mean Discrepancy, F-Test and T-Test methods on a number of challenging simulated high dimensional and small sample size data. We also perform two-sample hypothesis testing experiments on six cancer gene expression data sets and achieve zero type-I and type-II error results on all data sets.


Theory-guided Data Science: A New Paradigm for Scientific Discovery from Data

arXiv.org Artificial Intelligence

Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of data science models in enabling scientific discovery. The overarching vision of TGDS is to introduce scientific consistency as an essential component for learning generalizable models. Further, by producing scientifically interpretable models, TGDS aims to advance our scientific understanding by discovering novel domain insights. Indeed, the paradigm of TGDS has started to gain prominence in a number of scientific disciplines such as turbulence modeling, material discovery, quantum chemistry, bio-medical science, bio-marker discovery, climate science, and hydrology. In this paper, we formally conceptualize the paradigm of TGDS and present a taxonomy of research themes in TGDS. We describe several approaches for integrating domain knowledge in different research themes using illustrative examples from different disciplines. We also highlight some of the promising avenues of novel research for realizing the full potential of theory-guided data science.


Trend Analysis of Fragmented Time Series: Hypothesis Testing Based Adaptive Spline Filtering Method

#artificialintelligence

Missing data present significant challenges to trend analysis of time series. Straightforward approaches consisting of supplementing missing data with constant or zero values or with linear trends can severely degrade the quality of the trend analysis, which significantly reduces the reliability of the trend analysis. We present a robust adaptive approach to discover the trends from fragmented time series. The approach proposed in this paper is based on the HASF (Hypothesis-testing-based Adaptive Spline Filtering) trend analysis algorithm, which can accommodate non-uniform sampling and is therefore inherently robust to missing data. HASF adapts the nodes of the spline based on hypothesis testing and variance minimization, which adds to its robustness.


Data Science- Hypothesis Testing Using Minitab and R

@machinelearnbot

Formulating the Null and the alternate hypothesis for normality test; Choice of null hypothesis based on absence of action and the vice versa for alternate hypothesis; checking for normality in Minitab; interpreting the Q–Q plot; Comparing the computed'p' value with α (alpha) for taking the decision on whether or not to take the action; Step to performing the 1 sample Z test, selection of appropriate hypothesis in minitab.