Scientific Discovery
Your Guide to Master Hypothesis Testing in Statistics
This article was written by Sunil Ray. Sunil is a Business Analytics and Intelligence professional with deep experience. I started my career as a MIS professional and then made my way into Business Intelligence (BI) followed by Business Analytics, Statistical modeling and more recently machine learning. Each of these transition has required me to do a change in mind set on how to look at the data. But, one instance sticks out in all these transitions.
Hypothesis Testing in Unsupervised Domain Adaptation with Applications in Alzheimer's Disease
Zhou, Hao, Ithapu, Vamsi K., Ravi, Sathya Narayanan, Singh, Vikas, Wahba, Grace, Johnson, Sterling C.
Consider samples from two different data sources $\{\mathbf{x_s^i}\} \sim P_{\rm source}$ and $\{\mathbf{x_t^i}\} \sim P_{\rm target}$. We only observe their transformed versions $h(\mathbf{x_s^i})$ and $g(\mathbf{x_t^i})$, for some known function class $h(\cdot)$ and $g(\cdot)$. Our goal is to perform a statistical test checking if $P_{\rm source}$ = $P_{\rm target}$ while removing the distortions induced by the transformations. This problem is closely related to concepts underlying numerous domain adaptation algorithms, and in our case, is motivated by the need to combine clinical and imaging based biomarkers from multiple sites and/or batches, where this problem is fairly common and an impediment in the conduct of analyses with much larger sample sizes. We develop a framework that addresses this problem using ideas from hypothesis testing on the transformed measurements, where in the distortions need to be estimated {\it in tandem} with the testing. We derive a simple algorithm and study its convergence and consistency properties in detail, and we also provide lower-bound strategies based on recent work in continuous optimization. On a dataset of individuals at risk for neurological disease, our results are competitive with alternative procedures that are twice as expensive and in some cases operationally infeasible to implement.
Automated And Agile: The New Paradigm For Legal Service
Axiom, a legal staffing-turned-technology company, recently announced a five-year deal with Johnson & Johnson (J & J) to provide multi-shore contract management services to the pharmaceutical giant. Axiom will support J&J's global procurement contracting function, helping to standardize its vast trove of procurement agreements across a dozen contract types and 10 languages. This is not Axiom's lone big dollar, long-term contract with a major corporation. A couple years ago, it inked an eye-popping $73 million deal with Credit Suisse to process the bank's "master trading agreements." Axiom's metamorphosis from staffing to technology is emblematic of the maturing face and changing focus of legal service providers.
ALDI – A New Paradigm for Integrating Marketing Analytics with Data Science
Owing to the data deluge and the Cambrian explosion of machine learning techniques over the past decade, one might have expected the transformation of marketing strategy into a predominantly quantitative discipline by now. The fact that it hasn't happened yet, and the observation that marketing is still influenced by a lot of qualitative inputs can be ascribed to two reasons, in my opinion. The first and principal reason continues to be institutional inertia. Second, there is a significant communication and knowledge gap between data scientists and marketers, owing to their relative lack of familiarity with the other side's perspectives and paradigms. The successful marketer of the next decade is someone who is conversant with management theories of Kotler[1] as well as machine learning advances by Hinton[2]/LeCun[3]/ Ng[4].
Scientific discoveries inspire amid a turbulent 2016
A number of the notable science stories of the past year are, quite literally, out of this world. For me, the story of the year has to be the August discovery of an Earth-like planet orbiting the closest star to our own. The star, Proxima Centauri, is just 4.2 light-years from Earth. The planet circling that star has been named Proxima Centauri b. Proxima Centauri b was discovered by astronomers working on a project called Pale Red Dot, who reported that the planet lies in the star's habitable zone, meaning that it could possess water and, maybe, life.
The Bayesian New Statistics: Hypothesis Testing, Estimation, Meta-Analysis, and Power Analysis from a Bayesian Perspective
Many people have found the table above to be useful for understanding two conceptual distinctions in the practice of data analysis. The article that discusses the table, and many other issues, is now in press. The in-press version can be found at OSF and at SSRN. Abstract: In the practice of data analysis, there is a conceptual distinction between hypothesis testing, on the one hand, and estimation with quantified uncertainty, on the other hand. Among frequentists in psychology a shift of emphasis from hypothesis testing to estimation has been dubbed "the New Statistics" (Cumming, 2014).
This week's popular Artificial Intelligence news: December 15, 2016
Its been a busy week for . Here is our round up of the most popular articles. Kuhn's book The Structure of Scientific Revolution outlined an episodic model in which periods of "normal science" were interrupted by periods of "revolutionary science". Kuhn challenges us to consider new paradigms and to change the rules of the game, our standards and our best practices.Artificial intelligence (AI) and machine learning (ML) liberatingly delivers this new paradigm, putting the science back into security.It's quite clear that relying on endpoint protection solutions that only… Tagged In Computer Security Artificial Intelligence Machine Learning Malware Data Mining Dam High Frequency Trading Scientific Revolution CEVA creates new value by enhancing IoT and machine learning applications Tagged In Smartphone Compound Annual Growth Rate Artificial Intelligence Bluetooth NASDAQ Soft Bank LTE (telecommunication) Integrated Circuit Data Compression Big Data Machine Learning Artificial Neural Network Computer Vision ARM Architecture Ericsson Wilderness Software Framework Embedded System Session Initiation Protocol 3GPP Zig Bee Digital Signal Processing To get the right data, we could look at two different kinds of inputs: Explicit and implicit. Explicit means asking the user to provide information by asking straightforward questions such as, How much do you want to spend?
Characteristics of Good Visual Analytics and Data Discovery Tools
Visual Analytics and Data Discovery allow analysis of big data sets to find insights and valuable information. This is much more than just classical Business Intelligence (BI). See this article for more details and motivation: "Using Visual Analytics to Make Better Decisions: the Death Pill Exa...". Let's take a look at important characteristics to choose the right tool for your use cases. Several tools are available on the market for Visual Analytics and Data Discovery.
Nonparametric Detection of Anomalous Data Streams
Zou, Shaofeng, Liang, Yingbin, Poor, H. Vincent, Shi, Xinghua
A nonparametric anomalous hypothesis testing problem is investigated, in which there are totally n sequences with s anomalous sequences to be detected. Each typical sequence contains m independent and identically distributed (i.i.d.) samples drawn from a distribution p, whereas each anomalous sequence contains m i.i.d. samples drawn from a distribution q that is distinct from p. The distributions p and q are assumed to be unknown in advance. Distribution-free tests are constructed using maximum mean discrepancy as the metric, which is based on mean embeddings of distributions into a reproducing kernel Hilbert space. The probability of error is bounded as a function of the sample size m, the number s of anomalous sequences and the number n of sequences. It is then shown that with s known, the constructed test is exponentially consistent if m is greater than a constant factor of log n, for any p and q, whereas with s unknown, m should has an order strictly greater than log n. Furthermore, it is shown that no test can be consistent for arbitrary p and q if m is less than a constant factor of log n, thus the order-level optimality of the proposed test is established. Numerical results are provided to demonstrate that our tests outperform (or perform as well as) the tests based on other competitive approaches under various cases.
Key-Object – A New Paradigm in Search?
As we are all fond of saying, innovation follows pain points. Are we missing something in our uber-critical search capabilities that needs to be resolved? A colleague recently pointed me to a slim volume "Structured Search for Big Data" by Mikhail Gilula (published by Elsevier and available on Amazon) that argues that not only are our search tools deficient but that a complete revamp of the underlying key-word NoSQL DB structure is what's required. Use Google, Amazon, or any of the other life-critical search tools we've become so reliant upon and you are using key-word search on NoSQL. The pain that Gilula identifies is the length of time it takes the consumer to research and select complex merchandise for best deals resulting from the imprecision of the search results.