Scientific Discovery
Combining Multiple Hypothesis Testing with Machine Learning Increases the Statistical Power of Genome-wide Association Studies
The goal of genome-wide association studies (GWAS) (e.g. the WTCCC study1) is to examine the relationship between genetic markers such as single-nucleotide polymorphisms (SNPs) and individual traits, which are usually complex diseases or behavioral characteristics. Generally, a large number of statistical tests are performed in parallel, each SNP being individually tested for association2,3,4. The standard approach consists of computing individual, SNP-specific p-values corresponding to a statistical association test and comparing these p-values against some given significance threshold (say t*), meaning that precisely those SNPs with p-values smaller than t*are declared to be associated with the trait4,5,6. We refer to this approach as raw p-value thresholding (RPVT) and review some standard methods for choosing t*for the purpose of controlling multiple type I error rates (in particular, the family-wise error rate (FWER) and the expected number of false rejections (ENFR)) in the Methods Section. According to the GWAS catalog7,8 (last accessed 03-07-2015), the more than 1,400 GWAS published so far have led to the identification of more than 11,000 SNPs associated with about 800 human diseases and anthropometric traits with p-values using t* 1 10 5.
Key-Object – A New Paradigm in Search?
As we are all fond of saying, innovation follows pain points. Are we missing something in our uber-critical search capabilities that needs to be resolved? A colleague recently pointed me to a slim volume "Structured Search for Big Data" by Mikhail Gilula (published by Elsevier and available on Amazon) that argues that not only are our search tools deficient but that a complete revamp of the underlying key-word NoSQL DB structure is what's required. Use Google, Amazon, or any of the other life-critical search tools we've become so reliant upon and you are using key-word search on NoSQL. The pain that Gilula identifies is the length of time it takes the consumer to research and select complex merchandise for best deals resulting from the imprecision of the search results.
From America to Viagra: the art of finding what you're not looking for
STOCKHOLM – It is serendipity: from America to Viagra, history is full of great discoveries helped along by chance, as more than a century of Nobel prizes can attest. Among the chance discoveries that have been honored with the prestigious prize are X-rays (physics, 1901), penicillin (medicine, 1945), fullerenes that paved the way for nanotechnology (chemistry, 1996), conductive polymers (chemistry, 2000), and the bacteria responsible for ulcers (medicine, 2005). But, as the father of pasteurization Louis Pasteur noted in 1854, "In the fields of observation, chance only favors the prepared mind" -- a remark made in reference to the discovery of the link between electricity and magnetism by Danish scientist Hans Christian Orsted. Orsted happened to notice that a compass needle deflected from magnetic north when an electric current from a battery was switched on and off -- a pioneering discovery in electromagnetism. Like Pasteur, Dutch scientist Pek Van Andel also believes in the unexpected.
Hypothesis Testing is a Bad Idea (my talk at Warwick, England, 2:30pm Thurs 15 Sept)
This is the conference, and here's my talk (will do Google hangout, just as with my recent talks in Bern, Strasbourg, etc): Through a series of examples, we consider problems with classical hypothesis testing, whether performed using classical p-values or confidence intervals, Bayes factors, or Bayesian inference using noninformative priors. We locate the problem not in the use of any particular statistical method but rather with larger problems of deterministic thinking and a misguided version of Popperianism in which the rejection of a straw-man null hypothesis is taken as confirmation of a preferred alternative. We suggest solutions involving multilevel modeling and informative Bayesian inference. The post Hypothesis Testing is a Bad Idea (my talk at Warwick, England, 2:30pm Thurs 15 Sept) appeared first on Statistical Modeling, Causal Inference, and Social Science. The post Hypothesis Testing is a Bad Idea (my talk at Warwick, England, 2:30pm Thurs 15 Sept) appeared first on All About Statistics.
Why Science Should Stay Clear of Metaphysics - Issue 40: Learning
Philosophers of science are not known for agreeing with each other--contrariness is part of the job description. But for thousands of years, from Aristotle to Thomas Kuhn, those who study what science is have roughly categorized themselves into two basic camps: "realists" and "anti-realists." In philosophical terms, "anti-realists" or "empiricists" understand science as investigating the properties of observable objects via experiments. Empirical theories are constrained by the experimental results. "Realists," on the other hand, speculate more freely about the possible shape of the unobservable world, often designing mathematical explanations that cannot (yet) be tested. Isaac Newton was a realist, as are string theorists. Most scientists do not lose sleep worrying about philosophical divides. But maybe they should; Albert Einstein certainly did, as did Niels Bohr, and Erwin Schrödinger.
Conversational Marketing: A New Paradigm for Brands – Chatbots Magazine
There's one recurring obsession that keeps haunting almost every marketing and ad agency executive. It's in nearly every Powerpoint deck and on every marketer's mind… Millennials: the oh so documented and ubiquitous M Word!. For many years, most marketers (and their agencies) relied on the same old recipe. Brand awareness and reach were the only metrics that mattered. Problem is, the recipe started to turn sour.
Scientists reveal how LSD changes the way the brain processes language
Lauded by hippies, music heads and fans of all things psychedelic, acid has been used to blur the boundaries of reality and perception for decades. Experts have drawn parallels between the dissociative effects caused by the drug on the brain and psychiatric illnesses, with hope it could potentially be explored as a treatment. Now researchers have shown how the drug affects language and speech, reveal it may even enable users to be more creative. In the trials, participants took between 40 to 80 micrograms of the drug intravenously, which would be in the same range as the average tab of acid (illustrated). Researchers in Germany and the UK carried out trials in which participants were asked to name a number of pictures, either under the influence of acid or taking a placebo.
2at1RLo
From the era of the desktop app to the era of the web page to the era of the mobile app to the latest paradigm shift which seems to be happening now: the conversation. These providers will most likely sit at the center of an ecosystem which will handle NLP (Natural Language Processing), semantic analysis, and other core tasks such as location and calendar integration. Currently, there are "bits and pieces" for particulars like dialogs (IBM Dialog) and NLP (IBM AlchemyAPI) all the way to large sdk's for voice and digital assistants (Alexa, Siri, and Google). While the examples above are simplistic they do provide some structure and a view into the basic text lines of voice and chat applications.
A New Take on Data Discovery, Data Management, and its Relationships - DATAVERSITY
Having herself held senior roles in IT at Wall Street companies including Deutsche Bank and Morgan Stanley Smith Barney, Oksana Sokolovsky is quite familiar with the challenge of Data Management and data discovery. As co-founder and CEO of ROKITT, her goal was "to build a product that solves that challenge," she says. The challenge exists across large enterprises in multiple industries, but is often especially acute in those dealing with regulatory pressures and compliance requirements – healthcare, for instance, and of course, the financial sector. Basel Committee on Banking Supervision (BCBS) 239 compliance for effective risk data aggregation and reporting, for example, is a big driver of improved Data Management for global systemically important banks. In fact, a McKinsey & Company and Institute of International Finance survey showed that more than half of the world's biggest banks faced significant challenges meeting the January 1, 2016 deadline for compliance, with the Global Association of Risk Professionals commenting that "many institutions continue to struggle to fully implement the requirements across the business under the most demanding interpretation of those requirements."
EigenTransitions with Hypothesis Testing: The Anatomy of Urban Mobility
Zhang, Ke (University of Pittsburgh) | Lin, Yu-Ru (University of Pittsburgh) | Pelechrinis, Konstantinos (University of Pittsburgh)
Identifying the patterns in urban mobility is important for a variety of tasks such as transportation planning, urban resource allocation, emergency planning etc. This is evident from the large body of research on the topic, which has exploded with the vast amount of geo-tagged user-generated content from online social media. However, most of the existing work focuses on a specific setting, taking a statistical approach to describe and model the observed patterns. On the contrary in this work we introduce EigenTransitions, a spectrum-based, generic framework for analyzing spatio-temporal mobility datasets. EigenTransitions capture the anatomy of the aggregate and/or individuals’ mobility as a compact set of latent mobility patterns. Using a large corpus of geo-tagged content collected from Twitter, we utilize EigenTransitions to analyze the structure of urban mobility. In particular, we identify the EigenTransitions of a flow network between urban areas and derive hypothesis testing framework to evaluate urban mobility from both temporal and demographic perspectives. We further show how EigenTransitions not only identify latent mobility patterns, but also have the potential to support applications such as mobility prediction and inter-city comparisons. In particular, by identifying neighbors with similar latent mobility patterns and incorporating their historical transition behaviors, we proposed an EigenTransitions-based k-nearest neighbor algorithm, which can significantly improve the performance of individual mobility prediction. The proposed method is especially effective in “cold-start” scenarios where traditional methods are known to perform poorly.