Scientific Discovery
Google's AI is a "new paradigm" that unites humans and machines
Google is fully aware of artificial intelligence's (AI) potential -- DeepMind's AlphaGo AI is one of today's most well-known examples of its capabilities -- and in an earnings call this week, the company made it clear they believe the future of technology lies with AI. During the call, Sundar Pichai, CEO of Alphabet (Google's parent company), praised the company's decision to invest in AI early, highlighting the concept's trajectory from "a research project to something that can solve new problems for a billion people a day," according to an Inverse report. Pichai went on to note how Google's AI research is already producing products that utilize machine learning, such as the Google Clips camera that debuted earlier this month. "Even though we are in the early days of AI, we are already rethinking how to build products around machine learning," said Pichai. "It's a new paradigm compared to mobile-first software, and I'm thrilled how Google is leading the way."
From Distance Correlation to Multiscale Generalized Correlation
Shen, Cencheng, Priebe, Carey E., Vogelstein, Joshua T.
Understanding and developing a correlation measure that can detect general dependencies is not only imperative to statistics and machine learning, but also crucial to general scientific discovery in the big data age. We proposed the Multiscale Generalized Correlation (MGC) in Shen et al. 2017 as a novel correlation measure, which worked well empirically and helped a number of real data discoveries. But there is a wide gap with respect to the theoretical side, e.g., the population statistic, the convergence from sample to population, how well does the algorithmic Sample MGC perform, etc. To better understand its underlying mechanism, in this paper we formalize the population version of local distance correlations, MGC, and the optimal local scale between the underlying random variables, by utilizing the characteristic functions and incorporating the nearest-neighbor machinery. The population version enables a seamless connection with, and significant improvement to, the algorithmic Sample MGC, both theoretically and in practice, which further allows a number of desirable asymptotic and finite-sample properties to be proved and explored for MGC. The advantages of MGC are further illustrated via a comprehensive set of simulations with linear, nonlinear, univariate, multivariate, and noisy dependencies, where it loses almost no power against monotone dependencies while achieving superior performance against general dependencies.
Uncommon Hypothesis Tests to Debunk Common Misconceptions
I gave a talk about p-values and hypothesis testing at BIDS. Please check out my slides! P-values get a large share of the blame for the replication crisis in science. People take for granted that the tests they use work without justifying the leap from data to model. Often, reported p-values are erroneous because the underlying model doesn't accurately describe the way the data arose.
Open data from the Large Hadron Collider sparks new discovery
Back in 2014, CERN released the data from its Large Hadron Collider (LHC) experiments onto an online portal called the Open Data portal. It was an unprecedented move, making data from the LHC's experiments available to those who don't have access to a particle accelerator. It's not completely up-to-date; there's a three-year embargo on results, so, generally speaking, the most recent data being uploaded is from the year 2014. This was the first time results of any particle collider experiment have been released to the public, and now it's produced results. Last week, a team from MIT released an article in Physical Review Letters that used data from the Compact Muon Solenoid (CMS), one of the LHC's main detectors, to explain a feature within high-energy particle collisions.
The Fourth Paradigm: Data-Intensive Scientific Discovery - Microsoft Research
Increasingly, scientific breakthroughs will be powered by advanced computing capabilities that help researchers manipulate and explore massive datasets. The speed at which any given scientific discipline advances will depend on how well its researchers collaborate with one another, and with technologists, in areas of eScience such as databases, workflow management, visualization, and cloud computing technologies. In The Fourth Paradigm: Data-Intensive Scientific Discovery, the collection of essays expands on the vision of pioneering computer scientist Jim Gray for a new, fourth paradigm of discovery based on data-intensive science and offers insights into how it can be fully realized. "The individual essays--and The Fourth Paradigm as a whole--give readers a glimpse of the horizon for 21st-century research and, at their best, a peek at what lies beyond. "The impact of Jim Gray's thinking is continuing to get people to think in a new way about how data and software are redefining what it means to do science." "I often tell people working in eScience that they aren't in this field because they are visionaries or super-intelligent--it's because they care about science and they are alive now.
What Happens When Two Neutron Stars Collide? Scientific Revolution
Late last week, as some staff astronomers embarked on trips to see Monday's solar eclipse, two of NASA's space-based observatories--Hubble and Chandra X-ray--and at least two land-based telescopes scrambled to capture a far more explosive event. The astronomers who stayed behind trained their telescopes on a patch of sky where they hoped to find an astrophysical Rosetta stone: a cataclysmic event capable of producing electromagnetic signals on top of gravitational waves separately detected by the Advanced Laser Interferometer Gravitational-Wave Observatory (Advanced LIGO). Original story reprinted with permission from Quanta Magazine, an editorially independent publication of the Simons Foundation whose mission is to enhance public understanding of science by covering research developments and trends in mathematics and the physical and life sciences. The LIGO collaboration made headlines in February 2016 when it announced it had detected gravitational waves from two colliding black holes. Four months later, while still in its first observing run, the team confirmed the detection of a second black hole merger.
Bridging the gap between big banks and challengers
American physicist and philosopher Thomas Kuhn came up with idea of a "paradigm shift" in the 1960s to describe a scientific revolution โ a momentous discovery that fundamentally rewrites the laws of science, such as Galileo proving the Earth revolves around the Sun or Newton discovering gravity. It is not too much to say that finance is undergoing a paradigm shift today, driven by smartphones, financial technology startups, and trends such as blockchain and artificial intelligence. "Most of the change in the industry was quite incremental and what I regard as linear โ the introduction of ATMs, the introduction of credit cards, those kinds of things," says Antony Jenkins, former chief executive of Barclays. "When you look at what's happening now with mobile banking, it's a true transformation. The power in people's pockets enables them to change things in really quite a radical way."
Data Science Simplified Part 3: Hypothesis Testing
Application of hypothesis testing is predominant in Data Science. It is imperative to simplify and deconstruct it. Like a crime-fiction story, hypothesis testing, based on data, leads us from a novel suggestion to an effective proposition. Hypothesis originates from the Greek work hupo (under) and thesis(placing). It means an idea made from limited evidence. It is a starting point for further investigation.
Hypotheses testing on infinite random graphs
Drawing on some recent results that provide the formalism necessary to definite stationarity for infinite random graphs, this paper initiates the study of statistical and learning questions pertaining to these objects. Specifically, a criterion for the existence of a consistent test for complex hypotheses is presented, generalizing the corresponding results on time series. As an application, it is shown how one can test that a tree has the Markov property, or, more generally, to estimate its memory.
Two-sample Hypothesis Testing for Inhomogeneous Random Graphs
Ghoshdastidar, Debarghya, Gutzeit, Maurilio, Carpentier, Alexandra, von Luxburg, Ulrike
The study of networks leads to a wide range of high dimensional inference problems. In most practical scenarios, one needs to draw inference from a small population of large networks. The present paper studies hypothesis testing of graphs in this high-dimensional regime. We consider the problem of testing between two populations of inhomogeneous random graphs defined on the same set of vertices. We propose tests based on estimates of the Frobenius and operator norms of the difference between the population adjacency matrices. We show that the tests are uniformly consistent in both the "large graph, small sample" and "small graph, large sample" regimes. We further derive lower bounds on the minimax separation rate for the associated testing problems, and show that the constructed tests are near optimal.