Scientific Discovery
A unified framework for bandit multiple testing
Xu, Ziyu, Wang, Ruodu, Ramdas, Aaditya
In bandit multiple hypothesis testing, each arm corresponds to a different null hypothesis that we wish to test, and the goal is to design adaptive algorithms that correctly identify large set of interesting arms (true discoveries), while only mistakenly identifying a few uninteresting ones (false discoveries). One common metric in non-bandit multiple testing is the false discovery rate (FDR). We propose a unified, modular framework for bandit FDR control that emphasizes the decoupling of exploration and summarization of evidence. We utilize the powerful martingale-based concept of ``e-processes'' to ensure FDR control for arbitrary composite nulls, exploration rules and stopping times in generic problem settings. In particular, valid FDR control holds even if the reward distributions of the arms could be dependent, multiple arms may be queried simultaneously, and multiple (cooperating or competing) agents may be querying arms, covering combinatorial semi-bandit type settings as well. Prior work has considered in great detail the setting where each arm's reward distribution is independent and sub-Gaussian, and a single arm is queried at each step. Our framework recovers matching sample complexity guarantees in this special case, and performs comparably or better in practice. For other settings, sample complexities will depend on the finer details of the problem (composite nulls being tested, exploration algorithm, data dependence structure, stopping rule) and we do not explore these; our contribution is to show that the FDR guarantee is clean and entirely agnostic to these details.
Generalized Multivariate Signs for Nonparametric Hypothesis Testing in High Dimensions
Majumdar, Subhabrata, Chatterjee, Snigdhansu
High-dimensional data, where the dimension of the feature space is much larger than sample size, arise in a number of statistical applications. In this context, we construct the generalized multivariate sign transformation, defined as a vector divided by its norm. For different choices of the norm function, the resulting transformed vector adapts to certain geometrical features of the data distribution. Building up on this idea, we obtain one-sample and two-sample testing procedures for mean vectors of high-dimensional data using these generalized sign vectors. These tests are based on U-statistics using kernel inner products, do not require prohibitive assumptions, and are amenable to a fast randomization-based implementation. Through experiments in a number of data settings, we show that tests using generalized signs display higher power than existing tests, while maintaining nominal type-I error rates. Finally, we provide example applications on the MNIST and Minnesota Twin Studies genomic data.
A fuzzy take on the logical issues of statistical hypothesis testing
Booth, Matthew, Paillusson, Fabien
Statistical Hypothesis Testing (SHT) is a class of inference methods whereby one makes use of empirical data to test a hypothesis and often emit a judgment about whether to reject it or not. In this paper we focus on the logical aspect of this strategy, which is largely independent of the adopted school of thought, at least within the various frequentist approaches. We identify SHT as taking the form of an unsound argument from Modus Tollens in classical logic, and, in order to rescue SHT from this difficulty, we propose that it can instead be grounded in t-norm based fuzzy logics. We reformulate the frequentists' SHT logic by making use of a fuzzy extension of modus Tollens to develop a model of truth valuation for its premises. Importantly, we show that it is possible to preserve the soundness of Modus Tollens by exploring the various conventions involved with constructing fuzzy negations and fuzzy implications (namely, the S and R conventions). We find that under the S convention, it is possible to conduct the Modus Tollens inference argument using Zadeh's compositional extension and any possible t-norm. Under the R convention we find that this is not necessarily the case, but that by mixing R-implication with S-negation we can salvage the product t-norm, for example. In conclusion, we have shown that fuzzy logic is a legitimate framework to discuss and address the difficulties plaguing frequentist interpretations of SHT.
Statistics Fundamentals (7/9) Hypothesis Testing
Statistics Fundamentals (7/9) Hypothesis Testing Statistical Hypothesis Testing: Theory and Python Welcome to Statistics Fundamentals 7, Hypothesis Testing. This course is for beginners who are interested in statistical analysis. Description Welcome to Statistics Fundamentals 7, Hypothesis Testing. This course is for beginners who are interested in statistical analysis. And anyone who is not a beginner but wants to go over from the basics is also welcome!
Living in the wilderness: hypothesis testing in a world that disagrees with statistical theory
Sometimes it seems paradoxical to call the famous bell curve "normal". Among all the assumptions made by traditional statistical theory, the normality assumption is notorious for the frequency it doesn't hold. My aim in this article is to show a way to test hypotheses when the normality assumption of traditional hypothesis tests is violated. In this scenario, we can't rely on theoretical results, so we need to depart from theory's ivory tower and double the bet on our data. To get there, first I briefly review what hypothesis testing is, focusing on an intuitive grasp of the reasoning behind it (no equations allowed!). Then I proceed to a case study motivated by a business problem where the normality assumption doesn't hold. This makes matters concrete and will direct our discussion. After the problem is explained, I will show that bootstrapping is a good way to fill the gaps left by theory without changing anything in the reasoning at the heart of hypothesis testing. In particular, I will show that bootstrapping leads to the right conclusion about the test. I conclude this article with a critical evaluation of bootstrapping and similar methods, pointing out their pros and cons. Many data scientists have trouble understanding hypothesis testing.
The problem with 'follow your dream
I walked into my adviser's office, overflowing with frustration and confusion about the advice I had received at a recent career development workshop. It reiterated what I had heard so many times before: I should follow my dream, and if I didn't yet know what that was, I should live with career uncertainty until I figured it out. But as an international student working in the United States, taking time to explore wasn't an option for me. After listening to me rant, my adviser calmly looked across his desk. He told me that instead of focusing on finding a dream job, I should think about what I am good at and what makes me happy at least 80% of the time. This advice surprised me at first, but it ended up being exactly what I needed to hear. > โI should think about what I am good at and what makes me happy at least 80% of the time.โ I had spent the previous 22 years following my childhood dreamโbecoming a professor of marine biology. However, in grad school I saw how applying for grants is a constant source of worry for many professors. I realized I did not want to be responsible for the salaries of my hypothetical lab members. About 4 years into the program, I decided I did not want to pursue a career in research after all. I began to attend career panels, which all followed a worryingly similar template. I would walk into the room with other excited graduate students and collect my free cookies and coffee, confident that the panelists would have the magical answers I needed. Instead, they would talkโagainโabout following their dreams. The message: I just needed to find a new dream. It would mean taking time off from work to self-reflect and discover a new path. But I couldn't stay in the country without a visa. For most academic researchers, obtaining a university-sponsored visa is relatively straightforward. But outside of academia, it is infinitely more complex, requiring a company that has a job opening and is willing to foot the bill for a work visa. As well-meaning as the panelists were, they fell silent when I brought up this dilemma. I felt totally lost. Finally, I went to my adviser for help. We hadn't talked much about my career plans over the years, but I felt I needed a new perspective from someone who knew me well. When he offered his advice, I was taken aback at first. What happened to โif you love what you do, you'll never work a day in your lifeโ? My adviser assured me there is seldom such a job. Every job has its ugly bits. But as long as you're happy most of the time, you can struggle through the parts you don't like. He also said it was important to find a job I was good at, especially because my visa applications required me to make the case that I would benefit the country. I was relieved to finally have helpful, practical advice. But I discovered that finding overlap between what I like and what I'm good at was not easy. I love scuba diving, but the physical demands are a challenge for me. I'm good at teaching, as evidenced by my friends nagging me to teach them chemistry and microbiology during my high school and undergraduate years and getting rave reviews from my students when I was a teaching assistant, but I don't like repeating the same content every year. Through my teaching experience, however, I also learned that I love telling stories about science. Maybe science communication would offer the overlap I was looking for. To test the waters, during my โspare timeโ in grad school I started a blog about the history of scientific discoveries. I found that I loved the freedom to choose what to write about, and I never encountered a challenge I didn't enjoy. As for whether I was any good at it, the signs were promising. My writing got noticed, eventually by people at my institution, and I was given opportunities to write press releases and stories for the university's news bureau. After 3 years of writing, I was offered a position as a science writer. It's nothing like my childhood dream. But I am happyโmore than 80% of the time.
A learning theoretic perspective on local explainability
Going from left to right, we consider increasingly complex functions. These neighborhoods, in other words, need to become more and more disjoint as the function becomes more complex. Indeed, we quantify "disjointedness" of the neighborhoods via a term denoted by and relate it to the complexity of the function class, and subsequently, its generalization properties. There has been a growing interest in interpretable machine learning (IML), towards helping users better understand how their ML models behave. IML has become a particularly relevant concern especially as practitioners aim to apply ML in important domains such as healthcare [Caruana et al., '15], financial services [Chen et al., '18], and scientific discovery [Karpatne et al., '17]. While much of the work in IML has been qualitative and empirical, in our recent ICLR21 paper, we study how concepts in interpretability can be formally related to learning theory.
Selective Probabilistic Classifier Based on Hypothesis Testing
Germi, Saeed Bakhshi, Rahtu, Esa, Huttunen, Heikki
In this paper, we propose a simple yet effective method to deal with the violation of the Closed-World Assumption for a classifier. Previous works tend to apply a threshold either on the classification scores or the loss function to reject the inputs that violate the assumption. However, these methods cannot achieve the low False Positive Ratio (FPR) required in safety applications. The proposed method is a rejection option based on hypothesis testing with probabilistic networks. With probabilistic networks, it is possible to estimate the distribution of outcomes instead of a single output. By utilizing Z-test over the mean and standard deviation for each class, the proposed method can estimate the statistical significance of the network certainty and reject uncertain outputs. The proposed method was experimented on with different configurations of the COCO and CIFAR datasets. The performance of the proposed method is compared with the Softmax Response, which is a known top-performing method. It is shown that the proposed method can achieve a broader range of operation and cover a lower FPR than the alternative.
The Evolution of Data Catalogs: The Data Discovery Platform
As someone who has spent 13 years in the weeds of data, I witnessed the rise of the "data-driven" trend first hand. Before starting and selling my first data startup, I spent time as a statistical analyst building sales forecasting models in R, a software engineer creating data transformation jobs, and a product manager running A/B tests and analyzing user behaviors. What all these roles had in common was that they gave me an understanding that the context of data -- what it represents, how it was generated, when it was updated last, and the ways it could be joined with other datasets -- is essential to maximizing the data's potential and driving successful outcomes. However, accessing and understanding the context of data is quite difficult. This is because the context of data is often tribal knowledge, meaning it lives only in the brains of the engineers or analysts who have worked with it recently.
Why Computers Will Likely Never Perform Abductive Inferences
Humans, on the other hand, need none of this. On the basis of very limited or incomplete data, we nonetheless come to the right conclusion about many things (yes, we are fallible, but the miracle is that we are right so often). Noam Chomsky's entire claim to fame in linguistics really amounts to exploring this underdetermination problem, which he referred to as "the poverty of the stimulus." Humans pick up language despite very varied experiences with other human language speakers. Babies born in abusive and sensory deprived environments pick up language.