Goto

Collaborating Authors

 Genre


The Risk of Machine Learning

arXiv.org Machine Learning

Many applied settings in empirical economics involve simultaneous estimation of a large number of parameters. In particular, applied economists are often interested in estimating the effects of many-valued treatments (like teacher effects or location effects), treatment effects for many groups, and prediction models with many regressors. In these settings, machine learning methods that combine regularized estimation and data-driven choices of regularization parameters are useful to avoid over-fitting. In this article, we analyze the performance of a class of machine learning estimators that includes ridge, lasso and pretest in contexts that require simultaneous estimation of many parameters. Our analysis aims to provide guidance to applied researchers on (i) the choice between regularized estimators in practice and (ii) data-driven selection of regularization parameters. To address (i), we characterize the risk (mean squared error) of regularized estimators and derive their relative performance as a function of simple features of the data generating process. To address (ii), we show that data-driven choices of regularization parameters, based on Stein's unbiased risk estimate or on cross-validation, yield estimators with risk uniformly close to the risk attained under the optimal (unfeasible) choice of regularization parameters. We use data from recent examples in the empirical economics literature to illustrate the practical applicability of our results.


Bi-class classification of humpback whale sound units against complex background noise with Deep Convolution Neural Network

arXiv.org Machine Learning

Automatically detecting sound units of humpback whales in complex time-varying background noises is a current challenge for scientists. In this paper, we explore the applicability of Convolution Neural Network (CNN) method for this task. In the evaluation stage, we present 6 bi-class classification experimentations of whale sound detection against different background noise types (e.g., rain, wind). In comparison to classical FFT-based representation like spectrograms, we showed that the use of image-based pretrained CNN features brought higher performance to classify whale sounds and background noise.


Intraoperative margin assessment of human breast tissue in optical coherence tomography images using deep neural networks

arXiv.org Machine Learning

Objective: In this work, we perform margin assessment of human breast tissue from optical coherence tomography (OCT) images using deep neural networks (DNNs). This work simulates an intraoperative setting for breast cancer lumpectomy. Methods: To train the DNNs, we use both the state-of-the-art methods (Weight Decay and DropOut) and a newly introduced regularization method based on function norms. Commonly used methods can fail when only a small database is available. The use of a function norm introduces a direct control over the complexity of the function with the aim of diminishing the risk of overfitting. Results: As neither the code nor the data of previous results are publicly available, the obtained results are compared with reported results in the literature for a conservative comparison. Moreover, our method is applied to locally collected data on several data configurations. The reported results are the average over the different trials. Conclusion: The experimental results show that the use of DNNs yields significantly better results than other techniques when evaluated in terms of sensitivity, specificity, F1 score, G-mean and Matthews correlation coefficient. Function norm regularization yielded higher and more robust results than competing methods. Significance: We have demonstrated a system that shows high promise for (partially) automated margin assessment of human breast tissue, Equal error rate (EER) is reduced from approximately 12\% (the lowest reported in the literature) to 5\%\,--\,a 58\% reduction. The method is computationally feasible for intraoperative application (less than 2 seconds per image).


Learning Deep Features via Congenerous Cosine Loss for Person Recognition

arXiv.org Machine Learning

Person recognition aims at recognizing the same identity across time and space with complicated scenes and similar appearance. In this paper, we propose a novel method to address this task by training a network to obtain robust and representative features. The intuition is that we directly compare and optimize the cosine distance between two features - enlarging inter-class distinction as well as alleviating inner-class variance. We propose a congenerous cosine loss by minimizing the cosine distance between samples and their cluster centroid in a cooperative way. Such a design reduces the complexity and could be implemented via softmax with normalized inputs. Our method also differs from previous work in person recognition that we do not conduct a second training on the test subset. The identity of a person is determined by measuring the similarity from several body regions in the reference set. Experimental results show that the proposed approach achieves better classification accuracy against previous state-of-the-arts.


Errors-in-variables models with dependent measurements

arXiv.org Machine Learning

Suppose that we observe $y \in \mathbb{R}^n$ and $X \in \mathbb{R}^{n \times m}$ in the following errors-in-variables model: \begin{eqnarray*} y & = & X_0 \beta^* +\epsilon \\ X & = & X_0 + W, \end{eqnarray*} where $X_0$ is an $n \times m$ design matrix with independent subgaussian row vectors, $\epsilon \in \mathbb{R}^n$ is a noise vector and $W$ is a mean zero $n \times m$ random noise matrix with independent subgaussian column vectors, independent of $X_0$ and $\epsilon$. This model is significantly different from those analyzed in the literature in the sense that we allow the measurement error for each covariate to be a dependent vector across its $n$ observations. Such error structures appear in the science literature when modeling the trial-to-trial fluctuations in response strength shared across a set of neurons. Under sparsity and restrictive eigenvalue type of conditions, we show that one is able to recover a sparse vector $\beta^* \in \mathbb{R}^m$ from the model given a single observation matrix $X$ and the response vector $y$. We establish consistency in estimating $\beta^*$ and obtain the rates of convergence in the $\ell_q$ norm, where $q = 1, 2$. We show error bounds which approach that of the regular Lasso and the Dantzig selector in case the errors in $W$ are tending to 0. We analyze the convergence rates of the gradient descent methods for solving the nonconvex programs and show that the composite gradient descent algorithm is guaranteed to converge at a geometric rate to a neighborhood of the global minimizers: the size of the neighborhood is bounded by the statistical error in the $\ell_2$ norm. Our analysis reveals interesting connections between computational and statistical efficiency and the concentration of measure phenomenon in random matrix theory. We provide simulation evidence illuminating the theoretical predictions.


Membership Inference Attacks against Machine Learning Models

arXiv.org Machine Learning

We quantitatively investigate how machine learning models leak information about the individual data records on which they were trained. We focus on the basic membership inference attack: given a data record and black-box access to a model, determine if the record was in the model's training dataset. To perform membership inference against a target model, we make adversarial use of machine learning and train our own inference model to recognize differences in the target model's predictions on the inputs that it trained on versus the inputs that it did not train on. We empirically evaluate our inference techniques on classification models trained by commercial "machine learning as a service" providers such as Google and Amazon. Using realistic datasets and classification tasks, including a hospital discharge dataset whose membership is sensitive from the privacy perspective, we show that these models can be vulnerable to membership inference attacks. We then investigate the factors that influence this leakage and evaluate mitigation strategies.


Hankel Matrix Nuclear Norm Regularized Tensor Completion for $N$-dimensional Exponential Signals

arXiv.org Machine Learning

Signals are generally modeled as a superposition of exponential functions in spectroscopy of chemistry, biology and medical imaging. For fast data acquisition or other inevitable reasons, however, only a small amount of samples may be acquired and thus how to recover the full signal becomes an active research topic. But existing approaches can not efficiently recover $N$-dimensional exponential signals with $N\geq 3$. In this paper, we study the problem of recovering N-dimensional (particularly $N\geq 3$) exponential signals from partial observations, and formulate this problem as a low-rank tensor completion problem with exponential factor vectors. The full signal is reconstructed by simultaneously exploiting the CANDECOMP/PARAFAC structure and the exponential structure of the associated factor vectors. The latter is promoted by minimizing an objective function involving the nuclear norm of Hankel matrices. Experimental results on simulated and real magnetic resonance spectroscopy data show that the proposed approach can successfully recover full signals from very limited samples and is robust to the estimated tensor rank.


Ingenious: Lisa Feldman Barrett - Issue 46: Balance

Nautilus

Do you think you can read emotions like joy or anger in another person's face and actions? Read them because joy and anger are universal emotions and we all know what they look and feel like? Well, if so, says neuroscientist Lisa Feldman Barrett, you are winging it, guessing at best. Emotions like happiness and despair are not baked into our brains, waiting to be triggered by experiences in the world. Sure, we have a range of feelings, stimulated by our senses. But those feelings cannot be categorized as emotions innate in everyone. What we call emotions, Barrett says, are concepts constructed by our individual neural systems, molded by our cultures and past experiences. In her new and first book, How Emotions Are Made: The Secret Life of the Brain, based on years of research at her neuroscience lab at Northeastern University, Barrett spells out the "theory of constructed emotion."


Darwin Was a Slacker and You Should Be Too - Issue 46: Balance

Nautilus

When you examine the lives of history's most creative figures, you are immediately confronted with a paradox: They organize their lives around their work, but not their days. Figures as different as Charles Dickens, Henri Poincaré, and Ingmar Bergman, working in disparate fields in different times, all shared a passion for their work, a terrific ambition to succeed, and an almost superhuman capacity to focus. Yet when you look closely at their daily lives, they only spent a few hours a day doing what we would recognize as their most important work. The rest of the time, they were hiking mountains, taking naps, going on walks with friends, or just sitting and thinking. Their creativity and productivity, in other words, were not the result of endless hours of toil. Their towering creative achievements result from modest "working" hours. How did they manage to be so accomplished? Can a generation raised to believe that 80-hour workweeks are necessary for success learn something from the lives of the people who laid the foundations of chaos theory and topology or wrote Great Expectations? If some of history's greatest figures didn't put in immensely long hours, maybe the key to unlocking the secret of their creativity lies in understanding not just how they labored but how they rested, and how the two relate. Let's start by looking at the lives of two figures. They were both very accomplished in their fields.


Noise Is a Drug and New York Is Full of Addicts - Issue 46: Balance

Nautilus

As soon as the door slams, I slide to the floor in a cross-legged position and hold my breath. The room in which I have just barricaded myself looks a bit like Matilda's chokey; a single light bulb casts a sickly yellow glow about the room, its walls lined with triangle-shaped chunks of fiberglass straining against wire mesh. In 15 minutes I will leave this room for the cacophonous world of Manhattan. I should, theoretically, be appreciating this small respite for what it is. Even so, with every second, I feel as if I'm going deeper underwater. I am sitting in an anechoic chamber, the only one in New York City. Nestled in the hip, angled building of The Cooper Union for the Advancement of Science and Art, the anechoic chamber is where acoustics students, headed by the aptly-named Melody Baglione, conduct research--it's the equivalent of a zero-gravity chamber, only in this case, the variable is sound. The room is designed to be as noise-free as possible; its chunky walls completely absorb reflections of sound waves, and insulate the space within from all exterior sources of noise.