Genre
Signed Support Recovery for Single Index Models in High-Dimensions
Neykov, Matey, Lin, Qian, Liu, Jun S.
In this paper we study the support recovery problem for single index models $Y=f(\boldsymbol{X}^{\intercal} \boldsymbol{\beta},\varepsilon)$, where $f$ is an unknown link function, $\boldsymbol{X}\sim N_p(0,\mathbb{I}_{p})$ and $\boldsymbol{\beta}$ is an $s$-sparse unit vector such that $\boldsymbol{\beta}_{i}\in \{\pm\frac{1}{\sqrt{s}},0\}$. In particular, we look into the performance of two computationally inexpensive algorithms: (a) the diagonal thresholding sliced inverse regression (DT-SIR) introduced by Lin et al. (2015); and (b) a semi-definite programming (SDP) approach inspired by Amini & Wainwright (2008). When $s=O(p^{1-\delta})$ for some $\delta>0$, we demonstrate that both procedures can succeed in recovering the support of $\boldsymbol{\beta}$ as long as the rescaled sample size $\kappa=\frac{n}{s\log(p-s)}$ is larger than a certain critical threshold. On the other hand, when $\kappa$ is smaller than a critical value, any algorithm fails to recover the support with probability at least $\frac{1}{2}$ asymptotically. In other words, we demonstrate that both DT-SIR and the SDP approach are optimal (up to a scalar) for recovering the support of $\boldsymbol{\beta}$ in terms of sample size. We provide extensive simulations, as well as a real dataset application to help verify our theoretical observations.
A Unified Theory of Confidence Regions and Testing for High Dimensional Estimating Equations
Neykov, Matey, Ning, Yang, Liu, Jun S., Liu, Han
We propose a new inferential framework for constructing confidence regions and testing hypotheses in statistical models specified by a system of high dimensional estimating equations. We construct an influence function by projecting the fitted estimating equations to a sparse direction obtained by solving a large-scale linear program. Our main theoretical contribution is to establish a unified Z-estimation theory of confidence regions for high dimensional problems. Different from existing methods, all of which require the specification of the likelihood or pseudo-likelihood, our framework is likelihood-free. As a result, our approach provides valid inference for a broad class of high dimensional constrained estimating equation problems, which are not covered by existing methods. Such examples include, noisy compressed sensing, instrumental variable regression, undirected graphical models, discriminant analysis and vector autoregressive models. We present detailed theoretical results for all these examples. Finally, we conduct thorough numerical simulations, and a real dataset analysis to back up the developed theoretical results.
A Fuzzy Clustering Algorithm for the Mode Seeking Framework
The analysis of large and possibly high-dimensional datasets is becoming ubiquitous in the sciences. The long-term objective is to gain insight into the structure of measurement or simulation data, for a better understanding of the underlying physical phenomena at work. Clustering is one of the simplest ways of gaining such insight, by finding a suitable decomposition of the data into clusters such that data points within a same cluster share common (and, if possible, exclusive) properties. In this work, we are interested in the mode seeking approach to clustering. This approach assumes the data points to be drawn from some unknown probability distribution and defines the clusters as the basins of attraction of the maxima of the density, requiring a preliminary density estimation phase [7, 5, 10, 11, 13, 15]. The theoretical analysis of this clustering framework has drawn increasing attention recently, see 1 [6, 3, 9, 8, 2]. However, this (hard) clustering method provides a fairly limited knowledge on the structure of the data: while the partition into clusters is well understood, the interplay between clusters (respective locations, proximity relations, interactions) remains unknown. Identifying interfaces between clusters is the first step towards a higher-level understanding of the data, and it already plays a prominent role in some applications such as the study of the conformations space of a protein, where a fundamental question beyond the detection of metastable states is to understand when and how the protein can switch from one metastable state to another [12]. Hard clustering can be used in this context, for instance by defining the border between two clusters as the set of data points whose neighborhood (in the ambient space or in some neighborhood graph) intersects the two clusters, however this kind of information is by nature unstable with respect to perturbations of the data.
AI Boosts Cancer Screens to Nearly 100 Percent Accuracy
Diagnosing cancer is about to get more accurate, with the help of artificial intelligence. Pathologists have diagnosed diseases in more or less the same way for the past 100 years, by laboring over a microscope reviewing biopsy samples on little glass slides. Working almost robotically, they sift through millions of normal cells to identify just a few diseased ones. The task is tedious and prone to human error. But now, scientists and engineers have created a technique that uses artificial intelligence (AI) and can differentiate cancer cells from normal cells almost as well as a top-notch pathologist.
The AI that could predict the future of your relationships
Our inability to predict how people will interact with each other has led to many doomed relationships and awkward moments. But now MIT researchers have taught an algorithm to understand body language patterns in order to guess the next move between two individuals. Trained with YouTube videos and TV shows, it looks for outstretched arms, raised hands or prolonged stares to make predictions, which could help us avoid high-five fails and missed kisses. MIT researchers have taught an algorithm to understand body language patterns in order to guess the next move between two individuals. Researchers at MIT fed their predictive AI 600 hours' worth of videos so it could learn the next move of two characters – whether it would be a hug, kiss, handshake or high-five.
What Does Algorithmic Business Really Mean, Anyway?
Tomorrow's companies are going to rely on algorithms more than ever before, the firm said in its latest research note, which could lead to business models that morph and change automatically. It all sounds like something from science fiction, but what does it really mean? An algorithm is simply a set of instructions to follow when completing a process. Every piece of software is algorithmic, meaning that business has been using algorithms since the launch of the LEO I, so we might well ask why the term is being bandied about so breathlessly now. Whereas businesses used software logic in discrete ways to automate certain processes, in the new model algorithms become a more central part of the business, making decisions that couldn't easily be reached without them, and then even taking action on them automatically.
The Fourth Industrial Revolution : what is it and how it will impact Southeast Asia - Asean, Real Estate
The Fourth Industrial Revolution is about the convergence of automation, artificial intelligence and rising connectivity. In the next five to 10 years, the adoption of these emerging technologies has the potential to raise efficiency, productivity and income levels to improve quality of life in Southeast Asia. The Fourth Industrial Revolution is characterised by the fusion and amplification of emerging technology breakthroughs in artificial intelligence, automation and robotics. This will be multiplied by the extreme connectivity between billions of people with mobile devices with unprecedented access to data and knowledge. Southeast Asian countries have the potential to leapfrog ahead of other developing nations by embracing new technologies to transform how people work, live and play, according to a new report by real estate consultancy JLL.
How do songbirds learn their mating melodies? Scientists reveal clues.
Human babies seem to have a natural knack for languages. Fluency in any of the world's 6,500 languages comes within the first few years of life, without much apparent effort. A recent study of how songbirds learn their melodies seeks to shine light on the cognitive processes through which young birds learn and imitate vocal communication – insights which could lead to a better understanding of the development of human speech, as well. "A bird's baby song is really immature. There's no clear structure - it's more like a baby babbling, but then it becomes structured, like the tutor song [as the bird gets older]," Yoko Yazaki-Sugiyama, co-author of the study published in Nature Communications Tuesday, tells The Christian Science Monitor in a phone interview. The male zebra finch learns a complex song from his father, or tutor, in order to attract a female finch.
Twitter videos to become extra long as Vine also drops six-second limit and lets people make huge videos
Nasa has announced that it has found evidence of flowing water on Mars. Scientists have long speculated that Recurring Slope Lineae -- or dark patches -- on Mars were made up of briny water but the new findings prove that those patches are caused by liquid water, which it has established by finding hydrated salts. Several hundred camped outside the London store in Covent Garden. The 6s will have new features like a vastly improved camera and a pressure-sensitive "3D Touch" display
The first steps with Machine learning -- learning-ai
The learning that is being done is always based on some sort of observations or data, such as examples, direct experience, or instruction. For instance, you might wish to predict how much a user Bob will like a movie that he hasn't seen, based on her ratings of movies that he has seen. This means making informed guesses about some unobserved property of some object, based on observed properties of that object. Supervised learning is a type of machine learning algorithm that uses a known dataset (called the training dataset) to make predictions. The training dataset includes input data and response values.