Genre
Why Intel Is Tweaking Xeon Phi For Deep Learning
If there is anything that chip giant Intel has learned over the past two decades as it has gradually climbed to dominance in processing in the datacenter, it is ironically that one size most definitely does not fit all. As the tight co-design of hardware and software continues in all parts of the IT industry, we can expect fine-grained customization for very precise – and lucrative – workloads, like data analytics and machine learning, just to name two of the hottest areas today. Software will run most efficiently on hardware that is tuned for it, although we are used to thinking of that process in a mirror image, where programmers tweak their code to take advantage of the forward-looking features a chip maker conceives of four or five years before they are etched into its transistors and delivered as a product. The competition is fierce these days, and Intel has to move fast if it is to keep its compute hegemony in the datacenter. That is why at the Intel Developer Forum in San Francisco the company put a new path on the Knights family of many-core processors that will see the company deliver a version of this chip specifically tuned for machine learning workloads.
33 Corporations Working On Autonomous Vehicles
Want to receive a weekly deep dive into all things auto, transportation, & logistics tech? Click here to subscribe to our auto tech newsletter. Private companies working in auto tech are on pace to attract record levels of deals and funding in 2016, with autonomous driving startups leading the charge. As expectations around self-driving vehicles have risen, major corporations have ramped up their own initiatives, racing to deploy technology onto public roads. Using CB Insights' investment, acquisition, and partnership data, we identified 33 corporate groups involved in the development of advanced driver assistance systems and self-driving vehicles. They are a diverse group of players, ranging from automotive industry stalwarts to leading technology brands. The list is organized alphabetically (companies working on industrial autonomous vehicles were not included in this analysis).
MD Anderson Benches IBM Watson In Setback For Artificial Intelligence In Medicine
It was one of those amazing "we're living in the future" moments. In an October 2013 press release, IBM declared that MD Anderson, the cancer center that is part of the University of Texas, "is using the IBM Watson cognitive computing system for its mission to eradicate cancer." Well, now that future is past. The partnership between IBM and one of the world's top cancer research institutions is falling apart. The project is on hold, MD Anderson confirms, and has been since late last year.
Artificial intelligence 'to revolutionise higher education'
The use of artificial intelligence and the "next-generation" of virtual learning environments (VLEs) are two areas of technology that have been forecast to have a major impact on higher education in the future, according to the expert panel of a major new report. The NMC Horizon Report: 2017 Higher Education Edition is produced by the New Media Consortium – a community of hundreds of universities, colleges, museums and research organisations driving innovation across their campuses – and is the flagship publication of the NMC Horizon Project, which analyses emerging technology uptake in education. Artificial intelligence, the report notes, has the "potential to enhance online learning, adaptive learning software, and research processes in ways that more intuitively respond to and engage with students". Samantha Adams Becker, senior director of publications and communications at NMC and the report's editor, said that the higher education world was already seeing the initial benefits of AI, which was "very much driving" the adaptive learning field. "If you think about online courses where there may be hundreds of students, it's currently very difficult for a professor or instructor to maybe get a good grasp on how students not only are performing, but are feeling about the material…as they're lecturing or a video's playing," she said. "Virtual avatars and chatbots…have the ability to assess that on an individual level, and if the student seems stuck then maybe you can replay part of the video.
A Sparse Linear Model and Significance Test for Individual Consumption Prediction
Li, Pan, Zhang, Baosen, Weng, Yang, Rajagopal, Ram
Accurate prediction of user consumption is a key part not only in understanding consumer flexibility and behavior patterns, but in the design of robust and efficient energy saving programs as well. Existing prediction methods usually have high relative errors that can be larger than 30% and have difficulties accounting for heterogeneity between individual users. In this paper, we propose a method to improve prediction accuracy of individual users by adaptively exploring sparsity in historical data and leveraging predictive relationship between different users. Sparsity is captured by popular least absolute shrinkage and selection estimator, while user selection is formulated as an optimal hypothesis testing problem and solved via a covariance test. Using real world data from PG&E, we provide extensive simulation validation of the proposed method against well-known techniques such as support vector machine, principle component analysis combined with linear regression, and random forest. The results demonstrate that our proposed methods are operationally efficient because of linear nature, and achieve optimal prediction performance. Pan Li and Baosen Zhang are with the Department of Electrical Engineering, University of Washington, Seattle, WA, 98195, (email: {pli69, zhangbao}@uw.edu). Yang Weng and Ram Rajagopal are with the Civil and Environmental Department, Stanford University, Stanford, CA, 94035, (email: {yangweng, ramr}@stanford.edu). 2 Estimated consumption at time t. Estimated variance of the noise. Electric load forecasting is an important problem in the power engineering industry and have received extensive attention from both industry and academia over the last century. Many different forecasting techniques have been developed during this time. The authors in [1] present a comprehensive literature review on different methods related to load forecasting, from regression models to expert systems. Time series methods are further discussed in [2]. A thorough research on load and price forecasting is presented in [3]. A common theme among many of the established methods is that they are used to forecast relative large loads, from substations serving megawatts to transmission networks serving more than gigawatts of power [4]. Recent advances in technology such as smart meters, bidirectional communication capabilities and distributed energy resources have made individual households active participants in the power system. Many applications and programs based on these new technologies require estimating the future load of individual homes.
Social Learning and Diffusion of Pervasive Goods: An Empirical Study of an African App Store
Nia, Meisam Hejazi, Ratchford, Brian T., Bruce, Norris
In this study, the authors develop a structural model that combines a macro diffusion model with a micro choice model to control for the effect of social influence on the mobile app choices of customers over app stores. Social influence refers to the density of adopters within the proximity of other customers. Using a large data set from an African app store and Bayesian estimation methods, the authors quantify the effect of social influence and investigate the impact of ignoring this process in estimating customer choices. The findings show that customer choices in the app store are explained better by offline than online density of adopters and that ignoring social influence in estimations results in biased estimates. Furthermore, the findings show that the mobile app adoption process is similar to adoption of music CDs, among all other classic economy goods. A counterfactual analysis shows that the app store can increase its revenue by 13.6% through a viral marketing policy (e.g., a sharing with friends and family button).
Stochastic Composite Least-Squares Regression with convergence rate O(1/n)
Flammarion, Nicolas, Bach, Francis
We consider the minimization of composite objective functions composed of the expectation of quadratic functions and an arbitrary convex function. We study the stochastic dual averaging algorithm with a constant step-size, showing that it leads to a convergence rate of O(1/n) without strong convexity assumptions. This thus extends earlier results on least-squares regression with the Euclidean geometry to (a) all convex regularizers and constraints, and (b) all geome-tries represented by a Bregman divergence. This is achieved by a new proof technique that relates stochastic and deterministic recursions.
Interpreting Outliers: Localized Logistic Regression for Density Ratio Estimation
Yamada, Makoto, Liu, Song, Kaski, Samuel
We propose an inlier-based outlier detection method capable of both identifying the outliers and explaining why they are outliers, by identifying the outlier-specific features. Specifically, we employ an inlier-based outlier detection criterion, which uses the ratio of inlier and test probability densities as a measure of plausibility of being an outlier. For estimating the density ratio function, we propose a localized logistic regression algorithm. Thanks to the locality of the model, variable selection can be outlier-specific, and will help interpret why points are outliers in a high-dimensional space. Through synthetic experiments, we show that the proposed algorithm can successfully detect the important features for outliers. Moreover, we show that the proposed algorithm tends to outperform existing algorithms in benchmark datasets.
Maximally Correlated Principal Component Analysis
Soheil Feizi and David Tse Stanford University Abstract In the era of big data, reducing data dimensionality is critical in many areas of science. Widely used Principal Component Analysis (PCA) addresses this problem by computing a low dimensional data embedding that maximally explain variance of the data. However, PCA has two major weaknesses. Firstly, it only considers linear correlations among variables (features), and secondly it is not suitable for categorical data. We resolve these issues by proposing Maximally Correlated Principal Component Analysis (MCPCA). MCPCA computes transformations of variables whose covariance matrix has the largest Ky Fan norm. Variable transformations are unknown, can be nonlinear and are computed in an optimization. MCPCA can also be viewed as a multivariate extension of Maximal Correlation. For jointly Gaussian variables we show that the covariance matrix corresponding to the identity (or the negative of the identity) transformations majorizes covariance matrices of non-identity functions. Using this result we characterize global MCPCA optimizers for nonlinear functions of jointly Gaussian variables for every rank constraint. For categorical variables we characterize global MCPCA optimizers for the rank one constraint based on the leading eigenvector of a matrix computed using pairwise joint distributions. For a general rank constraint we propose a block coordinate descend algorithm and show its convergence to stationary points of the MCPCA optimization. We compare MCPCA with PCA and other state-of-the-art dimensionality reduction methods including Isomap, LLE, multilayer autoencoders (neural networks), kernel PCA, probabilistic PCA and diffusion maps on several synthetic and real datasets. We show that MCPCA consistently provides improved performance compared to other methods. 1 Introduction Let X 1 and X 2 be two mean zero and unit variance random variables. Pearson's correlation [1] defined as ρ Pearson(X 1,X 2) E [X 1X 2 ] (1.1) is a basic statistical parameter and plays a central role in many statistical and machine learning methods such as linear regression [2], principal component analysis [3], and support vector machines [4], partially owing to its simplicity and computational efficiency. Pearson's correlation however has two main weaknesses: firstly it only captures linear dependency between variables, and secondly for discrete (categorical) variables the value of Pearson's correlation depends somewhat arbitrarily on the labels. To overcome these weaknesses, Maximal Correlation (MC) has been proposed and 1 arXiv:1702.05471v2 MC tackles the two main drawbacks of the Pearson's correlation: it models a family of nonlinear relationships between the two variables.
Generative Temporal Models with Memory
Gemici, Mevlana, Hung, Chia-Chun, Santoro, Adam, Wayne, Greg, Mohamed, Shakir, Rezende, Danilo J., Amos, David, Lillicrap, Timothy
We consider the general problem of modeling temporal data with long-range dependencies, wherein new observations are fully or partially predictable based on temporally-distant, past observations. A sufficiently powerful temporal model should separate predictable elements of the sequence from unpredictable elements, express uncertainty about those unpredictable elements, and rapidly identify novel elements that may help to predict the future. To create such models, we introduce Generative T emporal Modelsaugmented with external memory systems. They are developed within the variational inference framework, which provides both a practical training methodology and methods to gain insight into the models' operation. We show, on a range of problems with sparse, long-term temporal dependencies, that these models store information from early in a sequence, and reuse this stored information efficiently. This allows them to perform substantially better than existing models based on well-known recurrent neural networks, like LSTMs. Many of the data sets we use in machine learning applications are sequential, whether these be natural language and speech processing data, streams of high-definition video, longitudinal time-series from medical diagnostics, or spatiotemporal data in climate forecasting. Generative Temporal Models (GTMs) are a core requirement for these applications. Generative Temporal Models are also important components of intelligent agents, as they permit counterfactual reasoning, physical predictions, robot localisation, and simulation-based planning among other capacities (Sutton, 1991; Deisenroth and Rasmussen, 2011; Watter et al., 2015; Levine and Abbeel, 2014; Assael et al., 2015). These tasks require models of high-dimensional observation sequences and contain complex, long temporal dependencies--requirements that most available GTMs are unable to fulfil. Developing such GTMs is the aim of this paper. Many GTMs--whether they are linear or nonlinear, deterministic or stochastic--assume that the underlying temporal dynamics is governed by low-order Markov transitions and use fixed-dimensional sufficient statistics. Examples of such models include Hidden Markov Models (Rabiner, 1989), and linear dynamical systems such as Kalman filters and their nonlinear extensions (Kalman, 1960; Ghahramani and Hinton, 1996; Krishnan et al., 2015). The fixed-order Markov assumption used in these models is insufficient for characterising many systems of practical relevance.