Genre
What Did You Miss at the Deep Learning Summit Last Week?
Media attending the event included BBC News, The Guardian, The Wall Street Journal, Bloomberg, VentureBeat, Digital Trends, Financial Times, Ars Technica and more. News coverage focused on a range of topics, exploring advancements in robotics, chatbot personalities, machine vision for understanding differences in language and culture, as well as startup acquisitions and funding. We've shared just a few of the great articles from the summit below. Why Data is the New Coal - The Guardian Deep learning needs to become more efficient if it is going to move from using data to categorise images of cats to diagnosing rare illnesses. Alex Hern reports on revelations in this area from speaker Neil Lawrence, the newly appointed Senior Principal Scientist at Amazon.
AWS Announces Availability of P2 Instances for Amazon EC2
With up to 16 NVIDIA Tesla K80 GPUs, P2 instances are the most powerful GPU instances available in the cloud. "The massive parallel floating point performance of Amazon EC2 P2 instances, combined with up to 64 vCPUs and 732 GB host memory, will enable customers to realize results faster and process larger datasets than was previously possible." P2 instances allow customers to build and deploy compute-intensive applications using the CUDA parallel computing platform or the OpenCL framework without up-front capital investments. To offer the best performance for these high performance computing applications, the largest P2 instance offers 16 GPUs with a combined 192 Gigabytes (GB) of video memory, 40,000 parallel processing cores, 70 teraflops of single precision floating point performance, over 23 teraflops of double precision floating point performance, and GPUDirect technology for higher bandwidth and lower latency peer-to-peer communication between GPUs. P2 instances also feature up to 732 GB of host memory, up to 64 vCPUs using custom Intel Xeon E5-2686 v4 (Broadwell) processors, dedicated network capacity for I/O operation, and enhanced networking through the Amazon EC2 Elastic Network Adaptor.
Multilevel and Mixed Models Fall 2016
Multilevel models are a class of regression models for data that have a hierarchical (or nested) structure. Common examples of such data structures are students nested within schools or classrooms, patients nested within hospitals, or survey respondents nested within countries. Using regression techniques that ignore this hierarchical structure (such as ordinary least squares) can lead to incorrect results because such methods assume that all observations are independent. Perhaps more important, using inappropriate techniques (like pooling or aggregating) prevents researchers from asking substantively interesting questions about how processes work at different levels. This two-day seminar provides an intensive introduction to multilevel models.
Funneled Bayesian Optimization for Design, Tuning and Control of Autonomous Systems
Bayesian optimization has become a fundamental global optimization algorithm in many problems where sample efficiency is of paramount importance. Recently, there has been proposed a large number of new applications in fields such as robotics, machine learning, experimental design, simulation, etc. In this paper, we focus on several problems that appear in robotics and autonomous systems: algorithm tuning, automatic control and intelligent design. All those problems can be mapped to global optimization problems. However, they become hard optimization problems. Bayesian optimization internally uses a probabilistic surrogate model (e.g.: Gaussian process) to learn from the process and reduce the number of samples required. In order to generalize to unknown functions in a black-box fashion, the common assumption is that the underlying function can be modeled with a stationary process. Nonstationary Gaussian process regression cannot generalize easily and it typically requires prior knowledge of the function. Some works have designed techniques to generalize Bayesian optimization to nonstationary functions in an indirect way, but using techniques originally designed for regression, where the objective is to improve the quality of the surrogate model everywhere. Instead optimization should focus on improving the surrogate model near the optimum. In this paper, we present a novel kernel function specially designed for Bayesian optimization, that allows nonstationary behavior of the surrogate model in an adaptive local region. In our experiments, we found that this new kernel results in an improved local search (exploitation), without penalizing the global search (exploration). We provide results in well-known benchmarks and real applications. The new method outperforms the state of the art in Bayesian optimization both in stationary and nonstationary problems.
Deep Learning Algorithms for Signal Recognition in Long Perimeter Monitoring Distributed Fiber Optic Sensors
In this paper, we show an approach to build deep learning algorithms for recognizing signals in distributed fiber optic monitoring and security systems for long perimeters. Synthesizing such detection algorithms poses a nontrivial research and development challenge, because these systems face stringent error (type I and II) requirements and operate in difficult signal-jamming environments, with intensive signal-like jamming and a variety of changing possible signal portraits of possible recognized events. To address these issues, we have developed a two-level event detection architecture, where the primary classifier is based on an ensemble of deep convolutional networks, can recognize 7 classes of signals and receives time-space data frames as input. Using real-life data, we have shown that the applied methods result in efficient and robust multiclass detection algorithms that have a high degree of adaptability.
HNP3: A Hierarchical Nonparametric Point Process for Modeling Content Diffusion over Social Media
Hosseini, Seyed Abbas, Khodadadi, Ali, Arabzade, Soheil, Rabiee, Hamid R.
This paper introduces a novel framework for modeling temporal events with complex longitudinal dependency that are generated by dependent sources. This framework takes advantage of multidimensional point processes for modeling time of events. The intensity function of the proposed process is a mixture of intensities, and its complexity grows with the complexity of temporal patterns of data. Moreover, it utilizes a hierarchical dependent nonparametric approach to model marks of events. These capabilities allow the proposed model to adapt its temporal and topical complexity according to the complexity of data, which makes it a suitable candidate for real world scenarios. An online inference algorithm is also proposed that makes the framework applicable to a vast range of applications. The framework is applied to a real world application, modeling the diffusion of contents over networks. Extensive experiments reveal the effectiveness of the proposed framework in comparison with state-of-the-art methods.
Deep unsupervised learning through spatial contrasting
Hoffer, Elad, Hubara, Itay, Ailon, Nir
Convolutional networks have marked their place over the last few years as the best performing model for various visual tasks. They are, however, most suited for supervised learning from large amounts of labeled data. Previous attempts have been made to use unlabeled data to improve model performance by applying unsupervised techniques. These attempts require different architectures and training methods. In this work we present a novel approach for unsupervised training of Convolutional networks that is based on contrasting between spatial regions within images. This criterion can be employed within conventional neural networks and trained using standard techniques such as SGD and back-propagation, thus complementing supervised methods.
Stealing Machine Learning Models via Prediction APIs
Tramèr, Florian, Zhang, Fan, Juels, Ari, Reiter, Michael K., Ristenpart, Thomas
Machine learning (ML) models may be deemed confidential due to their sensitive training data, commercial value, or use in security applications. Increasingly often, confidential ML models are being deployed with publicly accessible query interfaces. ML-as-a-service ("predictive analytics") systems are an example: Some allow users to train models on potentially sensitive data and charge others for access on a pay-per-query basis. The tension between model confidentiality and public access motivates our investigation of model extraction attacks. In such attacks, an adversary with black-box access, but no prior knowledge of an ML model's parameters or training data, aims to duplicate the functionality of (i.e., "steal") the model. Unlike in classical learning theory settings, ML-as-a-service offerings may accept partial feature vectors as inputs and include confidence values with predictions. Given these practices, we show simple, efficient attacks that extract target ML models with near-perfect fidelity for popular model classes including logistic regression, neural networks, and decision trees. We demonstrate these attacks against the online services of BigML and Amazon Machine Learning. We further show that the natural countermeasure of omitting confidence values from model outputs still admits potentially harmful model extraction attacks. Our results highlight the need for careful ML model deployment and new model extraction countermeasures.
Amazon's 2.5M 'Alexa Prize' seeks chatbot that can converse intelligently for 20 minutes
Amazon will award up to 2.5 million in prizes and sponsorships in a new annual competition called the "Alexa Prize"-- challenging university students to build a "socialbot" on the company's digital brain Alexa that can have an intelligent conversation about pop culture and news events. Amazon said teams can submit applications now, and the winners will be announced at AWS re:Invent in November 2017. The winning team will receive a 500,000 prize, and Amazon will give another 1 million to the winning team's university if its bot can converse coherently for 20 minutes. "Conversing for 20 minutes is difficult for most humans and an extraordinarily ambitious challenge for bots that are learning to converse like us," Dan Jurafsky, professor and chair of linguistics and professor of computer science at Stanford University. "The Alexa Prize will encourage student researchers to come up with great ideas for leveraging real-world conversational AI technologies like Alexa to create software that can converse as engagingly as humans."
Josh Fischel, founder of Music Tastes Good, dies at 47
Josh Fischel, the founder of last weekend's inaugural Music Tastes Good Festival in Long Beach, died Thursday afternoon of liver disease, according to festival organizers. The news came as a shock to the Long Beach and Southern California music communities, who just days ago saw Fischel overseeing the culmination of a life's work in local music. Though family and festival organizers knew he had been sick, no one knew how rapidly his disease would progress after the festival. Music Tastes Good was a three-day event in downtown Long Beach headlined by the likes of the Specials, Warpaint and the Squeeze, among many others. "He couldn't go more than a few feet in his golf cart without somebody stopping him to say'Hey, Josh!' " said Jon Halperin, the talent buyer and co-promoter of Music Tastes Good.