Deep Learning
SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair
Chen, Zimin, Kommrusch, Steve, Tufano, Michele, Pouchet, Louis-Noël, Poshyvanyk, Denys, Monperrus, Martin
This paper presents a novel end-to-end approach to program repair based on sequence-to-sequence learning. We devise, implement, and evaluate a system, called SequenceR, for fixing bugs based on sequence-to-sequence learning on source code. This approach uses the copy mechanism to overcome the unlimited vocabulary problem that occurs with big code. Our system is data-driven; we train it on 35,578 samples, carefully curated from commits to open-source repositories. We evaluate it on 4,711 independent real bug fixes, as well on the Defects4J benchmark used in program repair research. SequenceR is able to perfectly predict the fixed line for 950/4711 testing samples, and find correct patches for 14 bugs in Defects4J. It captures a wide range of repair operators without any domain-specific top-down design.
Deep Uncertainty Quantification: A Machine Learning Approach for Weather Forecasting
Wang, Bin, Lu, Jie, Yan, Zheng, Luo, Huaishao, Li, Tianrui, Zheng, Yu, Zhang, Guangquan
Weather forecasting is usually solved through numerical weather prediction (NWP), which can sometimes lead to unsatisfactory performance due to inappropriate setting of the initial states. In this paper, we design a data-driven method augmented by an effective information fusion mechanism to learn from historical data that incorporates prior knowledge from NWP. We cast the weather forecasting problem as an end-to-end deep learning problem and solve it by proposing a novel negative log-likelihood error (NLE) loss function. A notable advantage of our proposed method is that it simultaneously implements single-value forecasting and uncertainty quantification, which we refer to as deep uncertainty quantification (DUQ). Efficient deep ensemble strategies are also explored to further improve performance. This new approach was evaluated on a public dataset collected from weather stations in Beijing, China. Experimental results demonstrate that the proposed NLE loss significantly improves generalization compared to mean squared error (MSE) loss and mean absolute error (MAE) loss. Compared with NWP, this approach significantly improves accuracy by 47.76%, which is a state-of-the-art result on this benchmark dataset. The preliminary version of the proposed method won 2nd place in an online competition for daily weather forecasting.
Deep learning hope and hype: MIT Technology Review's Will Knight
Both the progress and the hype around cutting-edge machine learning techniques were on vivid display at the December 2018 NeurIPS Conference in Montreal, Quebec, says Will Knight, MIT Technology Review's senior editor for artificial intelligence. One big question hanging over the meeting, he says, was how to detect and reverse the sexism, racism, and other forms of bias that seep into machine-learning algorithms that train themselves using real-world data. Participants also previewed the coming generation of chips designed specifically to support deep learning--a field where US manufacturers face growing competition from China. Separately, Will looks to the most exciting AI trends for 2019, including the generative adversarial networks (GANs) being used to generate authentic-looking photos and videos. This episode is sponsored by PwC, a global consulting firm in 158 countries with more than 250,000 people. PwC transforms business outcomes and results, helping companies use digital and emerging tech to reimagine their business, from strategy and operations to tax and finance. In the second half of the show, Scott Likens, PwC's New Services and Emerging Tech Leader, shares details from a new PwC study on the main trends in artificial intelligence that business leaders need to know about in 2019. Business Lab is hosted by Elizabeth Bramson-Boudreau, the CEO and publisher of MIT Technology Review. The show is produced by Wade Roush, with editorial help from Mindy Blodgett. Will Knight: "China has never had a real chip industry. Making AI chips could change that." PwC 2019 AI Predictions: Six AI priorities you can't afford to ignore Elizabeth Bramson-Boudreau: From MIT Technology Review, I'm Elizabeth Bramson-Boudreau, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace.
New AI tech reshapes skin cancer detection
Created by FotoFinder Systems, Moleanalyzer pro is a portal that lets physicians confirm their skin cancer diagnosis using evaluation techniques, combining specialist expertise with AI and including the option of receiving a second opinion from international skin cancer experts. FotoFinder Systems Global Brand Director Kathrin Niemela told HITNA that the technology aims to aid skin cancer diagnoses. According to the Cancer Council Australia, every year skin cancers account for around 80 per cent of all newly diagnosed cancers in Australia, with GPs seeing more than a million patients per year for skin cancer. In addition, the Australian Government identified that there were 14,320 new cases of melanoma skin cancer diagnosed in 2018, accounting for 10.4 per cent of all new cancer cases diagnosed. "The earlier skin cancer is detected, the better the prognosis. The leisure behaviour of sunbathing in many parts of the world makes early detection of skin cancer more important worldwide," Niemela said.
Xfer: an open-source library for neural network transfer learning
Transfer learning is a set of techniques for reusing and repurposing already trained machine learning models in new situations. It brings particular advantages in the domain of deep learning, where training a model from scratch (rather than reusing an existing model) requires a lot of computational and data resources, as well as expertise. This blog post contains a quick overview of transfer learning through the introduction of Xfer, an open-source library that enables easy application and prototyping of transfer learning approaches. Neural networks are machine learning models that learn functions and patterns from data. They underpin numerous modern AI-enabled technologies with applications in conversational agents, self-driving cars, self-learning agents that play board games and many more.
AI Deep-Learning Program Places Steve Buscemi's Face On Jennifer Lawrence In Golden Globes Video
This is a video created using a deep-learning artificial intelligence program that placed Steve Buscemi's face on Jennifer Lawrence's head and body while she was speaking at the 2016 Golden Globe awards. So, if you were wondering if we've gone too far the answer is yes -- we're already over the edge of the cliff like Wyle E. Coyote and just haven't realized it yet. The moment we look down it's all over. Keep going for the video. Thanks to Allyson S, who agrees there are some things best left unseen.
Two AIs Go Head-to-Head on Atari's 'Breakout' to Test Deep Learning
It seems like every day brings a new AI more capable than the last. This was recently apparent with AlphaGo--it was pretty great at beating Breakout, then Google got involved and soon it was capable of beating the world's leading Go champion. To do this, AlphaGo uses what is known as'deep reinforcement learning'. For example, in Breakout, it will take raw image frames of the game as it's being played. Whether or not the ball is hitting the bricks in those frames will decide whether or not positive reinforcement is registered.
Learned Indexes for Dynamic Workloads
Tang, Chuzhe, Dong, Zhiyuan, Wang, Minjie, Wang, Zhaoguo, Chen, Haibo
The recent proposal of learned index structures opens up a new perspective on how traditional range indexes can be optimized. However, the current learned indexes assume the data distribution is relatively static and the access pattern is uniform, while real-world scenarios consist of skew query distribution and evolving data. In this paper, we demonstrate that the missing consideration of access patterns and dynamic data distribution notably hinders the applicability of learned indexes. To this end, we propose solutions for learned indexes for dynamic workloads (called Doraemon). To improve the latency for skew queries, Doraemon augments the training data with access frequencies. To address the slow model re-training when data distribution shifts, Doraemon caches the previously-trained models and incrementally fine-tunes them for similar access patterns and data distribution. Our preliminary result shows that, Doraemon improves the query latency by 45.1% and reduces the model re-training time to 1/20.
Parameter-Efficient Transfer Learning for NLP
Houlsby, Neil, Giurgiu, Andrei, Jastrzebski, Stanislaw, Morrone, Bruna, de Laroussilhe, Quentin, Gesmundo, Andrea, Attariyan, Mona, Gelly, Sylvain
Fine-tuning large pre-trained models is an effective transfer mechanism in NLP. However, in the presence of many downstream tasks, fine-tuning is parameter inefficient: an entire new model is required for every task. As an alternative, we propose transfer with adapter modules. Adapter modules yield a compact and extensible model; they add only a few trainable parameters per task, and new tasks can be added without revisiting previous ones. The parameters of the original network remain fixed, yielding a high degree of parameter sharing. To demonstrate adapter's effectiveness, we transfer the recently proposed BERT Transformer model to 26 diverse text classification tasks, including the GLUE benchmark. Adapters attain near state-of-the-art performance, whilst adding only a few parameters per task. On GLUE, we attain within 0.4% of the performance of full fine-tuning, adding only 3.6% parameters per task. By contrast, fine-tuning trains 100% of the parameters per task.
Asymmetric Valleys: Beyond Sharp and Flat Local Minima
He, Haowei, Huang, Gao, Yuan, Yang
Despite the non-convex nature of their loss functions, deep neural networks are known to generalize well when optimized with stochastic gradient descent (SGD). Recent work conjectures that SGD with proper configuration is able to find wide and flat local minima, which have been proposed to be associated with good generalization performance. In this paper, we observe that local minima of modern deep networks are more than being flat or sharp. Specifically, at a local minimum there exist many asymmetric directions such that the loss increases abruptly along one side, and slowly along the opposite side--we formally define such minima as asymmetric valleys. Under mild assumptions, we prove that for asymmetric valleys, a solution biased towards the flat side generalizes better than the exact minimizer. Further, we show that simply averaging the weights along the SGD trajectory gives rise to such biased solutions implicitly. This provides a theoretical explanation for the intriguing phenomenon observed by Izmailov et al. (2018). In addition, we empirically find that batch normalization (BN) appears to be a major cause for asymmetric valleys.