Goto

Collaborating Authors

 Deep Learning


Dynamic Local Regret for Non-convex Online Forecasting

arXiv.org Machine Learning

We consider online forecasting problems for non-convex machine learning models. Forecasting introduces several challenges such as (i) frequent updates are necessary to deal with concept drift issues since the dynamics of the environment change over time, and (ii) the state of the art models are non-convex models. We address these challenges with a novel regret framework. Standard regret measures commonly do not consider both dynamic environment and non-convex models. We introduce a local regret for non-convex models in a dynamic environment. We present an update rule incurring a cost, according to our proposed local regret, which is sublinear in time T. Our update uses time-smoothed gradients. Using a real-world dataset we show that our time-smoothed approach yields several benefits when compared with state-of-the-art competitors: results are more stable against new data; training is more robust to hyperparameter selection; and our approach is more computationally efficient than the alternatives.


Global Sparse Momentum SGD for Pruning Very Deep Neural Networks

arXiv.org Machine Learning

Deep Neural Network (DNN) is powerful but computationally expensive and memory intensive, thus impeding its practical usage on resource-constrained front-end devices. DNN pruning is an approach for deep model compression, which aims at eliminating some parameters with tolerable performance degradation. In this paper, we propose a novel momentum-SGD-based optimization method to reduce the network complexity by on-the-fly pruning. Concretely, given a global compression ratio, we categorize all the parameters into two parts at each training iteration which are updated using different rules. In this way, we gradually zero out the redundant parameters, as we update them using only the ordinary weight decay but no gradients derived from the objective function. As a departure from prior methods that require heavy human works to tune the layer-wise sparsity ratios, prune by solving complicated non-differentiable problems or finetune the model after pruning, our method is characterized by 1) global compression that automatically finds the appropriate per-layer sparsity ratios; 2) end-to-end training; 3) no need for a time-consuming re-training process after pruning; and 4) superior capability to find better winning tickets which have won the initialization lottery.


On the Cross-lingual Transferability of Monolingual Representations

arXiv.org Artificial Intelligence

State-of-the-art unsupervised multilingual models (e.g., multilingual BERT) have been shown to generalize in a zero-shot cross-lingual setting. This generalization ability has been attributed to the use of a shared subword vocabulary and joint training across multiple languages giving rise to deep multilingual abstractions. We evaluate this hypothesis by designing an alternative approach that transfers a monolingual model to new languages at the lexical level. More concretely, we first train a transformer-based masked language model on one language, and transfer it to a new language by learning a new embedding matrix with the same masked language modeling objective--freezing parameters of all other layers. This approach does not rely on a shared vocabulary or joint training. However, we show that it is competitive with multilingual BERT on standard cross-lingual classification benchmarks and on a new Cross-lingual Question Answering Dataset (XQuAD). Our results contradict common beliefs of the basis of the generalization ability of multilingual models and suggest that deep monolingual models learn some abstractions that generalize across languages. We also release XQuAD as a more comprehensive cross-lingual benchmark, which comprises 240 paragraphs and 1190 question-answer pairs from SQuAD v1.1 translated into ten languages by professional translators.


Deep Reinforcement Learning in HOL4

arXiv.org Artificial Intelligence

The paper describes an implementation of deep reinforcement learning through self-supervised learning within the proof assistant HOL4. A close interaction between the machine learning modules and the HOL4 library is achieved by the choice of tree neural networks (TNNs) as machine learning models and the internal use of HOL4 terms to represent tree structures of TNNs. Recursive improvement is possible when a given task is expressed as a search problem. In this case, a Monte Carlo Tree Search (MCTS) algorithm guided by a TNN can be used to explore the search space and produce better examples for training the next TNN. As an illustration, tasks over propositional and arithmetical terms, representative of fundamental theorem proving techniques, are specified and learned: truth estimation, end-to-end computation, term rewriting and term synthesis.


Computing Research at Tata Consultancy Services

Communications of the ACM

Further, TCS eats its own dog food--it has deployed a home-grown enterprise social media platform7 and more recently a deep-learning-based conversational system across all its 400K employees.c There have been many lessons learned along this journey; we mention some critical ones here: First, research initiatives have always preceded their applicability, and so continuing to invest in research areas seemingly unrelated to the current business pays off in initially unforeseen ways. For example, research in genome-based early prediction of rare diseases4 later enabled TCS' business to build genome analysis pipelines for pharma customers. Deep expertise in computational chemistry3 is now allowing TCS to design new chemical formulations and molecules for customers, a very different kind of service that could potentially expand the very scope of its core business in the future.


Micron Introduces Comprehensive AI Development Platform

#artificialintelligence

SAN FRANCISCO, Oct. 24, 2019 (GLOBE NEWSWIRE) -- MICRON INSIGHT -- Micron Technology, Inc. (MU), today announced a powerful new set of high-performance hardware and software tools for deep learning applications with the acquisition of FWDNXT, a software and hardware startup. When combined with advanced Micron memory, FWDNXT's (pronounced "forward next") artificial intelligence (AI) hardware and software technology enables Micron to explore deep learning solutions required for data analytics, particularly in IoT and edge computing. With this acquisition, Micron is integrating compute, memory, tools and software into a comprehensive AI development platform. This platform in turn provides the key building blocks required to explore innovative memory optimized for AI workloads. "FWDNXT is an architecture designed to create fast-time-to-market edge AI solutions through an extremely easy to use software framework with broad modeling support and flexibility," said Micron Executive Vice President and Chief Business Officer Sumit Sadana.


Your guide to artificial Intelligence and machine learning at re:Invent 2019 Amazon Web Services

#artificialintelligence

With less than 40 days to re:Invent 2019, the excitement is building up and we are looking forward to seeing you all soon! Continuing our journey on artificial intelligence and machine learning, we are bringing a lot of technical content this year, with over 200 breakout sessions, deep-dive chalk talks, hands-on exercises with workshops featuring Amazon SageMaker, AWS DeepRacer, and deep learning frameworks such as TensorFlow, PyTorch, and more. You'll hear from many customers including Vanguard, BBC, Autodesk, British Airways, Fannie Mae, Thermo Fisher, Intuit, and many more. We are also hosting the Machine Learning Summit again this year, where you will hear from researchers and entrepreneurs about the latest breakthroughs today and the future possibilities tomorrow. To get you started on planning, here are a few highlights for the AI and ML sessions from the re:Invent 2019 session catalog.


Keras vs. tf.keras: What's the difference in TensorFlow 2.0? - PyImageSearch

#artificialintelligence

In this tutorial you'll discover the difference between Keras and tf.keras, including what's new in TensorFlow 2.0. Today's tutorial is inspired from an email I received last Tuesday from PyImageSearch reader, Jeremiah. Hi Adrian, I saw that TensorFlow 2.0 was released a few days ago. TensorFlow developers seem to be promoting Keras, or rather, something called tf.keras, as the recommended high-level API for TensorFlow 2.0. But I thought Keras was its own separate package?


The Future of Computation for Machine Learning and Data Science

#artificialintelligence

Deep learning has become ubiquitous in the modern world, with wide-ranging applications in nearly every field. As might be expected, people have started to notice, and the hype behind deep learning continues to increase as its widespread adoption by businesses occurs. Deep learning has been hugely successful for some important tasks in speech recognition, computer vision, and text understanding. The major downside of deep learning is its computational intensity, requiring high-performance computational resources and long training times. For facial recognition and image reconstructions, this also means working with low-resolution images.


Spell - Join the Movement to Accelerate Machine Learning

#artificialintelligence

Community authors will receive $250 cloud GPU credits on the Spell platform in exchange for creating content that will live on the Spell Community site. This contribution from community authors helps grow Spell's library of data science, machine learning, and deep learning tutorials and educate others on the world of AI.