Goto

Collaborating Authors

 Deep Learning


Regularization via Adaptive Pairwise Label Smoothing

arXiv.org Machine Learning

Label Smoothing (LS) is an effective regularizer to improve the generalization of state-of-the-art deep models. For each training sample the LS strategy smooths the one-hot encoded training signal by distributing its distribution mass over the non ground-truth classes, aiming to penalize the networks from generating overconfident output distributions. This paper introduces a novel label smoothing technique called Pairwise Label Smoothing (PLS). The PLS takes a pair of samples as input. Smoothing with a pair of ground-truth labels enables the PLS to preserve the relative distance between the two truth labels while further soften that between the truth labels and the other targets, resulting in models producing much less confident predictions than the LS strategy. Also, unlike current LS methods, which typically require to find a global smoothing distribution mass through cross-validation search, PLS automatically learns the distribution mass for each input pair during training. We empirically show that PLS significantly outperforms LS and the baseline models, achieving up to 30% of relative classification error reduction. We also visually show that when achieving such accuracy gains the PLS tends to produce very low winning softmax scores.


Second-Order Guarantees in Federated Learning

arXiv.org Machine Learning

Federated learning is a useful framework for centralized learning from distributed data under practical considerations of heterogeneity, asynchrony, and privacy. Federated architectures are frequently deployed in deep learning settings, which generally give rise to non-convex optimization problems. Nevertheless, most existing analysis are either limited to convex loss functions, or only establish first-order stationarity, despite the fact that saddle-points, which are first-order stationary, are known to pose bottlenecks in deep learning. We draw on recent results on the second-order optimality of stochastic gradient algorithms in centralized and decentralized settings, and establish second-order guarantees for a class of federated learning algorithms.


Deep learning based numerical approximation algorithms for stochastic partial differential equations and high-dimensional nonlinear filtering problems

arXiv.org Machine Learning

In this article we introduce and study a deep learning based approximation algorithm for solutions of stochastic partial differential equations (SPDEs). In the proposed approximation algorithm we employ a deep neural network for every realization of the driving noise process of the SPDE to approximate the solution process of the SPDE under consideration. We test the performance of the proposed approximation algorithm in the case of stochastic heat equations with additive noise, stochastic heat equations with multiplicative noise, stochastic Black--Scholes equations with multiplicative noise, and Zakai equations from nonlinear filtering. In each of these SPDEs the proposed approximation algorithm produces accurate results with short run times in up to 50 space dimensions.


Aligning Hyperbolic Representations: an Optimal Transport-based approach

arXiv.org Machine Learning

Hyperbolic embeddings are state-of-the-art models to learn representations of data with an underlying hierarchical structure [18]. The hyperbolic space serves as a geometric prior to hierarchical structures, tree graphs, heavy-tailed distributions, e.g., scale-free, powerlaw [45]. A relevant tool to implement hyperbolic space algorithms is the Möbius gyrovector spaces or Gyrovector spaces [66]. Gyrovector spaces are an algebraic formalism, which leads to vector-like operations, i.e., gyrovector, in the Poincaré model of the hyperbolic space. Thanks to this formalism, we can quickly build estimators that are well-suited to perform end-to-end optimization [6]. Gyrovector spaces are essential to design the hyperbolic version of several machine learning algorithms, like Hyperbolic Neural Networks (HNN) [24], Hyperbolic Graph NN [36], Hyperbolic Graph Convolutional NN [12], learning latent feature representations [41, 46], word embeddings [62, 25], and image embeddings [29]. Modern machine learning algorithms rely on the availability to accumulate large volumes of data, often coming from various sources, e.g., acquisition devices or languages. However, these massive amounts of heterogeneous data can entangle downstream learning tasks since the data may follow different distributions. Alignment aims at building connections between two or more disparate data sets by aligning their underlying manifolds.


About contrastive unsupervised representation learning for classification and its convergence

arXiv.org Machine Learning

The aim of this work is to provide additional theoretical guarantees for contrastive learning (van den Oord et al., 2018), which corresponds to methods allowing to learn useful data representations in an unsupervised setting. Unsupervised representation learning was initially approached with a fair amount of success by training through the minimization of losses coming from "pretext" tasks, a technique known as self-supervision (Doersch and Zisserman, 2017), where labels can be automatically constructed. Notable examples of pretext tasks in computer vision include colorization (Zhang et al., 2016), transformation prediction (Gidaris et al., 2018; Dosovitskiy et al., 2014) or predicting patch relative positions (Doersch et al., 2015). Some theoretical guarantees (Lee et al., 2020) were recently proposed to support training on pretext tasks. Contrastive learning is also known to be very effective for pretraining supervised methods (Chen et al., 2020a,b; Grill et al., 2020; Caron et al., 2020), where we can observe that, quite surprisingly, the gap between unsupervised and supervised performance has been closed for tasks such as image classification: the use of a pretrained image encoder on top of simple classification layers, that are trained on a fraction of the labels available, allows to achieve an accuracy comparable to that of a fully supervised end-to-end training (Hénaff et al., 2019; Grill et al., 2020).


The Next Generation Of Artificial Intelligence

#artificialintelligence

For the second part of this article series, see here. It has only been 8 years since the modern era of deep learning began at the 2012 ImageNet competition. Progress in the field since then has been breathtaking and relentless. If anything, this breakneck pace is only accelerating. Five years from now, the field of AI will look very different than it does today.


AWS previews ultra-efficient AI instances for neural network training - SiliconANGLE

#artificialintelligence

The cloud giant is introducing the Gaudi instances at an opportune time. AI models are getting more complex, partially because enterprise machine learning initiatives are maturing and partially because research conducted by the likes of OpenAI is facilitating bigger neural network architectures. As neural networks grow in complexity, the amount of computing power necessary to train them is increasing and fueling demand for more efficient training infrastructure.


Deep learning helps robots grasp and move objects with ease

#artificialintelligence

In the past year, lockdowns and other COVID-19 safety measures have made online shopping more popular than ever, but the skyrocketing demand is leaving many retailers struggling to fulfill orders while ensuring the safety of their warehouse employees. Researchers at the University of California, Berkeley, have created new artificial intelligence software that gives robots the speed and skill to grasp and smoothly move objects, making it feasible for them to soon assist humans in warehouse environments. The technology is described in a paper published online today (Wednesday, Nov. 18) in the journal Science Robotics. Automating warehouse tasks can be challenging because many actions that come naturally to humans--like deciding where and how to pick up different types of objects and then coordinating the shoulder, arm and wrist movements needed to move each object from one location to another--are actually quite difficult for robots. Robotic motion also tends to be jerky, which can increase the risk of damaging both the products and the robots. "Warehouses are still operated primarily by humans, because it's still very hard for robots to reliably grasp many different objects," said Ken Goldberg, William S. Floyd Jr. Distinguished Chair in Engineering at UC Berkeley and senior author of the study.


Taiwanese team develops AI system that can detect pancreatic cancer - Focus Taiwan

#artificialintelligence

Taipei, Oct. 28 (CNA) A team at National Taiwan University Hospital (NTUH) has developed an artificial intelligence (AI) system that can identify tumors in the pancreas with an accuracy of over 90 percent. At a press conference on Tuesday where the team introduced the technology, NTUH doctor Liao Wei-chih (廖偉智) said that pancreatic cancer was the seventh deadliest type of cancer in Taiwan in 2019, causing nearly 2,500 deaths that year. The disease is extremely hard to detect, however, as patients experience no symptoms during the early stages, and studies have found that 40 percent of pancreatic tumors that are smaller than 2 centimeters are missed when doctors use CT scans, Liao said. This is because these small tumors do not look like lumps, but appear to be a thin layer of gray film, Liao explained, which is a challenge for even the most experienced of experts to identify. As a result of these difficulties, patients are often only diagnosed when the cancer has spread to other parts of the body, thus complicating treatment, he said.


8 Best Free Resources To Learn Deep Reinforcement Learning Using TensorFlow

#artificialintelligence

With the success of DeepMind's AlphaGo system defeating the world Go champion, reinforcement learning has achieved significant attention among researchers and developers. Deep reinforcement learning has become one of the most significant techniques in AI that is also being used by the researchers in order to attain artificial general intelligence. Below here is a list of 10 best free resources, in no particular order to learn deep reinforcement learning using TensorFlow. About: This tutorial "Introduction to RL and Deep Q Networks" is provided by the developers at TensorFlow. The topics include an introduction to deep reinforcement learning, the Cartpole Environment, introduction to DQN agent, Q-learning, Deep Q-Learning, DQN on Cartpole in TF-Agents and more.