Deep Learning
MINIROCKET: A Very Fast (Almost) Deterministic Transform for Time Series Classification
Dempster, Angus, Schmidt, Daniel F., Webb, Geoffrey I.
Until recently, the most accurate methods for time series classification were limited by high computational complexity. ROCKET achieves state-of-the-art accuracy with a fraction of the computational expense of most existing methods by transforming input time series using random convolutional kernels, and using the transformed features to train a linear classifier. We reformulate ROCKET into a new method, MINIROCKET, making it up to 75 times faster on larger datasets, and making it almost deterministic (and optionally, with additional computational expense, fully deterministic), while maintaining essentially the same accuracy. Using this method, it is possible to train and test a classifier on all of 109 datasets from the UCR archive to state-of-the-art accuracy in less than 10 minutes. MINIROCKET is significantly faster than any other method of comparable accuracy (including ROCKET), and significantly more accurate than any other method of even roughly-similar computational expense. As such, we suggest that MINIROCKET should now be considered and used as the default variant of ROCKET.
Unsupervised Learning of Global Factors in Deep Generative Models
Peis, Ignacio, Olmos, Pablo M., Artรฉs-Rodrรญguez, Antonio
We present a novel deep generative model based on non i.i.d. variational autoencoders that captures global dependencies among observations in a fully unsupervised fashion. In contrast to the recent semi-supervised alternatives for global modeling in deep generative models, our approach combines a mixture model in the local or data-dependent space and a global Gaussian latent variable, which lead us to obtain three particular insights. First, the induced latent global space captures interpretable disentangled representations with no user-defined regularization in the evidence lower bound (as in $\beta$-VAE and its generalizations). Second, we show that the model performs domain alignment to find correlations and interpolate between different databases. Finally, we study the ability of the global space to discriminate between groups of observations with non-trivial underlying structures, such as face images with shared attributes or defined sequences of digits images.
Quality-Diversity Optimization: a novel branch of stochastic optimization
Chatzilygeroudis, Konstantinos, Cully, Antoine, Vassiliades, Vassilis, Mouret, Jean-Baptiste
Traditional optimization algorithms search for a single global optimum that maximizes (or minimizes) the objective function. Multimodal optimization algorithms search for the highest peaks in the search space that can be more than one. Quality-Diversity algorithms are a recent addition to the evolutionary computation toolbox that do not only search for a single set of local optima, but instead try to illuminate the search space. In effect, they provide a holistic view of how high-performing solutions are distributed throughout a search space. The main differences with multimodal optimization algorithms are that (1) Quality-Diversity typically works in the behavioral space (or feature space), and not in the genotypic (or parameter) space, and (2) Quality-Diversity attempts to fill the whole behavior space, even if the niche is not a peak in the fitness landscape. In this chapter, we provide a gentle introduction to Quality-Diversity optimization, discuss the main representative algorithms, and the main current topics under consideration in the community. Throughout the chapter, we also discuss several successful applications of Quality-Diversity algorithms, including deep learning, robotics, and reinforcement learning.
Risk-Monotonicity via Distributional Robustness
Mhammedi, Zakaria, Husain, Hisham
Acquisition of data is a difficult task in most applications of Machine Learning (ML), and it is only natural that one hopes and expects lower populating risk (better performance) with increasing data points. It turns out, somewhat surprisingly, that this is not the case even for the most standard algorithms such as the Empirical Risk Minimizer (ERM). Non-monotonic behaviour of the risk and instability in training have manifested and appeared in the popular deep learning paradigm under the description of double descent. These problems not only highlight our lack of understanding of learning algorithms and generalization but rather render our efforts at data acquisition in vain. It is, therefore, crucial to pursue this concern and provide a characterization of such behaviour. In this paper, we derive the first consistent and risk-monotonic algorithms for a general statistical learning setting under weak assumptions, consequently resolving an open problem (Viering et al. 2019) on how to avoid non-monotonic behaviour of risk curves. Our algorithms make use of Distributionally Robust Optimization (DRO) -- a technique that has shown promise in other complications of deep learning such as adversarial training. Our work makes a significant contribution to the topic of risk-monotonicity, which may be key in resolving empirical phenomena such as double descent.
Implicit bias of deep linear networks in the large learning rate phase
Huang, Wei, Du, Weitao, Da Xu, Richard Yi, Liu, Chunrui
Most theoretical studies explaining the regularization effect in deep learning have only focused on gradient descent with a sufficient small learning rate or even gradient flow (infinitesimal learning rate). Such researches, however, have neglected a reasonably large learning rate applied in most practical applications. In this work, we characterize the implicit bias effect of deep linear networks for binary classification using the logistic loss in the large learning rate regime, inspired by the seminal work by Lewkowycz et al. [26] in a regression setting with squared loss. They found a learning rate regime with a large stepsize named the catapult phase, where the loss grows at the early stage of training and eventually converges to a minimum that is flatter than those found in the small learning rate regime. We claim that depending on the separation conditions of data, the gradient descent iterates will converge to a flatter minimum in the catapult phase. We rigorously prove this claim under the assumption of degenerate data by overcoming the difficulty of the non-constant Hessian of logistic loss and further characterize the behavior of loss and Hessian for non-separable data. Finally, we demonstrate that flatter minima in the space spanned by non-separable data along with the learning rate in the catapult phase can lead to better generalization empirically.
Top 5 Machine Learning Highlights from Formnext 2020 โ AMEXCI
Oqton keeps adding new intelligent features to their Artificial Intelligence powered platform, FactoryOS. Oqton is building partnerships with machine manufacturer that open their API, and has announced their partnership with EOS. We believe such system is key in moving toward a more automated workflow, and we are in the process of testing it out with our own machines. We hope in the future to be able to plug into it deep learning based modules for live defect analysis. Click here to find out more about this technology.
DeepMind's latest AI breakthrough could turbocharge drug discovery
While impressive, the technology wasn't yet capable of replacing the existing expensive and time-consuming experimental methods for determining what these proteins look like. However, its latest software comes close. In November, AlphaFold again outperformed all the other competing groups at CASP. The technology solved protein structures other labs had been working on for years. Scientists think the technology could have immense implications for the way proteins are studied.
What is Edge Machine Learning?
Edge Machine Learning (Edge ML) is one of the most talked-about tech advancements since the Internet of Things (IoT), and for a good reason. With the rise of IoT came an explosion of Smart Devices connected to the Cloud, but the network was not yet ready to support this surge in demand. Cloud networks were congested, and companies overlooked key issues with Cloud computing, such as security. So, what is Edge ML anyway? Edge ML is a technique by which Smart Devices can process data locally (either using local servers or at the device-level) using machine and deep learning algorithms, reducing reliance on Cloud networks.
Why some artificial intelligence is smart until it's dumb
Starfleet's star android, Lt. Commander Data, has been enlisted by his renegade android "brother" Lore to join a rebellion against humankind -- much to the consternation of Jean-Luc Picard, captain of the USS Enterprise. "The reign of biological life-forms is coming to an end," Lore tells Picard. "You, Picard, and those like you, are obsolete." In real life, the era of smart machines has already arrived. They haven't completely taken over the world yet, but they're off to a good start. "Machine learning" -- a sort of concrete subfield within the more nebulous quest for artificial intelligence -- has invaded numerous fields of human endeavor, from medical diagnosis to searching for new subatomic particles.
DeepMind AI Predicts Protein Structure
If you are even remotely interested in science, you will have probably already heard about DeepMind's latest leap. Their AI system Alphafold 2 has cracked predicting proteins' 3D structure. There are plenty of great articles about it. Since I have written about machine learning/AI in an earlier series of posts, I decided to write a brief post about this development as well. For more details, do check the Nature/New Scientist/DeepMind articles linked above.