Africa
The Amazing Ways Chinese Face Recognition Company Megvii (Face ) Uses AI And Machine Vision
Megvii Technology, a Chinese company, founded in 2011 and widely known for its Face system, is one of the world leaders in facial recognition and artificial intelligence technology. While they might be best known for Face, Megvii uses artificial intelligence and machine vision in a variety of amazing ways. Megvii was the concept conceived by friends and Tsinghua University graduates Yin Qui, Yang Mu, and Tang Wenbin. After tremendous success in China (especially since they were able to train algorithms from China's vast pool of data) with clients such as Ant Financial, Vivo (smartphones), Didi Chuxing (ride-sharing) and investments from Bank of China, the State-Owned Venture Capital Fund, China-Russian Investment Fund and other private investors including Ant Financial (Alibaba's payment affiliate), Megvii is ready to go global. They have projects slated in the coming year for Japan, Europe, the Middle East, Southeast Asia, and the United States and have secured a distributor in Thailand.
Using Deep Networks and Transfer Learning to Address Disinformation
Dhamani, Numa, Azunre, Paul, Gleason, Jeffrey L., Corcoran, Craig, Honke, Garrett, Kramer, Steve, Morgan, Jonathon
We also demonstrate the the detection of inflammatory, inauthentic, or otherwise ability to use this architecture to transfer knowledge nefarious communication. Character-level convolutional from labeled data in one domain to related neural networks (CNNs) are particularly well-suited for (supervised and unsupervised) tasks. Characterlevel this task--as opposed to a word-level model--because they neural networks and transfer learning are allow for non-vernacular discourse, misspelling, and other particularly valuable tools in the disinformation social media features (e.g., emoticons) to be learned without space because of the messy nature of social media, the constraint of fixed vocabularies (Zhang et al., 2015). We lack of labeled data, and the multi-channel tactics implement an adaptation of a neural network architecture of influence campaigns. We demonstrate their effectiveness recently demonstrated to be effective for text classification in several tasks relevant for detecting (Zhang et al., 2015; Jรณzefowicz et al., 2016). The method disinformation: spam emails, review bombing, is purely content-based and does not require any additional political sentiment, and conversation clustering.
Bayesian Tensorized Neural Networks with Automatic Rank Selection
Tensor decomposition is an effective approach to compress over-parameterized neural networks and to enable their deployment on resource-constrained hardware platforms. However, directly applying tensor compression in the training process is a challenging task due to the difficulty of choosing a proper tensor rank. In order to achieve this goal, this paper proposes a Bayesian tensorized neural network. Our Bayesian method performs automatic model compression via an adaptive tensor rank determination. We also present approaches for posterior density calculation and maximum a posteriori (MAP) estimation for the end-to-end training of our tensorized neural network. We provide experimental validation on a fully connected neural network, a CNN and a residual neural network where our work produces $7.4\times$ to $137\times$ more compact neural networks directly from the training.
Loss Surface Modality of Feed-Forward Neural Network Architectures
Bosman, Anna Sergeevna, Engelbrecht, Andries, Helbig, Mardรฉ
It has been argued in the past that high-dimensional neural networks do not exhibit local minima capable of trapping an optimisation algorithm. However, the relationship between loss surface modality and the neural architecture parameters, such as the number of hidden neurons per layer and the number of hidden layers, remains poorly understood. This study employs fitness landscape analysis to study the modality of neural network loss surfaces under various feed-forward architecture settings. An increase in the problem dimensionality is shown to yield a more searchable and more exploitable loss surface. An increase in the hidden layer width is shown to effectively reduce the number of local minima, and simplify the shape of the global attractor. An increase in the architecture depth is shown to sharpen the global attractor, thus making it more exploitable.
Doctor of Crosswise: Reducing Over-parametrization in Neural Networks
Dr. of Crosswise proposes a new architecture to reduce over-parametrization in Neural Networks. It introduces an operand for rapid computation in the framework of Deep Learning that leverages learned weights. The formalism is described in detail providing both an accurate elucidation of the mechanics and the theoretical implications.
Leader Stochastic Gradient Descent for Distributed Training of Deep Learning Models
Teng, Yunfei, Gao, Wenbo, Chalus, Francois, Choromanska, Anna, Goldfarb, Donald, Weller, Adrian
We consider distributed optimization under communication constraints for training deep learning models. We propose a new algorithm, whose parameter updates rely on two forces: a regular gradient step, and a corrective direction dictated by the currently best-performing worker (leader). Our method differs from the parameter-averaging scheme EASGD in a number of ways: (i) our objective formulation does not change the location of stationary points compared to the original optimization problem; (ii) we avoid convergence decelerations caused by pulling local workers descending to different local minima to each other (i.e. to the average of their parameters); (iii) our update by design breaks the curse of symmetry (the phenomenon of being trapped in poorly generalizing sub-optimal solutions in symmetric non-convex landscapes); and (iv) our approach is more communication efficient since it broadcasts only parameters of the leader rather than all workers. We provide theoretical analysis of the batch version of the proposed algorithm, which we call Leader Gradient Descent (LGD), and its stochastic variant (LSGD). Finally, we implement an asynchronous version of our algorithm and extend it to the multi-leader setting, where we form groups of workers, each represented by its own local leader (the best performer in a group), and update each worker with a corrective direction comprised of two attractive forces: one to the local, and one to the global leader (the best performer among all workers). The multi-leader setting is well-aligned with current hardware architecture, where local workers forming a group lie within a single computational node and different groups correspond to different nodes. For training convolutional neural networks, we empirically demonstrate that our approach compares favorably to state-of-the-art baselines.
Introduction to Anomaly Detection using Machine Learning with a Case Study
A common need when you are analyzing real-world data-sets is determining which data point stand out as being different to all others data points. Such data points are known as anomalies. This article was originally published on Medium by Davis David. In this article, you will learn a couple of Machine Learning-Based Approaches for Anomaly Detection and then show how to apply one of these approaches to solve a specific use case for anomaly detection (Credit Fraud detection) in part two. A common need when you analyzing real-world data-sets is determining which data point stand out as being different to all others data points.
Armed with artificial intelligence, scientists take on climate change
Science needs to understand and predict how climate change--and the growing onslaught of hurricanes, fires, and floods it's bringing--affects tropical forests. Will the forests respond to the assault with shorter trees? Will they store less carbon, or support less tree and plant diversity and fewer wildlife species? To better understand the effects a changing climate will have on tropical forests, Maria Uriarte, Columbia University professor of ecology, evolution, and environmental biology, needs to analyze images of forests. These bird's-eye view images are the size of a postage stamp.
Predicting Sparse Clients' Actions with CPOPT-Net in the Banking Environment
Charlier, Jeremy, State, Radu, Hilger, Jean
The digital revolution of the banking system with evolving European regulations have pushed the major banking actors to innovate by a newly use of their clients' digital information. Given highly sparse client activities, we propose CPOPT-Net, an algorithm that combines the CP canonical tensor decomposition, a multidimensional matrix decomposition that factorizes a tensor as the sum of rank-one tensors, and neural networks. CPOPT-Net removes efficiently sparse information with a gradient-based resolution while relying on neural networks for time series predictions. Our experiments show that CPOPT-Net is capable to perform accurate predictions of the clients' actions in the context of personalized recommendation. CPOPT-Net is the first algorithm to use non-linear conjugate gradient tensor resolution with neural networks to propose predictions of financial activities on a public data set.
Deep Fuzzy Systems
ABSTRACT-An investigation of deep fuzzy systems is presented in this paper. A deep fuzzy system is represented by recursive fuzzy systems from an input terminal to output terminal. Recursive fuzzy systems are sequences of fuzzy grade memberships obtained using fuzzy transmition functions and recursive calls to fuzzy systems. A recursive fuzzy system which calls a fuzzy system times includes fuzzy chains to evaluate the final grade membership of this recursive system. A connection matrix which includes recursive calls are used to represent recursive fuzzy systems.