Genre
Revolution AI: Why everyone wants in to Montreal's deep-learning hub
All eyes are on Montreal these days as a hub for deep learning. "Clearly it's a place where everybody wants to be if we want to tap into that talent," says Nagraj Kashyap, corporate vice-president of Microsoft Ventures in San Francisco. Montreal's pre-eminence as a deep learning centre can largely be attributed to the efforts of Yoshua Bengio, considered to be one of the three "co-fathers" of deep learning technology. Bengio not only engaged in cutting-edge research at the Université de Montréal long before deep learning was considered viable; his work has spawned an ecosystem that many say is unrivalled in the artificial intelligence (AI) world. That ecosystem includes the Montreal Institute for Learning Algorithms (MILA) which has been funded by government and private sector parties, including Google and Microsoft, among other tech notables.
Flipboard on Flipboard
In the future, it's likely that many aspects of human society will be controlled -- either partly or wholly -- by artificial intelligence. AI computer agents could manage systems from the quotidian (e.g., traffic lights) to the complex (e.g., a nation's whole economy), but leaving aside the problem of whether or not they can do their jobs well, there is another challenge: will these agents be able to play nice with one another? What happens if one AI's aims conflict with another's? Will they fight, or work together? Google's AI subsidiary DeepMind has been exploring this problem in a new study published today.
Elementary, my dear Watson! IBM now using AI platform to solve cybercrimes
IBM's Watson AI and IoT-fuelled supercomputer finally has the role it was born for, as it is now being recruited to tackle cybercrime. Since it was first presented to the world as a contestant on the gameshow Jeopardy!, IBM Watson has evolved 10-fold. From its new internet of things (IoT) headquarters in Munich, it is being used not only to develop the brains for future autonomous vehicles, but also to influence decision-making, from healthcare to smart cities. Now, similar to the Dr Watson character in the legendary Sherlock Holmes books, IBM Watson is to be recruited to solve crimes – specifically, cybercrimes. Over the past year, Watson has been trained in the language of cybersecurity, ingesting more than 1m security documents.
China's first 'deep learning lab' intensifies challenge to US in artificial intelligence race
Beijing has given the green light for the creation of China's very first'national laboratory for deep learning', in a move that could help the country to surpass the United States in developing artificial intelligence (AI). The National Development and Reform Commission (NDRC) recently approved the plan to set up a national engineering'lab' for researching and implementing deep learning technologies. The lab will not have a physical presence, instead taking the form of a research network predominantly based online. Regarded as one of the most exciting and fastest-growing areas of AI, deep learning - a subdivision of machine learning - involves feeding data through virtual neural networks designed to mimic the human brain's decision-making process, in order to solve problems and recognise images and sounds. It is seen by many as the key to elevating AI to something approximating human intelligence, and is already credited with major breakthroughs in technologies such as voice recognition in smartphones.
Exponential equity in an age of abundance: a modest proposal
Nine years ago, renowned futurist and Google Chief Engineer, Ray Kurzweil, and X-Prize Foundation Chair, Peter Diamandis, founded Singularity University in order to explore and explain the potential of exponential technologies to address our great global challenges. Last week I had the privilege to attend their Executive Programme. During our week long bootcamp, we were exposed to future technological trends in robotics, artificial intelligence, neuroscience, digital manufacturing, digital medicine and space exploration, to name but a few. Underpinning spectacular advances across all of these areas is the explosion of computational power, as predicted by Moore's Law. For example, genomes that initially cost $3bn to sequence now cost $1,000 and are on their way to cost $0.01.
EOMM: An Engagement Optimized Matchmaking Framework
Chen, Zhengxing, Xue, Su, Kolen, John, Aghdaie, Navid, Zaman, Kazi A., Sun, Yizhou, El-Nasr, Magy Seif
Matchmaking connects multiple players to participate in online player-versus-player games. Current matchmaking systems depend on a single core strategy: create fair games at all times. These systems pair similarly skilled players on the assumption that a fair game is best player experience. We will demonstrate, however, that this intuitive assumption sometimes fails and that matchmaking based on fairness is not optimal for engagement. In this paper, we propose an Engagement Optimized Matchmaking (EOMM) framework that maximizes overall player engagement. We prove that equal-skill based matchmaking is a special case of EOMM on a highly simplified assumption that rarely holds in reality. Our simulation on real data from a popular game made by Electronic Arts, Inc. (EA) supports our theoretical results, showing significant improvement in enhancing player engagement compared to existing matchmaking methods.
Preconditioned Stochastic Gradient Descent
Stochastic gradient descent (SGD) still is the workhorse for many practical problems. However, it converges slow, and can be difficult to tune. It is possible to precondition SGD to accelerate its convergence remarkably. But many attempts in this direction either aim at solving specialized problems, or result in significantly more complicated methods than SGD. This paper proposes a new method to estimate a preconditioner such that the amplitudes of perturbations of preconditioned stochastic gradient match that of the perturbations of parameters to be optimized in a way comparable to Newton method for deterministic optimization. Unlike the preconditioners based on secant equation fitting as done in deterministic quasi-Newton methods, which assume positive definite Hessian and approximate its inverse, the new preconditioner works equally well for both convex and non-convex optimizations with exact or noisy gradients. When stochastic gradient is used, it can naturally damp the gradient noise to stabilize SGD. Efficient preconditioner estimation methods are developed, and with reasonable simplifications, they are applicable to large scaled problems. Experimental results demonstrate that equipped with the new preconditioner, without any tuning effort, preconditioned SGD can efficiently solve many challenging problems like the training of a deep neural network or a recurrent neural network requiring extremely long term memories.
liquidSVM: A Fast and Versatile SVM package
Steinwart, Ingo, Thomann, Philipp
liquidSVM is a package written in C++ that provides SVM-type solvers for various classification and regression tasks. Because of a fully integrated hyper-parameter selection, very carefully implemented solvers, multi-threading and GPU support, and several built-in data decomposition strategies it provides unprecedented speed for small training sizes as well as for data sets of tens of millions of samples. Besides the C++ API and a command line interface, bindings to R, MATLAB, Java, Python, and Spark are available. We present a brief description of the package and report experimental comparisons to other SVM packages.
Scalable Inference for Nested Chinese Restaurant Process Topic Models
Chen, Jianfei, Zhu, Jun, Lu, Jie, Liu, Shixia
Nested Chinese Restaurant Process (nCRP) topic models are powerful nonparametric Bayesian methods to extract a topic hierarchy from a given text corpus, where the hierarchical structure is automatically determined by the data. Hierarchical Latent Dirichlet Allocation (hLDA) is a popular instance of nCRP topic models. However, hLDA has only been evaluated at small scale, because the existing collapsed Gibbs sampling and instantiated weight variational inference algorithms either are not scalable or sacrifice inference quality with mean-field assumptions. Moreover, an efficient distributed implementation of the data structures, such as dynamically growing count matrices and trees, is challenging. In this paper, we propose a novel partially collapsed Gibbs sampling (PCGS) algorithm, which combines the advantages of collapsed and instantiated weight algorithms to achieve good scalability as well as high model quality. An initialization strategy is presented to further improve the model quality. Finally, we propose an efficient distributed implementation of PCGS through vectorization, pre-processing, and a careful design of the concurrent data structures and communication strategy. Empirical studies show that our algorithm is 111 times more efficient than the previous open-source implementation for hLDA, with comparable or even better model quality. Our distributed implementation can extract 1,722 topics from a 131-million-document corpus with 28 billion tokens, which is 4-5 orders of magnitude larger than the previous largest corpus, with 50 machines in 7 hours.
A Unified Parallel Algorithm for Regularized Group PLS Scalable to Big Data
de Micheaux, Pierre Lafaye, Liquet, Benoit, Sutton, Matthew
Partial Least Squares (PLS) methods have been heavily exploited to analyse the association between two blocs of data. These powerful approaches can be applied to data sets where the number of variables is greater than the number of observations and in presence of high collinearity between variables. Different sparse versions of PLS have been developed to integrate multiple data sets while simultaneously selecting the contributing variables. Sparse modelling is a key factor in obtaining better estimators and identifying associations between multiple data sets. The cornerstone of the sparsity version of PLS methods is the link between the SVD of a matrix (constructed from deflated versions of the original matrices of data) and least squares minimisation in linear regression. We present here an accurate description of the most popular PLS methods, alongside their mathematical proofs. A unified algorithm is proposed to perform all four types of PLS including their regularised versions. Various approaches to decrease the computation time are offered, and we show how the whole procedure can be scalable to big data sets.