Education
5 free Data Science courses you can take online during the lockdown
The Great Learning Academy offers a course called'Introduction to R'. This course is for anyone who is a beginner and wants to understand the field of data science. R is a comprehensive statistical and graphical programming language which is fast gaining popularity among data analysts. Learners will also receive a certificate from Great learning post the completion of the program.
Robust Non-Linear Matrix Factorization for Dictionary Learning, Denoising, and Clustering
Fan, Jicong, Yang, Chengrun, Udell, Madeleine
Low dimensional nonlinear structure abounds in datasets across computer vision and machine learning. Kernelized matrix factorization techniques have recently been proposed to learn these nonlinear structures from partially observed data, with impressive empirical performance, by observing that the image of the matrix in a sufficiently large feature space is low-rank. However, these nonlinear methods fail in the presence of noise or outliers. In this work, we propose a new robust nonlinear factorization method called Robust Non-Linear Matrix Factorization (RNLMF). RNLMF constructs a dictionary for the data space by factoring a kernelized feature space; a noisy matrix can then be decomposed as the sum of a sparse noise matrix and a clean data matrix that lies in a low dimensional nonlinear manifold. RNLMF is robust to noise and outliers and scales to matrices with thousands of rows and columns. Empirically, RNLMF achieves noticeable improvements over baseline methods in denoising and clustering.
A learning problem whose consistency is equivalent to the non-existence of real-valued measurable cardinals
We show that the $k$-nearest neighbour learning rule is universally consistent in a metric space $X$ if and only if it is universally consistent in every separable subspace of $X$ and the density of $X$ is less than every real-measurable cardinal. In particular, the $k$-NN classifier is universally consistent in every metric space whose separable subspaces are sigma-finite dimensional in the sense of Nagata and Preiss if and only if there are no real-valued measurable cardinals. The latter assumption is relatively consistent with ZFC, however the consistency of the existence of such cardinals cannot be proved within ZFC. Our results were inspired by an example sketched by C\'erou and Guyader in 2006 at an intuitive level of rigour.
Vocabulary Alignment in Openly Specified Interactions
Chocron, Paula Daniela (Hutoma) | Schorlemmer, Marco
The problem of achieving common understanding between agents that use different vocabularies has been mainly addressed by techniques that assume the existence of shared external elements, such as a meta-language or a physical environment. In this article, we consider agents that use different vocabularies and only share knowledge of how to perform a task, given by the specification of an interaction protocol. We present a framework that lets agents learn a vocabulary alignment from the experience of interacting. Unlike previous work in this direction, we use open protocols that constrain possible actions instead of defining procedures, making our approach more general. We present two techniques that can be used either to learn an alignment from scratch or to repair an existent one, and we evaluate their performance experimentally.
Ensemble Learning of Coarse-Grained Molecular Dynamics Force Fields with a Kernel Approach
Wang, Jiang, Chmiela, Stefan, Mรผller, Klaus-Robert, Noรจ, Frank, Clementi, Cecilia
Gradient-domain machine learning (GDML) is an accurate and efficient approach to learn a molecular potential and associated force field based on the kernel ridge regression algorithm. Here, we demonstrate its application to learn an effective coarse-grained (CG) model from all-atom simulation data in a sample efficient manner. The coarse-grained force field is learned by following the thermodynamic consistency principle, here by minimizing the error between the predicted coarse-grained force and the all-atom mean force in the coarse-grained coordinates. Solving this problem by GDML directly is impossible because coarse-graining requires averaging over many training data points, resulting in impractical memory requirements for storing the kernel matrices. In this work, we propose a data-efficient and memory-saving alternative. Using ensemble learning and stratified sampling, we propose a 2-layer training scheme that enables GDML to learn an effective coarse-grained model. We illustrate our method on a simple biomolecular system, alanine dipeptide, by reconstructing the free energy landscape of a coarse-grained variant of this molecule. Our novel GDML training scheme yields a smaller free energy error than neural networks when the training set is small, and a comparably high accuracy when the training set is sufficiently large.
Time Efficiency in Optimization with a Bayesian-Evolutionary Algorithm
Lan, Gongjin, Tomczak, Jakub M., Roijers, Diederik M., Eiben, A. E.
Not all generate-and-test search algorithms are created equal. Bayesian Optimization (BO) invests a lot of computation time to generate the candidate solution that best balances the predicted value and the uncertainty given all previous data, taking increasingly more time as the number of evaluations performed grows. Evolutionary Algorithms (EA) on the other hand rely on search heuristics that typically do not depend on all previous data and can be done in constant time. Both the BO and EA community typically assess their performance as a function of the number of evaluations. However, this is unfair once we start to compare the efficiency of these classes of algorithms, as the overhead times to generate candidate solutions are significantly different. We suggest to measure the efficiency of generate-and-test search algorithms as the expected gain in the objective value per unit of computation time spent. We observe that the preference of an algorithm to be used can change after a number of function evaluations. We therefore propose a new algorithm, a combination of Bayesian optimization and an Evolutionary Algorithm, BEA for short, that starts with BO, then transfers knowledge to an EA, and subsequently runs the EA. We compare the BEA with BO and the EA. The results show that BEA outperforms both BO and the EA in terms of time efficiency, and ultimately leads to better performance on well-known benchmark objective functions with many local optima. Moreover, we test the three algorithms on nine test cases of robot learning problems and here again we find that BEA outperforms the other algorithms.
Microsoft introduces DigiGirlz AI Class, brings Artificial Intelligence to high-school girls
Microsoft wants high school girls around the world to develop their understanding of Artificial Intelligence. To achieve this, the company is introducing the DigiGirlz AI Class in partnership with Avanade and Accenture. The program gives girls access to a collection of tutorials, games and technological resources. Designed in a way that it equips them with the skills and knowledge to apply the technology in real life. Participants of the program will be treated to workshops led by industry experts.
UC Berkeley researchers open-source RAD to improve any reinforcement learning algorithm
In an accompanying paper, the authors say this module can improve any existing reinforcement learning algorithm and that RAD achieves better compute and data efficiency than Google AI's PlaNet, as well as recently released cutting-edge algorithms like DeepMind's Dreamer and SLAC from UC Berkeley and DeepMind. RAD achieves state-of-the-art results on common benchmarks and matches or beats every baseline in terms of performance and data efficiency across 15 DeepMind control environments, the researchers say. It does this in part by applying data augmentations for visual observations. Coauthors of the paper on RAD include Michael "Misha" Laskin, Kimin Lee, and Berkeley AI Research codirector and Covariant founder Pieter Abbeel. RAD was released Thursday on preprint repository arXiv.
Nigel Willson joins Marktechpost.com as Chief Advisory Board Member
TUSTIN, Calif., May 2, 2020 /PRNewswire-PRWeb/ -- Nigel Willson joins Marktechpost.com Nigel is a Global Speaker, Influencer, and Advisor on Artificial Intelligence and Co-founder of We and AI. He is ranked amongst the top AI Influencers in the World and as Co-Founder of We and AI (a non-government organization) is working to raise awareness of the risks and rewards of AI and helping to give humanity a voice in the age of machines. Marktechpost.com is a California-based Artificial Intelligence platform for the latest updates in machine learning, deep learning, and data science research. The theme of the platform is set in such a way that AI and Data Science professionals can share their knowledge and suggestions with the AI and Data Science aspirants.